diff --git a/internal/mcpserver/mcpserver.go b/internal/mcpserver/mcpserver.go index 8d28381..b849d9d 100644 --- a/internal/mcpserver/mcpserver.go +++ b/internal/mcpserver/mcpserver.go @@ -21,9 +21,18 @@ import ( // instructions is returned at initialize and concatenated into the agent's // system prompt. It can't carry the context bundle itself (the project isn't // known yet), so it just teaches the agent to call get_context first. -const instructions = `lard is a long-term memory layer. At the start of a session, before doing project work, call get_context with the current workspace's git origin remote (gitRemote) and absolute path (path). It returns the user profile, an index of all memory subjects, and this project's area file if identified. - -Read other subjects with memory_read when their description looks relevant. Persist durable facts the user states with memory_append (one fact) or memory_write (full rewrite; read first for the version token). Skip transient detail.` +const instructions = `lard is the user's long-term memory: markdown subject files about the user (profile), their projects (areas/), cross-cutting interests and preferences (topics/), and people. + +Read: +1. At session start, before any project work, call get_context with the workspace's git origin remote (gitRemote) and absolute path (path). It returns the profile, this project's area file, and an index of every other subject (path + description + aliases). +2. Treat what you get as established context: let the profile and area shape defaults and choices instead of asking the user again. +3. When the task plausibly touches a topic or another project in the index, memory_read that subject before working. Descriptions are retrieval keys. + +Write: +- When the user states a durable fact — a decision, convention, deploy target, host, version, preference — persist it in the same session with memory_append. One specific statement per call, routed to the right subject: project specifics → its area, cross-project preferences → a topic, identity → profile, people → people/. +- Carry the specifics: "deploys to Cloudflare Pages at s.dunkirk.sh" beats "has a deployment". +- Skip anything true only for the current task, and anything the repo already shows. +- memory_write is for full rewrites; memory_read first for the version token.` // New builds the MCP server backed by the HTTP API server's store. func New(api *httpapi.Server) *mcp.Server { diff --git a/internal/pipeline/extract.go b/internal/pipeline/extract.go index 7140432..2d50722 100644 --- a/internal/pipeline/extract.go +++ b/internal/pipeline/extract.go @@ -29,26 +29,28 @@ const extractSystem = `You extract durable facts about a user and their projects The single most important skill is telling DURABLE facts apart from EPHEMERAL task chatter. A coding session is mostly task-local ("this test is flaky", "rename foo to bar") that must NOT become memory. Only emit facts that will still be true and useful in a future session. +DURABLE MEANS CONCRETE. Memory is worth nothing if it stays vague. Keep the specifics that survive the session: repo names and URLs, domains, deploy targets and paths, hosts and ports, framework/library choices and versions, config conventions, file locations, API endpoints, credentials locations (never the secrets themselves), and the reasoning behind decisions. Prefer "deploys to Cloudflare Pages at s.dunkirk.sh" over "has a deployment"; "prefers Bun and Bun.serve over Node" over "prefers modern tooling". If the user explains WHY they chose something, fold the reason in. + SUBJECTS. Every fact belongs to exactly one subject, chosen by the NATURE of the fact (not where it was said): - kind "profile" (name "profile"): durable identity only — name, role, employer, education, location, contact, pronouns. The test: "still true in 3 months, independent of any project?" Keep this SMALL. Anything dated, "currently", or tied to one project does NOT go here. -- kind "area": one subject per project or ongoing thing. name = a short slug of the project (e.g. "crush", "battleship-arena"). Facts about what a project is, its conventions, decisions, architecture. "user is a developer of X" is an area fact for X, NOT profile. +- kind "area": one subject per project or ongoing thing. name = a short slug of the project (e.g. "crush", "battleship-arena"). Facts about what a project is, its stack, conventions, decisions, architecture, deployment, and status. "user is a developer of X" is an area fact for X, NOT profile. - kind "topic": cross-cutting domain facts spanning projects. name = the domain slug (e.g. "software-projects", "ctf-security", "frc-robotics", "hardware", "photography"). Durable preferences and skills that aren't project-specific go here (e.g. "prefers Bun over React" → software-projects). -- kind "people": one subject per person. name = a slug of their name. +- kind "people": one subject per person. name = a slug of their name. Facts about who they are and how the user works with them. You are given the EXISTING SUBJECTS (name + description + aliases). ALWAYS route a fact to an existing subject when it fits — match on name or alias — instead of inventing a near-duplicate. Only create a new subject when nothing fits; then give it a short "description" (one line: what it covers) and optional "aliases". Rules: - Only facts about the USER or the PROJECT, grounded in the user's own words. Never about the assistant. -- Durable only. Skip anything true only for this task or hour. -- One clear statement per fact; group tightly-related details into one fact rather than splitting hairs. -- Prefer durable phrasing over specifics that go stale ("meeting-heavy mornings" > "10:00 standup"). +- Durable only. Skip anything true only for this task or hour — bug fixes in progress, one-off commands, "currently debugging X". +- One clear statement per fact, but group tightly-related details into ONE fact rather than splitting hairs ("uses Bun, Drizzle, and Lit" is one stack fact, not three). +- Prefer durable phrasing over specifics that go stale ("meeting-heavy mornings" > "10:00 standup"), but never use that as an excuse to drop load-bearing specifics. - Calibrate to evidence: one mention → "mentioned X once", not "X expert". - Tag sensitivity: if a fact touches race, ethnicity, religion, sexual orientation, gender identity, immigration, disability, illness/health, politics, sexual history, abuse, finances, criminal history, real-time location, biometric data, or a date of birth, set "sensitivity" to that category. It will be dropped. - Provenance "tag" is always "stated" here (these are the user's own words). - Keep facts descriptive, never prescriptive instructions that would suppress honest feedback. Output a JSON array (no prose) of objects: -{"text": "the durable fact", +{"text": "the durable fact, with its specifics", "subjectKind": "profile" | "area" | "topic" | "people", "subjectName": "slug", "description": "one-line subject description (only if creating a new subject)", diff --git a/internal/pipeline/synthesize.go b/internal/pipeline/synthesize.go index 9d4ecfd..a361ef4 100644 --- a/internal/pipeline/synthesize.go +++ b/internal/pipeline/synthesize.go @@ -9,16 +9,18 @@ import ( "github.com/taciturnaxolotl/lard/internal/types" ) -const synthesizeSystem = `You maintain one memory file about a single subject. You are given the subject's kind, its current file body (may be empty), and a list of facts gathered from the user's sessions. Rewrite the file body as clean, durable prose. +const synthesizeSystem = `You maintain one memory file about a single subject. You are given the subject's kind, its current file body (may be empty), and a list of facts gathered from the user's sessions. Rewrite the file body so a future assistant reading it knows the subject without any other context. Rules: -- Produce a tight set of bullet points, each a durable fact. Merge related facts into single coherent bullets; do not just concatenate. +- Produce a tight set of bullet points, each a self-contained fact. Merge related facts into single coherent bullets; do not just concatenate. +- Keep the specifics. Names, URLs, domains, hosts, paths, versions, stack components, deploy targets, and the reasoning behind decisions are the value of memory — merging bullets must never mean stripping them out. "Cloudflare Workers-based PWA deployed on Cloudflare Pages at s.dunkirk.sh" is the right level; "a web project" is not. - Preserve everything in the current body that is not contradicted; integrate the new facts. This is a revise-in-place, not a fresh write — respect prior content (it may include the user's own edits). - When new facts supersede old ones (a project changed stack, a role changed), replace the stale statement rather than keeping both. -- Prefer durable phrasing over specifics that go stale. Drop task-local noise that slipped through. +- Prefer durable phrasing over specifics that go stale, but never use that as an excuse to drop load-bearing specifics. Drop task-local noise that slipped through. +- For an "area" subject, aim to cover, when known: what the project is, its stack and key libraries, where it runs/deploys, conventions and decisions (with reasons where given), and current goals or open threads. - Every bullet starts with a provenance tag in brackets: [stated] (user said it), [observed], or [inferred]. Use [stated] unless the input fact says otherwise. - For a "profile" subject: keep ONLY durable identity (name, role, education, location, contact, pronouns). Move anything project-specific out (omit it — it lives elsewhere). -- Keep it concise. A subject file is an overview, not a transcript. Aim for the smallest set of bullets that captures the durable truth. +- Keep it concise — a memory file, not a transcript — but favor completeness over brevity when they conflict. A rich 15-bullet subject beats a vague 4-bullet one. Output ONLY the file body (the bullet lines). No frontmatter, no headings, no preamble.`