diff --git a/.agents/skills/observing-misaligned-landings/SKILL.md b/.agents/skills/observing-misaligned-landings/SKILL.md index ad94a371..59466389 100644 --- a/.agents/skills/observing-misaligned-landings/SKILL.md +++ b/.agents/skills/observing-misaligned-landings/SKILL.md @@ -72,35 +72,49 @@ The checks are deliberately stricter than a successful push: The top-level deterministic state is `ready`, `blocked`, `landed`, or `inconsistent`. Exit `0` means `ready` or `landed`; `blocked` exits `1`, and `inconsistent` exits `2`. Unknown or unreadable evidence is preserved and fails -closed instead of being collapsed into absence. A dirty primary checkout is an -ownership warning unless it touches the observed landing surface; a later clean -landing is reported as superseding the candidate instead of contradicting it. +closed instead of being collapsed into absence. Before landing, any dirty primary +checkout blocks the existing landing path until its owner resolves it. Candidate paths +carry an explicit merge-base-to-candidate scope while that delta remains provable; +once current main contains the candidate, the original delta becomes unknown rather +than being mislabeled empty. Absent or different scope fails closed instead of proving +dirt disjoint. NUL-delimited Git evidence compares both literal +sides of renames and copies, enumerates every untracked descendant, and treats either +path-component ancestor direction as overlap. Git path bytes survive UTF-8 decoding +through surrogate escape. After exact remote, gate, and cleanup proof establishes a +completed landing, later primary dirt remains an ownership warning rather than +rewriting history. A later clean descendant is reported as superseding the candidate +instead of contradicting it. ## Ask Trace only about anomalies Deterministic checks decide classification. Agent interpretation is optional, supplemental, and only when non-benign anomalies remain. Ordinary public-edge -convergence and explicitly superseded landings do not invoke it by themselves. +convergence and `current_main_unpublished` on an already proved historical landing do +not invoke it by themselves. -Install the pinned SDK once for this checked-in skill, then pass the receipt: +Install the locked dependencies with the package-manager version named by this skill, +then pass the receipt: ```bash -pnpm --dir "$SKILL" install --frozen-lockfile +corepack pnpm@10.20.0 --dir "$SKILL" install --frozen-lockfile node "$SKILL/scripts/interpret-anomalies.mjs" /tmp/misaligned-landing.json ``` The interpreter uses `@letta-ai/letta-agent-sdk` to find exactly one retained `letta/auto` specialist named `Misaligned Landing Observer`, provisioning it only when -absent. More than one exact-name match fails closed. The specialist stays discoverable -for exact-name reuse, with empty memory, MemFS disabled, no tools or skills, and a strict -zero-action prompt. +absent. More than one exact-name match fails closed. The specialist is hidden from +ordinary agent surfaces but remains available to exact-name API lookup, with empty +memory, MemFS disabled, no tools or skills, and a strict zero-action prompt. Each observation uses SDK `prompt()` with `stateless: true`, which opens a fresh conversation for that retained agent rather than resuming an earlier receipt and loads no persisted MemFS state. Only minimized evidence crosses the repository boundary: local paths, remote URLs, task paths, and raw read errors are removed. Interpretation may explain likely convergence, stale evidence, or project-operation inconsistency, but cannot alter the receipt or its deterministic -state. +state. The pinned transport contract resolves one explicit organization computer, +opens separate control and stream sockets, starts a strict stateless runtime with no +skills, synchronizes it, and sends the turn with empty client-tool authority and +interactive tools excluded. If no non-benign anomaly remains, the interpreter returns `status: not-needed` without loading the SDK or contacting Letta. If credentials, SDK installation, exact @@ -111,9 +125,14 @@ separately; never downgrade or upgrade the deterministic landing result. ```bash python3 tools/test_landing_observer.py -node --test .agents/skills/observing-misaligned-landings/tests/*.test.mjs +corepack pnpm@10.20.0 --dir .agents/skills/observing-misaligned-landings install --frozen-lockfile +corepack pnpm@10.20.0 --dir .agents/skills/observing-misaligned-landings test ``` -The fixture proves byte-stable clean receipts, exact anomaly codes, optional-SDK -fail-closed behavior, and that observing leaves the target repository's Git bytes -unchanged. +The fixture proves byte-stable clean receipts, exact anomaly codes and dirty-path +ownership, complete Telegram status/message-id semantics, a clean descendant after +candidate cleanup, optional-SDK fail-closed behavior, the real pinned SDK's +REST/WebSocket request shape through a hermetic fake transport, and that observing +leaves the target repository's Git bytes unchanged. The project gate performs the frozen install before +the Node suite; missing dependencies cannot silently turn transport coverage into a +skip. diff --git a/.agents/skills/observing-misaligned-landings/package.json b/.agents/skills/observing-misaligned-landings/package.json index 935e69b9..31090946 100644 --- a/.agents/skills/observing-misaligned-landings/package.json +++ b/.agents/skills/observing-misaligned-landings/package.json @@ -9,6 +9,10 @@ "test": "node --test tests/*.test.mjs" }, "dependencies": { - "@letta-ai/letta-agent-sdk": "0.7.1" + "@letta-ai/letta-agent-sdk": "0.7.1", + "@letta-ai/letta-client": "1.12.1" + }, + "devDependencies": { + "ws": "8.21.3" } } diff --git a/.agents/skills/observing-misaligned-landings/pnpm-lock.yaml b/.agents/skills/observing-misaligned-landings/pnpm-lock.yaml index 69448e78..7aded413 100644 --- a/.agents/skills/observing-misaligned-landings/pnpm-lock.yaml +++ b/.agents/skills/observing-misaligned-landings/pnpm-lock.yaml @@ -11,6 +11,13 @@ importers: '@letta-ai/letta-agent-sdk': specifier: 0.7.1 version: 0.7.1(ink@7.1.1(react@18.2.0))(react-dom@19.2.8(react@18.2.0))(zod@4.4.3) + '@letta-ai/letta-client': + specifier: 1.12.1 + version: 1.12.1 + devDependencies: + ws: + specifier: 8.21.3 + version: 8.21.3 packages: diff --git a/.agents/skills/observing-misaligned-landings/scripts/interpret-anomalies.mjs b/.agents/skills/observing-misaligned-landings/scripts/interpret-anomalies.mjs index 90ebac13..76fd6649 100755 --- a/.agents/skills/observing-misaligned-landings/scripts/interpret-anomalies.mjs +++ b/.agents/skills/observing-misaligned-landings/scripts/interpret-anomalies.mjs @@ -10,7 +10,10 @@ export const INTERPRETATION_SCHEMA = export const MODEL = "letta/auto"; export const SPECIALIST_NAME = "Misaligned Landing Observer"; const MAX_RECEIPT_BYTES = 256 * 1024; -const BENIGN_ANOMALY_CODES = new Set(["public_edge_converging"]); +const BENIGN_ANOMALY_CODES = new Set([ + "current_main_unpublished", + "public_edge_converging", +]); const INTERPRETER_SYSTEM_PROMPT = [ "You are the retained Misaligned Landing Observer interpreting exactly one minimized deterministic landing receipt in a fresh stateless conversation.", "You have no tools, skills, repository access, or authority to mutate anything.", @@ -69,8 +72,9 @@ export function validateReceipt(value) { } export function needsInterpretation(receipt) { - return validateReceipt(receipt).anomalies.some( - (row) => !BENIGN_ANOMALY_CODES.has(row.code), + const value = validateReceipt(receipt); + return value.anomalies.some( + (row) => value.state !== "landed" || !BENIGN_ANOMALY_CODES.has(row.code), ); } @@ -209,8 +213,13 @@ function agentId(agent) { : null; } -export async function findOrCreateSpecialist(client) { - const listed = await client.agents.list({ name: SPECIALIST_NAME, limit: 100 }); +export async function findOrCreateSpecialist(client, lookup = client.agents) { + const page = await lookup.list({ + limit: 100, + name: SPECIALIST_NAME, + show_hidden_agents: true, + }); + const listed = Array.isArray(page) ? page : page?.items; if (!Array.isArray(listed)) { throw new Error("Letta agent lookup returned an unsupported result"); } @@ -229,6 +238,7 @@ export async function findOrCreateSpecialist(client) { baseTools: [], description: "Retained zero-tool specialist for interpreting deterministic Misaligned landing anomalies", + hidden: true, memfs: false, memory: [], model: MODEL, @@ -244,7 +254,13 @@ export async function findOrCreateSpecialist(client) { export async function interpret( receipt, env = process.env, - sdkLoader = () => import("@letta-ai/letta-agent-sdk"), + sdkLoader = async () => { + const [{ LettaAgentClient }, { default: Letta }] = await Promise.all([ + import("@letta-ai/letta-agent-sdk"), + import("@letta-ai/letta-client"), + ]); + return { Letta, LettaAgentClient }; + }, ) { const validated = validateReceipt(receipt); if (!needsInterpretation(validated)) { @@ -254,19 +270,24 @@ export async function interpret( reason: validated.anomalies.length === 0 ? "deterministic receipt has no anomalies" - : "deterministic receipt has only benign public-edge convergence", + : "deterministic receipt has only benign publication-state anomalies", schema: INTERPRETATION_SCHEMA, status: "not-needed", }; } - const { LettaAgentClient } = await sdkLoader(); + const { Letta, LettaAgentClient } = await sdkLoader(); + const apiKey = env.LETTA_API_KEY ?? env.LETTA_CLOUD_API_KEY; const clientOptions = { backend: "cloud", - ...(env.LETTA_API_KEY ? { apiKey: env.LETTA_API_KEY } : {}), + ...(apiKey ? { apiKey } : {}), ...(env.LETTA_BASE_URL ? { apiBaseUrl: env.LETTA_BASE_URL } : {}), }; const client = new LettaAgentClient(clientOptions); - const specialist = await findOrCreateSpecialist(client); + const lookup = new Letta({ + ...(apiKey ? { apiKey } : {}), + ...(env.LETTA_BASE_URL ? { baseURL: env.LETTA_BASE_URL } : {}), + }); + const specialist = await findOrCreateSpecialist(client, lookup.agents); const result = await client.prompt(buildPrompt(validated), specialist.agentId, { allowedTools: [], permissionMode: "strict", diff --git a/.agents/skills/observing-misaligned-landings/scripts/observe-landing.py b/.agents/skills/observing-misaligned-landings/scripts/observe-landing.py index ce837e4d..2b4cc65e 100755 --- a/.agents/skills/observing-misaligned-landings/scripts/observe-landing.py +++ b/.agents/skills/observing-misaligned-landings/scripts/observe-landing.py @@ -25,6 +25,7 @@ TITLE = "Misaligned — WORK / THINK / LIE" DEFAULT_SITE_URL = "https://cameron.tngl.io/misaligned/" SHA_RE = re.compile(r"^[0-9a-f]{40}$") SOURCE_OUTPUT_RE = re.compile(r"(?:source=)([0-9a-f]{40}|missing)") +CHANGED_PATHS_SCOPE = "merge-base-to-candidate" @dataclass(frozen=True) @@ -85,8 +86,9 @@ def command_runner( completed = subprocess.run( command, cwd=cwd, + encoding="utf-8", env=env or read_only_env(), - text=True, + errors="surrogateescape", stdout=subprocess.PIPE, stderr=subprocess.PIPE, timeout=timeout, @@ -163,10 +165,75 @@ def public_source(output: str) -> str | None: return None -def list_output(result: CommandResult) -> list[str]: - if result.returncode != 0: +def nul_records(output: str, description: str) -> list[str]: + if not output: return [] - return sorted(line for line in result.stdout.splitlines() if line) + records = output.split("\0") + if records[-1] != "": + raise ValueError(f"{description} is not NUL terminated") + records.pop() + return records + + +def porcelain_v1_z_paths(output: str) -> list[str]: + """Return every literal path named by `git status --porcelain=v1 -z`.""" + records = nul_records(output, "porcelain status") + paths: set[str] = set() + index = 0 + while index < len(records): + record = records[index] + if len(record) < 4 or record[2] != " " or not record[3:]: + raise ValueError("porcelain status has a malformed entry") + status = record[:2] + paths.add(record[3:]) + index += 1 + if "R" in status or "C" in status: + if index >= len(records) or not records[index]: + raise ValueError("porcelain rename/copy entry is missing its source path") + paths.add(records[index]) + index += 1 + return sorted(paths) + + +def diff_name_status_z_paths(output: str) -> list[str]: + """Return both literal path sides named by `git diff --name-status -z`.""" + records = nul_records(output, "name-status diff") + paths: set[str] = set() + index = 0 + while index < len(records): + status = records[index] + index += 1 + if not status or status[0] not in "ACDMRTUXB": + raise ValueError("name-status diff has a malformed status") + path_count = 2 if status[0] in {"R", "C"} else 1 + if index + path_count > len(records): + raise ValueError("name-status diff is missing a path") + for path in records[index : index + path_count]: + if not path: + raise ValueError("name-status diff has an empty path") + paths.add(path) + index += path_count + return sorted(paths) + + +def overlapping_candidate_paths(touching: list[str], changed: list[str]) -> list[str]: + """Return candidate paths equal to, above, or beneath a dirty Git path.""" + conflicts: set[str] = set() + for dirty in touching: + dirty_root = dirty.rstrip("/") + if not dirty_root: + continue + for candidate in changed: + candidate_root = candidate.rstrip("/") + if not candidate_root: + continue + if ( + dirty_root == candidate_root + or dirty_root.startswith(f"{candidate_root}/") + or candidate_root.startswith(f"{dirty_root}/") + ): + conflicts.add(candidate) + return sorted(conflicts) def count_pair(result: CommandResult) -> tuple[int | None, int | None]: @@ -228,8 +295,26 @@ def live_worktree(runner: Runner, path: Path | None) -> dict[str, object]: ) branch = git(runner, canonical, "branch", "--show-current") evidence["branch"] = branch.stdout.strip() if branch.returncode == 0 else None - status = git(runner, canonical, "status", "--porcelain", "--untracked-files=normal") - evidence["dirty"] = bool(status.stdout) if status.returncode == 0 else None + status = git( + runner, + canonical, + "status", + "--porcelain=v1", + "-z", + "--untracked-files=all", + ) + if status.returncode == 0: + try: + touching = porcelain_v1_z_paths(status.stdout) + except ValueError: + evidence["dirty"] = None + evidence["touching"] = None + else: + evidence["dirty"] = bool(touching) + evidence["touching"] = touching + else: + evidence["dirty"] = None + evidence["touching"] = None upstream = git(runner, canonical, "rev-parse", "--abbrev-ref", "@{upstream}") evidence["upstream"] = upstream.stdout.strip() if upstream.returncode == 0 else None return evidence @@ -377,6 +462,7 @@ def isolated_remote_snapshot( """Read remote history and page bytes in a disposable bare repository.""" snapshot: dict[str, object] = { "changed_paths": None, + "changed_paths_scope": None, "main_contains_expected": None, "main_read_error": None, "main_revision": None, @@ -532,16 +618,27 @@ def isolated_remote_snapshot( if behind is None or ahead is None: snapshot["main_read_error"] = clean_line(counts) return snapshot + # Once main contains the candidate, current history no longer proves the + # candidate's original pre-landing delta. Preserve that scope as unknown + # rather than mislabeling an empty candidate-side diff as exact evidence. + if snapshot["main_contains_expected"] is not False: + return snapshot changed = git( runner, tmp, "diff", - "--name-only", + "--name-status", + "-z", f"refs/observer/main...{candidate_ref}", + "--", ) - snapshot["changed_paths"] = ( - list_output(changed) if changed.returncode == 0 else None - ) + if changed.returncode == 0: + try: + snapshot["changed_paths"] = diff_name_status_z_paths(changed.stdout) + except ValueError: + snapshot["changed_paths"] = None + else: + snapshot["changed_paths_scope"] = CHANGED_PATHS_SCOPE return snapshot @@ -595,11 +692,17 @@ def read_receipt(path: Path | None, kind: str) -> dict[str, object]: channel = payload.get("channel") revision = payload.get("revision") raw_message_id = payload.get("message_id") - message_id = str(raw_message_id) if isinstance(raw_message_id, (str, int)) else None + message_id = ( + str(raw_message_id) + if isinstance(raw_message_id, (str, int)) and not isinstance(raw_message_id, bool) + else None + ) revision_valid = isinstance(revision, str) and bool(SHA_RE.fullmatch(revision)) message_valid = ( - status != "delivered" - or (isinstance(message_id, str) and message_id.isdigit()) + isinstance(message_id, str) + and bool(re.fullmatch(r"[1-9][0-9]*", message_id)) + if status == "delivered" + else raw_message_id is None ) valid = ( channel == "telegram" @@ -759,6 +862,40 @@ def classify( elif not owner_matches: append_once(ownership, "task_ownership_unproven", "the current agent cannot be matched to the target run owner") + primary_dirty_blocker: str | None = None + if primary.get("dirty") is True: + touching = primary.get("touching") + changed_paths = repository.get("changed_paths") + paths_are_exact = ( + isinstance(touching, list) + and all(isinstance(path, str) and bool(path) for path in touching) + and isinstance(changed_paths, list) + and all(isinstance(path, str) and bool(path) for path in changed_paths) + and repository.get("changed_paths_scope") == CHANGED_PATHS_SCOPE + ) + overlap = overlapping_candidate_paths(touching, changed_paths) if paths_are_exact else [] + if overlap: + primary_dirty_blocker = "primary_checkout_candidate_path_overlap" + append_once( + ownership, + "primary_checkout_candidate_path_overlap", + "the primary checkout has dirt on paths changed by the candidate; ownership must be resolved before landing", + ) + elif paths_are_exact: + primary_dirty_blocker = "primary_checkout_dirty" + append_once( + ownership, + "unrelated_primary_checkout_dirty", + "the primary checkout is dirty on paths outside the candidate delta", + ) + else: + primary_dirty_blocker = "primary_checkout_dirty_paths_unproven" + append_once( + ownership, + "primary_checkout_dirty_paths_unproven", + "the primary checkout is dirty, but exact path ownership could not be compared with the candidate delta", + ) + cleanup_complete = run_status == "ok" and target.get("present") is False exact_remote = repository.get("remote_main_is_expected") is True current_pages = ( @@ -779,12 +916,6 @@ def classify( ) if authoritative_landing and cleanup_complete and exact_gate_pass: - if primary.get("dirty") is True: - append_once( - ownership, - "unrelated_primary_checkout_dirty", - "the current primary checkout is dirty, but immutable remote and cleanup proof already establish this landing", - ) if superseded_landing and ( pages.get("snapshot_matches_ref") is not True or pages.get("title_contract") is not True @@ -809,7 +940,7 @@ def classify( ) if channel.get("valid") is True and channel.get("status") == "delivered": pass - elif channel.get("valid") is True and channel.get("status") in {"unavailable", "failed"}: + elif channel.get("valid") is True and channel.get("status") == "unavailable": append_once( missing, "telegram_delivery_pending", @@ -818,7 +949,24 @@ def classify( append_once( anomalies, "telegram_runtime_unavailable", - "the supplied channel receipt does not prove outbound Telegram delivery", + "the supplied channel receipt explicitly records an unavailable Telegram runtime", + ) + elif channel.get("valid") is True and channel.get("status") == "failed": + append_once( + missing, + "telegram_delivery_pending", + "retry or otherwise complete user-visible Telegram delivery", + ) + append_once( + anomalies, + "telegram_delivery_failed", + "the supplied channel receipt explicitly records a failed Telegram delivery attempt", + ) + elif channel.get("valid") is True and channel.get("status") == "unknown": + append_once( + missing, + "telegram_delivery_status_unknown", + "determine whether Telegram delivery occurred and retain a real message id if it did", ) else: append_once( @@ -876,13 +1024,6 @@ def classify( f"Remote main contains {exact_revision[:12] if exact_revision else 'the candidate'}, but its completed landing receipt is incomplete.", ) - if primary.get("dirty") is True: - append_once( - ownership, - "unrelated_primary_checkout_dirty", - "the primary checkout is dirty; land from the owned worktree without disturbing it", - ) - ready_conditions: list[tuple[str, bool]] = [] target_present = target.get("present") is True ready_conditions.append(("target_worktree_missing", target_present)) @@ -910,6 +1051,8 @@ def classify( ready_conditions.append(("gate_receipt_invalid", False)) elif gate.get("status") != "pass": ready_conditions.append(("landing_gate_failed", False)) + if primary_dirty_blocker is not None: + ready_conditions.append((primary_dirty_blocker, False)) summaries = { "target_worktree_missing": "restore or identify the owned target worktree", "candidate_head_mismatch": "make the target worktree name the expected candidate", @@ -922,6 +1065,9 @@ def classify( "gate_receipt_missing": "supply an exact-revision landing-gate receipt", "gate_receipt_invalid": "replace the malformed gate receipt", "landing_gate_failed": "run and pass the landing gate on the exact candidate", + "primary_checkout_candidate_path_overlap": "resolve primary-checkout ownership on paths changed by the candidate before using the existing landing path", + "primary_checkout_dirty": "have the owner resolve unrelated primary-checkout dirt before using the existing landing path", + "primary_checkout_dirty_paths_unproven": "recover exact primary dirty-path evidence and resolve its ownership before using the existing landing path", "generated_work_orders_stale": "regenerate the work-order projection", "generated_ledgers_stale": "regenerate the corpus ledger indexes", } @@ -1020,6 +1166,7 @@ def observe( "candidate_head_stable": None, "hinted_root_matches": hinted_root_matches, "changed_paths": None, + "changed_paths_scope": None, "local_main_revision": None, "read_error": None, "readable": root_result.returncode == 0, @@ -1076,6 +1223,7 @@ def observe( ssh_command = ssh_result.stdout.strip() if ssh_result.returncode == 0 else None snapshot: dict[str, object] = { "changed_paths": None, + "changed_paths_scope": None, "main_contains_expected": None, "main_read_error": None, "main_revision": None, @@ -1118,6 +1266,7 @@ def observe( ) changed_paths = snapshot.get("changed_paths") repository["changed_paths"] = changed_paths if isinstance(changed_paths, list) else None + repository["changed_paths_scope"] = snapshot.get("changed_paths_scope") target["behind"] = snapshot.get("target_behind") target["ahead"] = snapshot.get("target_ahead") diff --git a/.agents/skills/observing-misaligned-landings/tests/interpret-anomalies.test.mjs b/.agents/skills/observing-misaligned-landings/tests/interpret-anomalies.test.mjs index 789a5859..894555ad 100644 --- a/.agents/skills/observing-misaligned-landings/tests/interpret-anomalies.test.mjs +++ b/.agents/skills/observing-misaligned-landings/tests/interpret-anomalies.test.mjs @@ -74,13 +74,30 @@ test("receipt validation rejects non-object evidence", () => { test("only non-benign anomalies need interpretation", () => { const clean = receipt({ state: "landed", anomalies: [] }); assert.equal(needsInterpretation(clean), false); - const converging = receipt({ - state: "landed", - anomalies: [ - { code: "public_edge_converging", summary: "The public edge is converging." }, - ], - }); - assert.equal(needsInterpretation(converging), false); + for (const benign of [ + { code: "public_edge_converging", summary: "The public edge is converging." }, + { + code: "current_main_unpublished", + summary: "The historical candidate is landed while newer main awaits publication.", + }, + ]) { + const landed = receipt({ state: "landed", anomalies: [benign] }); + assert.equal(needsInterpretation(landed), false); + } + assert.equal( + needsInterpretation( + receipt({ + state: "inconsistent", + anomalies: [ + { + code: "current_main_unpublished", + summary: "A contradictory receipt reused a landed-only anomaly code.", + }, + ], + }), + ), + true, + ); assert.equal(needsInterpretation(receipt()), true); }); @@ -129,7 +146,12 @@ test("specialist lookup reuses the exact retained name", async () => { agentId: "agent-exact", created: false, }); - assert.deepEqual(calls, [["list", { name: SPECIALIST_NAME, limit: 100 }]]); + assert.deepEqual(calls, [ + [ + "list", + { limit: 100, name: SPECIALIST_NAME, show_hidden_agents: true }, + ], + ]); }); test("specialist creation is empty-memory, zero-tool, and persistent", async () => { @@ -154,12 +176,13 @@ test("specialist creation is empty-memory, zero-tool, and persistent", async () assert.equal(options.name, SPECIALIST_NAME); assert.equal(options.model, MODEL); assert.equal(options.memfs, false); - assert.equal(options.hidden, undefined); + assert.equal(options.hidden, true); assert.deepEqual(options.memory, []); assert.deepEqual(options.baseTools, []); assert.deepEqual(Object.keys(options).sort(), [ "baseTools", "description", + "hidden", "memfs", "memory", "model", @@ -169,14 +192,17 @@ test("specialist creation is empty-memory, zero-tool, and persistent", async () assert.equal(calls.some(([kind]) => kind === "delete"), false); }); -test("installed SDK translates specialist creation to an empty-memory zero-tool agent", async (t) => { +test("installed SDK finds hidden specialists and translates zero-tool creation", async (t) => { + let Letta; let LettaAgentClient; try { ({ LettaAgentClient } = await import("@letta-ai/letta-agent-sdk")); + ({ default: Letta } = await import("@letta-ai/letta-client")); } catch (error) { if ( error?.code === "ERR_MODULE_NOT_FOUND" && - String(error.message).includes("@letta-ai/letta-agent-sdk") + (String(error.message).includes("@letta-ai/letta-agent-sdk") || + String(error.message).includes("@letta-ai/letta-client")) ) { t.skip("pinned SDK dependencies are not installed in this checkout"); return; @@ -184,17 +210,23 @@ test("installed SDK translates specialist creation to an empty-memory zero-tool throw error; } - let observedRequest = null; + const observedRequests = []; const server = http.createServer((request, response) => { const chunks = []; request.on("data", (chunk) => chunks.push(chunk)); request.on("end", () => { - observedRequest = { - body: JSON.parse(Buffer.concat(chunks).toString("utf8")), + const bytes = Buffer.concat(chunks); + observedRequests.push({ + authorization: request.headers.authorization, + body: bytes.length > 0 ? JSON.parse(bytes.toString("utf8")) : null, method: request.method, url: request.url, - }; + }); response.writeHead(200, { "content-type": "application/json" }); + if (request.method === "GET") { + response.end("[]"); + return; + } response.end(JSON.stringify({ id: "agent-sdk-contract", name: SPECIALIST_NAME })); }); }); @@ -208,27 +240,263 @@ test("installed SDK translates specialist creation to an empty-memory zero-tool apiKey: "hermetic-contract-key", backend: "cloud", }); - const specialist = await findOrCreateSpecialist({ - agents: { list: async () => [] }, - createAgent: (options) => client.createAgent(options), + const lookup = new Letta({ + apiKey: "hermetic-contract-key", + baseURL: `http://127.0.0.1:${address.port}`, }); + const specialist = await findOrCreateSpecialist(client, lookup.agents); assert.deepEqual(specialist, { agentId: "agent-sdk-contract", created: true }); } finally { await new Promise((resolve) => server.close(resolve)); } - assert.ok(observedRequest); - assert.equal(observedRequest.method, "POST"); - assert.equal(observedRequest.url, "/v1/agents/"); - assert.equal(observedRequest.body.name, SPECIALIST_NAME); - assert.equal(observedRequest.body.model, MODEL); - assert.equal(observedRequest.body.hidden, undefined); - assert.deepEqual(observedRequest.body.memory_blocks, []); - assert.deepEqual(observedRequest.body.tools, []); - assert.equal(observedRequest.body.include_base_tools, false); - assert.equal(observedRequest.body.include_base_tool_rules, false); - assert.deepEqual(observedRequest.body.tags, ["origin:letta-code"]); - assert.match(observedRequest.body.system, /no tools, skills, repository access/); + assert.equal(observedRequests.length, 2); + const [lookupRequest, createRequest] = observedRequests; + assert.equal(lookupRequest.method, "GET"); + const lookupUrl = new URL(lookupRequest.url, "http://127.0.0.1"); + assert.equal(lookupUrl.pathname, "/v1/agents/"); + assert.equal(lookupUrl.searchParams.get("limit"), "100"); + assert.equal(lookupUrl.searchParams.get("name"), SPECIALIST_NAME); + assert.equal(lookupUrl.searchParams.get("show_hidden_agents"), "true"); + assert.equal(createRequest.method, "POST"); + assert.equal(createRequest.url, "/v1/agents/"); + assert.equal(createRequest.body.name, SPECIALIST_NAME); + assert.equal(createRequest.body.model, MODEL); + assert.equal(createRequest.body.hidden, true); + assert.deepEqual(createRequest.body.memory_blocks, []); + assert.deepEqual(createRequest.body.tools, []); + assert.equal(createRequest.body.include_base_tools, false); + assert.equal(createRequest.body.include_base_tool_rules, false); + assert.deepEqual(createRequest.body.tags, ["origin:letta-code"]); + assert.match(createRequest.body.system, /no tools, skills, repository access/); + for (const request of observedRequests) { + assert.equal(request.authorization, "Bearer hermetic-contract-key"); + } +}); + +test("installed SDK carries the zero-authority prompt through Cloud transport", { timeout: 10_000 }, async (t) => { + let LettaAgentClient; + let WebSocket; + let WebSocketServer; + try { + ({ LettaAgentClient } = await import("@letta-ai/letta-agent-sdk")); + ({ WebSocket, WebSocketServer } = await import("ws")); + } catch (error) { + if ( + error?.code === "ERR_MODULE_NOT_FOUND" && + (String(error.message).includes("@letta-ai/letta-agent-sdk") || + String(error.message).includes("package 'ws'")) + ) { + t.skip("pinned SDK test dependencies are not installed in this checkout"); + return; + } + throw error; + } + + const agentId = "agent-sdk-transport"; + const conversationId = "conv-sdk-transport"; + const connectionId = "connection-sdk-transport"; + const connectionName = "observer-computer"; + const runtime = { + agent_id: agentId, + conversation_id: conversationId, + }; + const restRequests = []; + const socketRequests = []; + const frames = { control: [], stream: [] }; + let streamSocket = null; + + const server = http.createServer((request, response) => { + const chunks = []; + request.on("data", (chunk) => chunks.push(chunk)); + request.on("end", () => { + const bodyBytes = Buffer.concat(chunks); + restRequests.push({ + body: bodyBytes.length > 0 ? JSON.parse(bodyBytes.toString("utf8")) : null, + method: request.method, + url: request.url, + }); + response.setHeader("content-type", "application/json"); + if ( + request.method === "POST" && + request.url === `/v1/conversations/?agent_id=${agentId}` + ) { + response.end(JSON.stringify({ id: conversationId, agent_id: agentId })); + return; + } + if (request.method === "GET" && request.url === "/v1/environments?limit=100") { + response.end( + JSON.stringify({ + connections: [ + { + id: "environment-decoy", + connectedAt: 1, + connectionId: "connection-decoy", + connectionName: "another-computer", + deviceId: "device-decoy", + firstSeenAt: 1, + lastHeartbeat: 1, + lastSeenAt: 1, + organizationId: "organization-sdk-transport", + podId: "pod-decoy", + }, + { + id: "environment-sdk-transport", + connectedAt: 1, + connectionId, + connectionName, + deviceId: "device-sdk-transport", + firstSeenAt: 1, + lastHeartbeat: 1, + lastSeenAt: 1, + organizationId: "organization-sdk-transport", + podId: "pod-sdk-transport", + }, + ], + hasNextPage: false, + }), + ); + return; + } + response.statusCode = 404; + response.end(JSON.stringify({ error: `unexpected request: ${request.method} ${request.url}` })); + }); + }); + const sockets = new WebSocketServer({ noServer: true }); + sockets.on("connection", (socket, request) => { + const url = new URL(request.url, "http://127.0.0.1"); + const channel = url.searchParams.get("channel"); + assert.ok(channel === "control" || channel === "stream"); + socketRequests.push({ + authorization: request.headers.authorization, + channel, + url, + }); + if (channel === "stream") streamSocket = socket; + socket.on("message", (bytes) => { + const message = JSON.parse(bytes.toString()); + frames[channel].push(message); + if (channel !== "control") return; + if (message.type === "runtime_start") { + socket.send( + JSON.stringify({ + agent: { id: agentId, model: MODEL, model_settings: null, tools: [] }, + conversation: { id: conversationId, agent_id: agentId }, + request_id: message.request_id, + runtime, + success: true, + type: "runtime_start_response", + }), + ); + return; + } + if (message.type === "input") { + assert.ok(streamSocket); + streamSocket.send( + JSON.stringify({ + delta: { + content: "Pinned transport diagnosis.", + id: "message-sdk-transport", + message_type: "assistant_message", + run_id: "run-sdk-transport", + }, + runtime, + type: "stream_delta", + }), + ); + streamSocket.send( + JSON.stringify({ + delta: { + message_type: "stop_reason", + run_id: "run-sdk-transport", + stop_reason: "end_turn", + }, + runtime, + type: "stream_delta", + }), + ); + } + }); + }); + server.on("upgrade", (request, socket, head) => { + sockets.handleUpgrade(request, socket, head, (webSocket) => { + sockets.emit("connection", webSocket, request); + }); + }); + await new Promise((resolve) => server.listen(0, "127.0.0.1", resolve)); + const address = server.address(); + assert.ok(address && typeof address === "object"); + + let result; + try { + const client = new LettaAgentClient({ + WebSocket, + apiBaseUrl: `http://127.0.0.1:${address.port}`, + apiKey: "hermetic-transport-key", + backend: "cloud", + computer: { name: connectionName }, + requestTimeoutMs: 5_000, + }); + result = await client.prompt("Interpret this minimized receipt.", agentId, { + allowedTools: [], + permissionMode: "strict", + skillSources: [], + stateless: true, + tools: [], + toolset: { base: "none" }, + }); + } finally { + for (const socket of sockets.clients) socket.terminate(); + await new Promise((resolve) => sockets.close(resolve)); + await new Promise((resolve) => server.close(resolve)); + } + + assert.equal(result.success, true); + assert.equal(result.result, "Pinned transport diagnosis."); + assert.equal(result.conversationId, conversationId); + assert.equal(result.stopReason, "end_turn"); + assert.deepEqual( + restRequests.map(({ method, url }) => [method, url]), + [ + ["POST", `/v1/conversations/?agent_id=${agentId}`], + ["GET", "/v1/environments?limit=100"], + ], + ); + assert.equal(socketRequests.length, 2); + assert.deepEqual( + socketRequests.map(({ channel }) => channel).sort(), + ["control", "stream"], + ); + for (const request of socketRequests) { + assert.equal(request.authorization, "Bearer hermetic-transport-key"); + assert.equal( + request.url.pathname, + `/v1/environments/${connectionId}/status/ws`, + ); + assert.equal(request.url.searchParams.get("agentId"), agentId); + assert.equal(request.url.searchParams.get("conversationId"), conversationId); + } + + const commands = frames.control.filter(({ type }) => type !== "ack" && type !== "ping"); + assert.deepEqual(commands.map(({ type }) => type), ["runtime_start", "sync", "input"]); + const [runtimeStart, sync, input] = commands; + assert.equal(runtimeStart.agent_id, agentId); + assert.equal(runtimeStart.conversation_id, conversationId); + assert.equal(runtimeStart.mode, "strict"); + assert.equal(runtimeStart.stateless, true); + assert.deepEqual(runtimeStart.skill_sources, []); + assert.equal(Object.hasOwn(runtimeStart, "external_tools"), false); + assert.deepEqual(sync.runtime, runtime); + assert.equal(sync.recover_approvals, true); + assert.equal(sync.force_device_status, true); + assert.deepEqual(input.runtime, runtime); + assert.equal(input.payload.kind, "create_message"); + assert.equal(input.payload.messages.length, 1); + assert.equal(input.payload.messages[0].role, "user"); + assert.equal(input.payload.messages[0].content, "Interpret this minimized receipt."); + assert.deepEqual(input.payload.client_tool_allowlist, []); + assert.deepEqual(input.payload.client_toolset, { base: "none" }); + assert.equal(input.payload.exclude_interactive_tools, true); }); test("duplicate exact-name specialists fail closed", async () => { @@ -258,6 +526,15 @@ test("clean and benign receipts avoid SDK loading", async () => { { code: "public_edge_converging", summary: "The public edge is converging." }, ], }), + receipt({ + state: "landed", + anomalies: [ + { + code: "current_main_unpublished", + summary: "A newer main revision is not yet published by current pages.", + }, + ], + }), ]) { const result = await interpret(value, {}, async () => { loaded = true; @@ -275,12 +552,6 @@ test("anomaly interpretation uses retained specialist and a fresh zero-authority class FakeClient { constructor(options) { calls.push(["construct", options]); - this.agents = { - list: async (query) => { - calls.push(["list", query]); - return [{ id: "agent-observer", name: SPECIALIST_NAME }]; - }, - }; } async prompt(prompt, agentId, options) { @@ -293,16 +564,43 @@ test("anomaly interpretation uses retained specialist and a fresh zero-authority }; } } + class FakeLetta { + constructor(options) { + calls.push(["lookup-construct", options]); + this.agents = { + list: async (query) => { + calls.push(["list", query]); + return { items: [{ id: "agent-observer", name: SPECIALIST_NAME }] }; + }, + }; + } + } const result = await interpret( receipt(), { LETTA_API_KEY: "key", LETTA_BASE_URL: "https://letta.example" }, - async () => ({ LettaAgentClient: FakeClient }), + async () => ({ Letta: FakeLetta, LettaAgentClient: FakeClient }), ); assert.equal(result.status, "ok"); assert.equal(result.agent_id, "agent-observer"); assert.equal(result.agent_name, SPECIALIST_NAME); assert.equal(result.agent_created, false); assert.equal(result.conversation_id, "conv-one-shot"); + assert.deepEqual(calls.find(([kind]) => kind === "construct"), [ + "construct", + { + apiBaseUrl: "https://letta.example", + apiKey: "key", + backend: "cloud", + }, + ]); + assert.deepEqual(calls.find(([kind]) => kind === "lookup-construct"), [ + "lookup-construct", + { apiKey: "key", baseURL: "https://letta.example" }, + ]); + assert.deepEqual(calls.find(([kind]) => kind === "list"), [ + "list", + { limit: 100, name: SPECIALIST_NAME, show_hidden_agents: true }, + ]); const promptCall = calls.find(([kind]) => kind === "prompt"); assert.ok(promptCall); assert.equal(promptCall[2], "agent-observer"); diff --git a/tools/check.sh b/tools/check.sh index 9fca29f0..6639bc3c 100755 --- a/tools/check.sh +++ b/tools/check.sh @@ -381,12 +381,17 @@ step "landing observer fixtures" if ! python3 tools/test_landing_observer.py; then fail=1 fi -if command -v node >/dev/null 2>&1; then - if ! node --test .agents/skills/observing-misaligned-landings/tests/interpret-anomalies.test.mjs; then - fail=1 - fi +observer_skill=.agents/skills/observing-misaligned-landings +if ! command -v corepack >/dev/null 2>&1; then + echo "FAIL: corepack is required for pinned landing observer fixtures." + fail=1 +elif ! corepack pnpm@10.20.0 --dir "$observer_skill" install --frozen-lockfile; then + echo "FAIL: install pinned landing observer dependencies" + fail=1 +elif ! node --test "$observer_skill"/tests/interpret-anomalies.test.mjs; then + fail=1 else - echo "SKIP: node unavailable; optional interpreter fixtures not run." + echo "PASS: pinned landing observer SDK fixtures" fi if [ "$rust_gate" -eq 1 ]; then diff --git a/tools/test_landing_observer.py b/tools/test_landing_observer.py index 055f942e..038a858b 100644 --- a/tools/test_landing_observer.py +++ b/tools/test_landing_observer.py @@ -104,7 +104,7 @@ def deployment_payload( def channel_payload( revision: str, status: str = "delivered", - message_id: str | None = "491", + message_id: object | None = "491", ) -> dict[str, object]: payload: dict[str, object] = { "channel": "telegram", @@ -132,6 +132,7 @@ def base_evidence(revision: str = "1" * 40, main: str = "2" * 40) -> dict[str, o "primary_worktree": { "dirty": False, "present": True, + "touching": [], }, "project_status": { "consistent": True, @@ -175,6 +176,8 @@ def base_evidence(revision: str = "1" * 40, main: str = "2" * 40) -> dict[str, o }, "repository": { "candidate_head_stable": True, + "changed_paths": ["candidate.txt"], + "changed_paths_scope": landing_observer.CHANGED_PATHS_SCOPE, "hinted_root_matches": True, "local_main_revision": main, "readable": True, @@ -250,6 +253,139 @@ class ReceiptParsingTests(unittest.TestCase): self.assertEqual(receipt["valid"], False) self.assertIn("invalid fields", receipt["read_error"]) + def test_channel_status_controls_whether_a_message_id_may_exist(self) -> None: + revision = "1" * 40 + cases = [ + ("delivered", "491", True), + ("delivered", 491, True), + ("delivered", None, False), + ("delivered", "", False), + ("delivered", "0", False), + ("delivered", "-491", False), + ("delivered", "+491", False), + ("delivered", "0491", False), + ("delivered", "491.0", False), + ("delivered", "٤٩١", False), + ("delivered", -491, False), + ("delivered", 491.0, False), + ("delivered", True, False), + ("unknown", None, True), + ("failed", None, True), + ("unavailable", None, True), + ("unknown", "491", False), + ("failed", "491", False), + ("unavailable", "491", False), + ] + with tempfile.TemporaryDirectory() as raw_tmp: + path = Path(raw_tmp) / "channel.json" + for status, message_id, expected_valid in cases: + with self.subTest(status=status, message_id=message_id): + path.write_text( + json.dumps(channel_payload(revision, status, message_id)), + encoding="utf-8", + ) + receipt = landing_observer.read_receipt(path, "channel") + self.assertEqual(receipt["valid"], expected_valid) + + def test_channel_receipt_rejects_every_unsupported_identity_field(self) -> None: + revision = "1" * 40 + cases = [ + channel_payload(revision, "pending", None), + {**channel_payload(revision), "channel": "slack"}, + channel_payload("1" * 39), + channel_payload("G" * 40), + ] + with tempfile.TemporaryDirectory() as raw_tmp: + path = Path(raw_tmp) / "channel.json" + for payload in cases: + with self.subTest(payload=payload): + path.write_text(json.dumps(payload), encoding="utf-8") + receipt = landing_observer.read_receipt(path, "channel") + self.assertEqual(receipt["valid"], False) + self.assertIn("invalid fields", receipt["read_error"]) + + +class PorcelainPathParsingTests(unittest.TestCase): + def test_literal_paths_include_both_sides_of_renames_and_copies(self) -> None: + output = ( + " M plain.txt\0" + "?? path with spaces.txt\0" + "R renamed destination.txt\0renamed source.txt\0" + " C copied\nline.txt\0copy source.txt\0" + ) + self.assertEqual( + landing_observer.porcelain_v1_z_paths(output), + sorted( + [ + "plain.txt", + "path with spaces.txt", + "renamed destination.txt", + "renamed source.txt", + "copied\nline.txt", + "copy source.txt", + ] + ), + ) + + def test_malformed_or_non_nul_status_fails_closed(self) -> None: + for output in (" M plain.txt", "R destination.txt\0", "bad\0"): + with self.subTest(output=output): + with self.assertRaises(ValueError): + landing_observer.porcelain_v1_z_paths(output) + + def test_diff_paths_are_literal_and_include_both_rename_sides(self) -> None: + output = ( + "M\0plain.txt\0" + "R100\0renamed source.txt\0renamed\ndestination.txt\0" + "C75\0copy source.txt\0copied destination.txt\0" + ) + self.assertEqual( + landing_observer.diff_name_status_z_paths(output), + sorted( + [ + "plain.txt", + "renamed source.txt", + "renamed\ndestination.txt", + "copy source.txt", + "copied destination.txt", + ] + ), + ) + + def test_path_component_overlap_covers_both_ancestor_directions(self) -> None: + self.assertEqual( + landing_observer.overlapping_candidate_paths( + ["untracked tree/", "candidate-parent/dirty.txt", "unrelated.txt"], + [ + "untracked tree/nested/file.txt", + "candidate-parent", + "untracked treehouse/file.txt", + "candidate.txt", + ], + ), + ["candidate-parent", "untracked tree/nested/file.txt"], + ) + + def test_command_runner_preserves_non_utf8_git_path_bytes(self) -> None: + with tempfile.TemporaryDirectory() as raw_tmp: + root = Path(raw_tmp) + git(root, "init", "-q") + raw_name = b"non-utf8-\xff.txt" + descriptor = os.open(os.fsencode(root) + b"/" + raw_name, os.O_WRONLY | os.O_CREAT, 0o600) + os.close(descriptor) + evidence = landing_observer.live_worktree( + landing_observer.command_runner, + root, + ) + self.assertEqual(evidence["dirty"], True) + self.assertEqual(evidence["touching"], [os.fsdecode(raw_name)]) + + def test_malformed_or_non_nul_diff_fails_closed(self) -> None: + for output in ("M\0plain.txt", "R100\0source.txt\0", "?\0path.txt\0"): + with self.subTest(output=output): + with self.assertRaises(ValueError): + landing_observer.diff_name_status_z_paths(output) + class ClassificationMatrixTests(unittest.TestCase): def classify(self, evidence: dict[str, object], revision: str = "1" * 40): @@ -325,15 +461,68 @@ class ClassificationMatrixTests(unittest.TestCase): ) self.assertEqual(anomalies, []) - def test_dirty_primary_is_ownership_warning_while_ready(self) -> None: + def test_unrelated_dirty_primary_blocks_existing_landing_path(self) -> None: evidence = base_evidence() evidence["primary_worktree"]["dirty"] = True + evidence["primary_worktree"]["touching"] = ["someone-elses-file.txt"] state, missing, ownership, anomalies, _ = self.classify(evidence) - self.assertEqual(state, "ready") - self.assertEqual(codes(missing), {"land_candidate"}) + self.assertEqual(state, "blocked") + self.assertEqual(codes(missing), {"primary_checkout_dirty"}) self.assertEqual(codes(ownership), {"unrelated_primary_checkout_dirty"}) self.assertEqual(anomalies, []) + def test_candidate_path_overlap_is_not_called_unrelated(self) -> None: + evidence = base_evidence() + evidence["primary_worktree"].update( + dirty=True, + touching=["candidate.txt", "someone-elses-file.txt"], + ) + state, missing, ownership, anomalies, _ = self.classify(evidence) + self.assertEqual(state, "blocked") + self.assertEqual(codes(missing), {"primary_checkout_candidate_path_overlap"}) + self.assertEqual(codes(ownership), {"primary_checkout_candidate_path_overlap"}) + self.assertEqual(anomalies, []) + + def test_primary_path_ancestor_in_either_direction_is_candidate_overlap(self) -> None: + cases = [ + (["candidate-dir/"], ["candidate-dir/nested.txt"]), + (["candidate-dir/nested.txt"], ["candidate-dir"]), + ] + for touching, changed_paths in cases: + with self.subTest(touching=touching, changed_paths=changed_paths): + evidence = base_evidence() + evidence["primary_worktree"].update(dirty=True, touching=touching) + evidence["repository"]["changed_paths"] = changed_paths + state, missing, ownership, anomalies, _ = self.classify(evidence) + self.assertEqual(state, "blocked") + self.assertEqual(codes(missing), {"primary_checkout_candidate_path_overlap"}) + self.assertEqual(codes(ownership), {"primary_checkout_candidate_path_overlap"}) + self.assertEqual(anomalies, []) + + def test_dirty_primary_rejects_candidate_paths_without_exact_scope(self) -> None: + for scope in (None, "candidate-to-main", ""): + with self.subTest(scope=scope): + evidence = base_evidence() + evidence["primary_worktree"].update( + dirty=True, + touching=["someone-elses-file.txt"], + ) + evidence["repository"]["changed_paths_scope"] = scope + state, missing, ownership, anomalies, _ = self.classify(evidence) + self.assertEqual(state, "blocked") + self.assertEqual(codes(missing), {"primary_checkout_dirty_paths_unproven"}) + self.assertEqual(codes(ownership), {"primary_checkout_dirty_paths_unproven"}) + self.assertEqual(anomalies, []) + + def test_dirty_primary_with_unknown_paths_fails_closed(self) -> None: + evidence = base_evidence() + evidence["primary_worktree"].update(dirty=True, touching=None) + state, missing, ownership, anomalies, _ = self.classify(evidence) + self.assertEqual(state, "blocked") + self.assertEqual(codes(missing), {"primary_checkout_dirty_paths_unproven"}) + self.assertEqual(codes(ownership), {"primary_checkout_dirty_paths_unproven"}) + self.assertEqual(anomalies, []) + def test_landed_requires_remote_pages_gate_terminal_run_and_cleanup(self) -> None: state, missing, ownership, anomalies, summary = self.classify(landed_evidence()) self.assertEqual(state, "landed") @@ -398,15 +587,49 @@ class ClassificationMatrixTests(unittest.TestCase): self.assertEqual(codes(missing), {"telegram_delivery_pending"}) self.assertEqual(codes(anomalies), {"telegram_runtime_unavailable"}) + def test_failed_telegram_does_not_claim_runtime_unavailability(self) -> None: + evidence = landed_evidence() + evidence["receipts"]["channel"].update( + message_id=None, + status="failed", + ) + state, missing, _, anomalies, _ = self.classify(evidence) + self.assertEqual(state, "landed") + self.assertEqual(codes(missing), {"telegram_delivery_pending"}) + self.assertEqual(codes(anomalies), {"telegram_delivery_failed"}) + + def test_unknown_telegram_remains_uncertain_not_failed_or_unavailable(self) -> None: + evidence = landed_evidence() + evidence["receipts"]["channel"].update( + message_id=None, + status="unknown", + ) + state, missing, _, anomalies, _ = self.classify(evidence) + self.assertEqual(state, "landed") + self.assertEqual(codes(missing), {"telegram_delivery_status_unknown"}) + self.assertEqual(anomalies, []) + def test_unrelated_dirty_primary_does_not_undo_completed_owned_cleanup(self) -> None: evidence = landed_evidence() - evidence["primary_worktree"]["dirty"] = True + evidence["primary_worktree"].update( + dirty=True, + touching=["someone-elses-file.txt"], + ) state, missing, ownership, anomalies, _ = self.classify(evidence) self.assertEqual(state, "landed") self.assertEqual(missing, []) self.assertEqual(codes(ownership), {"unrelated_primary_checkout_dirty"}) self.assertEqual(anomalies, []) + def test_landed_overlap_is_labeled_without_undoing_remote_cleanup_proof(self) -> None: + evidence = landed_evidence() + evidence["primary_worktree"].update(dirty=True, touching=["candidate.txt"]) + state, missing, ownership, anomalies, _ = self.classify(evidence) + self.assertEqual(state, "landed") + self.assertEqual(missing, []) + self.assertEqual(codes(ownership), {"primary_checkout_candidate_path_overlap"}) + self.assertEqual(anomalies, []) + def test_superseded_historical_landing_remains_landed(self) -> None: evidence = landed_evidence() newer = "9" * 40 @@ -714,6 +937,14 @@ class IntegrationTests(unittest.TestCase): self.assertEqual(codes(receipt["missing_steps"]), {"land_candidate"}) self.assertEqual(receipt["anomalies"], []) self.assertEqual(receipt["evidence"]["target_worktree"]["present"], True) + self.assertEqual( + receipt["evidence"]["repository"]["changed_paths"], + ["README.md"], + ) + self.assertEqual( + receipt["evidence"]["repository"]["changed_paths_scope"], + landing_observer.CHANGED_PATHS_SCOPE, + ) def test_completed_landing_reads_cleaned_path_and_exact_receipts(self) -> None: with tempfile.TemporaryDirectory() as raw_tmp: @@ -740,6 +971,57 @@ class IntegrationTests(unittest.TestCase): fixture.revision, ) + def test_clean_descendant_after_cleanup_preserves_historical_landing(self) -> None: + with tempfile.TemporaryDirectory() as raw_tmp: + fixture = LandingFixture(Path(raw_tmp)) + fixture.land() + (fixture.root / "README.md").write_text( + "fixture descendant after candidate cleanup\n", + encoding="utf-8", + ) + git(fixture.root, "add", "README.md") + git(fixture.root, "commit", "-q", "-m", "fixture descendant") + descendant = git(fixture.root, "rev-parse", "HEAD") + git(fixture.root, "push", "-q", "origin", "main") + git(fixture.root, "update-ref", "refs/remotes/origin/main", descendant) + fixture.write_page(descendant, "fixture descendant pages") + fixture.write_receipt( + fixture.channel_receipt, + channel_payload(fixture.revision), + ) + + before = filesystem_fingerprint(fixture.base) + observed = fixture.observe( + channel=fixture.channel_receipt, + extra_env={"OBSERVER_PUBLIC_SOURCE": descendant}, + ) + after = filesystem_fingerprint(fixture.base) + + self.assertEqual(observed.returncode, 0, observed.stderr) + self.assertEqual(before, after) + receipt = json.loads(observed.stdout) + self.assertEqual(receipt["state"], "landed") + self.assertEqual(receipt["missing_steps"], []) + self.assertEqual(receipt["anomalies"], []) + self.assertEqual( + receipt["evidence"]["repository"]["remote_main_revision"], + descendant, + ) + self.assertEqual( + receipt["evidence"]["repository"]["remote_main_contains_expected"], + True, + ) + self.assertEqual( + receipt["evidence"]["pages"]["source_revision"], + descendant, + ) + self.assertIsNone( + receipt["evidence"]["repository"]["changed_paths"], + ) + self.assertIsNone( + receipt["evidence"]["repository"]["changed_paths_scope"], + ) + def test_remote_pages_can_prove_landing_while_public_edge_lags(self) -> None: with tempfile.TemporaryDirectory() as raw_tmp: fixture = LandingFixture(Path(raw_tmp)) @@ -817,7 +1099,7 @@ class IntegrationTests(unittest.TestCase): {"gate_revision_mismatch", "project_status_inconsistent"}, ) - def test_dirty_primary_is_only_an_ownership_warning_after_cleanup(self) -> None: + def test_dirty_primary_scope_is_unproven_but_does_not_erase_cleanup(self) -> None: with tempfile.TemporaryDirectory() as raw_tmp: fixture = LandingFixture(Path(raw_tmp)) fixture.land() @@ -833,7 +1115,7 @@ class IntegrationTests(unittest.TestCase): self.assertEqual(receipt["state"], "landed") self.assertEqual( codes(receipt["ownership_warnings"]), - {"unrelated_primary_checkout_dirty"}, + {"primary_checkout_dirty_paths_unproven"}, ) diff --git a/wiki/log/2026-08-12-landing-observer-trust-boundary.md b/wiki/log/2026-08-12-landing-observer-trust-boundary.md new file mode 100644 index 00000000..105cdb52 --- /dev/null +++ b/wiki/log/2026-08-12-landing-observer-trust-boundary.md @@ -0,0 +1,72 @@ +# Landing observation earns its certainty at the transport boundary + +``` +Type: log +``` + +## Why + +The first one-shot observer had the right authority boundary but still trusted four +soft edges. It called every dirty primary checkout unrelated without comparing exact +paths, collapsed explicit Telegram failure into channel-runtime unavailability, +provisioned its retained specialist on ordinary agent surfaces, and tested the +interpreter mostly through a hand-shaped SDK facade. The Node suite could also skip +those SDK checks when dependencies had not already been installed. + +Those gaps did not give the observer a mutation channel, but they let it report more +certainty than its instruments had earned. This pass hardens the existing observer; +it does not add another landing step or widen interpretation authority. + +## Implemented + +- Primary and candidate path evidence now comes from NUL-delimited Git status and + name-status output. Literal spaces and newlines survive, and both source and + destination sides of renames and copies participate in overlap. Malformed output + becomes unknown instead of a guessed path set. Surrogate escape preserves non-UTF-8 + Git path bytes. The candidate delta is explicitly scoped from merge base to candidate + while it remains provable; once current main contains the candidate, that original + delta is unknown rather than mislabeled empty. Every untracked descendant is enumerated, and path-component ancestors overlap in + either direction without prefix-colliding siblings. +- Before landing, any unresolved primary-checkout dirt blocks the existing landing + path. Candidate overlap gets its own exact ownership code; unrelated dirt is named + only after a real disjoint comparison. Exact remote, gate, and cleanup proof still + preserves a completed landing if the primary is dirtied later. +- Channel receipts now distinguish `unknown`, explicit delivery `failed`, and runtime + `unavailable`. Only the last state claims that Telegram transport could not be + reached. `delivered` requires one positive decimal message id; all non-delivered + states reject one rather than carrying contradictory proof. +- Newly provisioned `Misaligned Landing Observer` specialists are hidden while exact + API name lookup still reuses the retained identity. Empty memory, disabled MemFS, + no server or client tools, no skills, strict permission mode, stateless fresh + conversations, and minimized evidence remain binding. +- A hermetic local HTTP/WebSocket fixture now drives the real pinned + `@letta-ai/letta-agent-sdk`. It observes conversation creation, explicit computer + resolution, separate control and stream sockets, strict stateless runtime startup, + empty skill sources, post-start synchronization, the exact zero-authority input + payload, and a terminal SDK result. +- `ws` is pinned as a development dependency. The project gate now runs a frozen + install through exact pnpm `10.20.0` before the Node suite, so absent packages can + no longer turn transport coverage into a skip. +- A clean descendant after candidate cleanup is exercised end to end, while + `current_main_unpublished` remains a benign deterministic historical-publication + fact that does not spend an Agent SDK turn by itself. + +## Verification + +Focused Python fixtures cover literal and malformed status/diff paths, rename/copy +ownership, component-aware ancestor overlap, non-UTF-8 path bytes, pre-landing primary +dirt, completed-landing descendants, and complete Telegram receipt states/message ids. Node fixtures cover +hidden provisioning plus the installed SDK's actual REST/WebSocket boundary and +terminal result. The generated corpus indexes and the exact project landing gate +verify the reconciled candidate. + +## Not done + +This does not create delivery receipts, retry Telegram, make agent interpretation +mandatory, modify `tools/task.sh`, or let the observer repair, land, deploy, notify, +schedule, or clean anything it reads. + +**Defense:** unknown is not absent, and a label is not a receipt. Exact path bytes, +distinct channel states, hidden zero-authority provisioning, and the real pinned +transport each prove only what they can actually reach. Deterministic landing law +remains outside the specialist and outside this observer. diff --git a/wiki/log/DEVLOG.md b/wiki/log/DEVLOG.md index 36132453..1abd2ede 100644 --- a/wiki/log/DEVLOG.md +++ b/wiki/log/DEVLOG.md @@ -16,6 +16,11 @@ add or amend a session log, then re-run the generator. - Intent: (see session log) - Log: [wiki/log/2026-08-12-one-shot-landing-observer.md](2026-08-12-one-shot-landing-observer.md) +## 2026-08-12 - Landing observation earns its certainty at the transport boundary + +- Intent: (see session log) +- Log: [wiki/log/2026-08-12-landing-observer-trust-boundary.md](2026-08-12-landing-observer-trust-boundary.md) + ## 2026-08-11 - Territorial Personas core - Intent: (see session log) diff --git a/wiki/log/decisions/2026-08-12.md b/wiki/log/decisions/2026-08-12.md index da8dd374..ebee0641 100644 --- a/wiki/log/decisions/2026-08-12.md +++ b/wiki/log/decisions/2026-08-12.md @@ -27,6 +27,25 @@ Type: log conversation over minimized evidence, and its answer cannot change the receipt. - Receipt production and `tools/task.sh` integration remain outside this slice. The observer accepts optional exact receipts without inventing or backfilling them. +- Dirty-path ownership is literal evidence, not a clean/dirty guess. The observer + parses NUL-delimited status and candidate diffs, includes both rename/copy sides, + binds candidate paths to the exact merge-base-to-candidate scope only while current + main does not contain the candidate, preserving the original delta as unknown after + landing rather than mislabeled empty. It enumerates all untracked descendants, and compares path-component ancestors in both directions. It + blocks a pre-landing candidate on any unresolved primary-checkout dirt while naming + exact overlap separately. Once remote, gate, and cleanup receipts prove a + landing, later primary dirt can warn about ownership but cannot erase that result. +- Telegram receipt states preserve their actual meaning: `unknown` stays uncertain, + `failed` means an attempted delivery failed, and `unavailable` means the outbound + runtime was unavailable. None may borrow another state's explanation or carry a + message id; only `delivered` carries one positive decimal message id. +- The retained specialist is hidden from ordinary agent surfaces. The project gate + performs a frozen install with pnpm `10.20.0`, then drives the real pinned SDK + through a hermetic HTTP/WebSocket service to prove conversation creation, explicit + computer resolution, dual sockets, strict stateless no-skill startup, synchronization, + empty client-tool authority, interactive-tool exclusion, and terminal completion. + Clean descendants after owned cleanup preserve historical landing proof, while their + benign current-publication lag does not invoke the specialist by itself. ### REJECTED @@ -42,6 +61,13 @@ Type: log - **Poll until the public edge agrees.** Remote publication and edge convergence are separate facts; one read is enough unless a real deployment failure is being diagnosed. +- **Call all primary-checkout dirt unrelated, or infer path ownership from quoted + line-oriented Git output.** Both hide rename/copy and unusual-filename overlap. +- **Collapse unknown, failed, and unavailable Telegram delivery into one runtime + anomaly.** Those are three different claims and only one names runtime reach. +- **Mock only the interpreter's SDK facade or skip the Node suite when dependencies + are absent.** Facade tests cannot prove what the pinned SDK puts on its actual REST + and WebSocket boundaries. Owners: [repository skills](../../process/repository-skills.md) and [landing workflow](../../process/workflows.md#git-conventions). diff --git a/wiki/process/repository-skills.md b/wiki/process/repository-skills.md index ebf4c881..e01efaca 100644 --- a/wiki/process/repository-skills.md +++ b/wiki/process/repository-skills.md @@ -40,7 +40,7 @@ Claude mirror is needed. | `design-companion` | `.agents/skills/design-companion/SKILL.md` | Explore and explain unsettled game-design choices in conversation without mutating the repository; hand adopted choices to `design-session`. | | `design-session` | `.agents/skills/design-session/SKILL.md`
`.claude/skills/design-session/SKILL.md` | Capture affirmed design decisions into their owning law/spec pages, decision history, and session trace. | | `playtesting-misaligned` | `.agents/skills/playtesting-misaligned/SKILL.md` | Run evidence-bearing naive and informed playtests against the current player surface and corpus. | -| `observing-misaligned-landings` | `.agents/skills/observing-misaligned-landings/SKILL.md` | Read one exact candidate across isolated remote snapshots, public/site and project-operation evidence, classify it without mutation, and optionally ask one retained exact-name, zero-tool Letta specialist in a fresh conversation to interpret only non-benign anomalies. | +| `observing-misaligned-landings` | `.agents/skills/observing-misaligned-landings/SKILL.md` | Read one exact candidate across isolated remote snapshots, exact dirty-path ownership, public/site and project-operation evidence, classify it without mutation, and optionally ask one retained hidden exact-name, zero-tool Letta specialist in a fresh conversation to interpret only non-benign anomalies. | | `session-wrap` | `.agents/skills/session-wrap/SKILL.md`
`.claude/skills/session-wrap/SKILL.md` | Finish or preserve owned work, read exact current project state from live authorities, and report one bounded next step without maintaining a parallel checked-in handoff ledger. | | `tick` | `.agents/skills/tick/SKILL.md`
`.claude/skills/tick/SKILL.md` | Invoke the project's bounded stewardship heartbeat: audit one corpus slice, act on one finding, and leave a trace. | @@ -67,12 +67,21 @@ Claude mirror is needed. project-status, and issue authorities; it does not maintain another tracked current-state file or hard-code one agent harness's memory tool. 6. A read-only observer fingerprints the target repository before and after the - observation, isolates remote fetches outside the repository, preserves unknown - evidence, and cannot mutate or repair the state it classifies. + observation, isolates remote fetches outside the repository, derives literal + dirty and candidate paths from NUL-delimited Git output including both rename/copy + sides, records the exact merge-base-to-candidate delta scope only while current main + does not contain the candidate (otherwise preserving it as unknown), enumerates + untracked descendants, compares path-component ancestors in both directions, preserves + unknown evidence, and cannot mutate or repair the state it classifies. Before + landing, unresolved primary-checkout dirt blocks the existing path; after exact + remote and cleanup proof, later dirt or a clean descendant cannot erase history. 7. Optional agent interpretation cannot override deterministic state. It reuses exactly one retained `Misaligned Landing Observer` with empty memory, MemFS and - tools disabled, opens a fresh conversation over minimized evidence for each - observation, provisions only when absent, and fails closed on duplicate exact names. + tools disabled, hidden from ordinary agent surfaces, opens a fresh conversation + over minimized evidence for each observation, provisions only when absent, and + fails closed on duplicate exact names. The pinned SDK transport is exercised + end-to-end against a hermetic fake HTTP/WebSocket service after a frozen install; + missing dependencies cannot skip that gate. ## Defense @@ -80,4 +89,8 @@ The landing observer is registered here because it is a checked-in operational entry point, while the landing workflow remains owned by [workflows.md](workflows.md#git-conventions). Keeping deterministic evidence and optional interpretation subordinate to that law prevents a convenient observer from -becoming a second landing authority or a mutation doorway. +becoming a second landing authority or a mutation doorway. Exact scoped path +comparison, complete channel status/message-id semantics, benign historical-publication +handling, hidden provisioning, and real pinned-transport coverage close the remaining +places where an observer could claim more certainty or authority than its instruments +earned. diff --git a/wiki/process/workflows.md b/wiki/process/workflows.md index a1eec3fb..bfab7e3b 100644 --- a/wiki/process/workflows.md +++ b/wiki/process/workflows.md @@ -291,8 +291,26 @@ into the project, alter refs, push, deploy, notify, schedule, or clean worktrees Unknown evidence fails closed. A later coherent landing supersedes an older candidate rather than retroactively making that older landing inconsistent. Optional Letta interpretation is subordinate and cannot change the state. It reuses one retained -exact-name, empty-memory, zero-tool specialist but opens a fresh conversation for each -observation; duplicate exact names fail closed. +exact-name, hidden, empty-memory, zero-tool specialist but opens a fresh conversation +for each observation; duplicate exact names fail closed. Before a landing, NUL-safe +status and diff evidence compares both literal sides of renames and copies: any dirty +primary checkout blocks the existing path until its owner resolves it, and overlap +with the candidate is named exactly rather than guessed unrelated. The candidate delta +is explicitly scoped from merge base to candidate while the candidate remains +unlanded; once current main contains it, the original delta is unknown rather than +mislabeled empty. Missing or different scope cannot prove disjoint ownership. Non-UTF-8 path bytes survive through surrogate escape, every +untracked descendant is enumerated, and a dirty path-component ancestor in either +direction overlaps. Immutable remote, gate, and cleanup proof still preserves a +completed landing if the primary becomes dirty later or a clean descendant supersedes +it. Telegram receipts keep `unknown`, explicit delivery failure, and an unavailable +channel runtime distinct: only `delivered` carries a positive message id, while all +three non-delivered states carry none. +The project gate installs the skill lockfile with its pinned pnpm `10.20.0` before +running Node fixtures. Those fixtures drive the real pinned Agent SDK through a +hermetic fake HTTP/WebSocket transport and prove conversation creation, explicit +computer resolution, dual sockets, strict stateless startup without skills, +post-start synchronization, empty client-tool authority, interactive-tool exclusion, +and a terminal result. Do not poll the edge merely because its one public read still serves an older source. Starlight is the documentation renderer. Its sync copies root `DESIGN.md` to @@ -412,7 +430,11 @@ sleep 8 && kill %1; grep -iE "panic|ERROR" /tmp/bevy.log The observer reads the same exact revision and authorities already required by this workflow, but cannot produce any landing effect. Isolated remote snapshots prevent its proof from depending on stale local tracking refs; pre/post repository -fingerprints make the read-only boundary testable; explicit unknowns prevent missing -transport, project-status, receipt, or public evidence from becoming a false green. -Keeping agent interpretation optional and unable to override the four-state classifier -preserves the executable workflow as the only landing authority. +fingerprints make the read-only boundary testable; exact NUL-delimited path ownership +prevents primary-checkout dirt from being waved through as unrelated; explicit +unknowns prevent missing transport, channel, project-status, receipt, candidate-delta +scope, or public evidence from becoming a false green. Completed historical landings +and their benign current-publication lag remain local deterministic facts rather than +unnecessary agent prompts. Keeping agent interpretation hidden, +zero-authority, transport-tested, optional, and unable to override the four-state +classifier preserves the executable workflow as the only landing authority.