From bb5a35e2642c200d4008d67ddd68877720e8cd88 Mon Sep 17 00:00:00 2001 From: Patrick Singletary Date: Mon, 7 Sep 2026 11:40:38 -0400 Subject: [PATCH] agent-harness: instruction-flow & model-change risk analysis (2026-09-07) --- .../INSTRUCTION-FLOW-ANALYSIS.md | 272 ++++++++++++++++++ 1 file changed, 272 insertions(+) create mode 100644 docs/agent-harness/INSTRUCTION-FLOW-ANALYSIS.md diff --git a/docs/agent-harness/INSTRUCTION-FLOW-ANALYSIS.md b/docs/agent-harness/INSTRUCTION-FLOW-ANALYSIS.md new file mode 100644 index 0000000..1ce8e07 --- /dev/null +++ b/docs/agent-harness/INSTRUCTION-FLOW-ANALYSIS.md @@ -0,0 +1,272 @@ +# Instruction Flow & Model-Change Risk Analysis — Hermes Harness + +**Date:** 2026-09-07 +**Scope:** Why repo-level `AGENTS.md` instructions stop being honored after model +changes, plus a risk register for anything else that shifts when the model changes. +**Grounding:** installed Hermes source (v0.21.0, local 57f05e21, 2026.8.31), +`~/.hermes/state.db`, `~/.hermes/config.yaml`, `SOUL.md`, hooks, memory store, +per-repo context files, Zed settings. Every claim below traces to one of those. + +--- + +## 1. Headline findings (TL;DR) + +1. **The harness honors `.hermes.md` over `AGENTS.md`, and only one of them at a + time.** In timer, slicer, sifter, and acct-tools — repos that carry BOTH files — + Hermes loads only the ~900-byte `.hermes.md` boilerplate at session start. The + substantive `AGENTS.md` (deploy flows, guardrails, source-of-truth pointers) + never enters the system prompt. Your own convention ("AGENTS.md primary, + .hermes.md secondary") is the exact inverse of Hermes semantics, and it is + written into `_shared/AGENTS.md`, memory, and `AGENTS.common.md`. + +2. **Context files load from the session's working directory only.** Every CLI + session in state.db runs with `cwd = /Users/patricksingletary` (home) — not a + repo — so at startup *no* project context file loads at all in CLI sessions. + `_shared/AGENTS.md` content reached past sessions only as mid-session + "subdirectory hints" appended to tool results, which are advisory context, not + startup rules. + +3. **A model change is not what broke loading — it is what broke adherence.** + State.db shows every session since 2026-08-13/14 on `deepseek/deepseek-v4-flash` + / `-0731` (before that: v4-pro, kimi-k3 on 08-08, poolside free models in July). + Rules were *visible* to earlier, stronger models; the flash class follows + tail-end instructions in a very large system prompt more loosely, and the + same prompt is assembled the same way for every model. Symptom looks like + "instructions not honored," mechanism is (1) never loaded or (2) weakly followed. + +4. **Two loading paths disagree.** Startup loading (`agent/prompt_builder.py`) + prefers `.hermes.md` and silently drops `AGENTS.md` when both exist. Mid-session + hint loading (`agent/subdirectory_hints.py`) prefers `AGENTS.override.md` / + `AGENTS.md` and never loads `.hermes.md`. The same repo therefore yields + *different rule files* depending on whether instructions arrive at startup + (system prompt) or as a tool-result hint. + +5. **Documentation about this is stale in two places that matter:** + - The installed `hermes-agent` skill's `references/project-context-files.md` + says "AGENTS.md — cwd only, first match wins." The installed build actually + merges an `AGENTS.md` chain from git root down to cwd, supports + `AGENTS.override.md`, and adds per-touch subdirectory hints. Reasoning from + the skill misdiagnoses this setup. + - `_shared/docs/MODEL_STRATEGY_GUIDE.md` describes a tier system + (smart_chat.sh/budget_monitor.sh, laguna 8K context, openrouter models) that + no longer matches `config.yaml` (nous provider, deepseek v4 flash family). + +--- + +## 2. What was inspected + +| Path | Role | +|---|---| +| `~/.hermes/config.yaml` (8.4 KB) | Harness settings: model default, reasoning, compression, hooks, toolsets | +| `~/.hermes/SOUL.md` (667 B) | Identity slot — always loaded, no project rules | +| `~/.hermes/agent-hooks/plan-inbox-hook.sh` | pre_llm_call hook, first-turn plan menu | +| `~/.hermes/agent-hooks/zedra-agent-hooks.sh` | Zedra passthrough on 6 lifecycle events | +| `~/.hermes/hermes-agent/agent/prompt_builder.py` | Startup context-file loader (source of truth) | +| `~/.hermes/hermes-agent/agent/subdirectory_hints.py` | Mid-session context-hint loader (source of truth) | +| `~/.hermes/hermes-agent/acp_adapter/session.py` | Zed ACP session/cwd handling | +| `~/.hermes/state.db` | Session history: source, model, cwd, titles, message evidence | +| `~/.hermes/memories/MEMORY.md`, `USER.md` | Persistent per-turn injected memory (2,200 / 1,375 char caps) | +| `~/.config/zed/settings.json` | Zed `agent_servers.hermes-agent` = `hermes acp` | +| 13x `~/dev/*/AGENTS.md` + 4x `.hermes.md` | Per-repo instruction files (all < 3 KB — no truncation caps hit) | +| `~/dev/_shared/docs/*` | Shared convention docs incl. MODEL_STRATEGY_GUIDE.md | + +--- + +## 3. How instructions actually reach a session + +### 3.1 Startup — system prompt (prompt_builder.py) + +`build_context_files_prompt()` loads from the session cwd, in this exact +precedence — **only ONE project context source loads, first found wins**: + +1. `.hermes.md` / `HERMES.md` — nearest match walking cwd up to the git root + (no git root → cwd only). +2. `AGENTS.md` chain — per directory, first of `AGENTS.override.md` / + `AGENTS.md` / `agents.md`, from git root down to cwd; merged sections, dedup + of identical content, one budget for the whole chain. +3. `CLAUDE.md` — cwd only. +4. `.cursorrules` + `.cursor/rules/*.mdc` — cwd only. + +`SOUL.md` from the Hermes home is independent and ALWAYS appended. All project +context renders under a `# Project Context` header stating the files "should be +followed." Caps: 20,000 chars static floor; dynamic cap = 0.06 × context length +× 4 chars/token (≈314 KB for the 1.31M-token v4-flash-0731), ceiling 500 KB. +None of the files here approach a cap. + +### 3.2 Mid-session — subdirectory hints (subdirectory_hints.py) + +When a tool call references a path under cwd (read_file `path`, terminal +commands, workdirs), Hermes walks that directory and up to 5 ancestors and, on +first visit, appends that directory's first context file +(`AGENTS.override.md` > `AGENTS.md` > `CLAUDE.md` > `.cursorrules` — note: +**`.hermes.md` is not in this list**) to the *tool result* — head+tail truncated +past 32,000 chars, content-digest deduped. This preserves the prompt-cache +prefix (nothing in the system prompt changes mid-conversation). Observed live in +this session: reading `_shared/AGENTS.md` and `timer/.hermes.md` each pulled the +target repo's `AGENTS.md` in as "Subdirectory context discovered: ...". + +Consequence: a repo's real rules reach the model only if the agent *happens to +touch that directory* during the session, and they arrive wrapped in tool output +without the "should be followed" framing of §3.1. + +### 3.3 Always-on layers + +- Memory + user profile (MEMORY.md / USER.md) injected every turn, char-capped. +- Full skills catalog rendered into the system prompt from a 53 KB snapshot file + (rebuilt when the catalog changes). +- `plan-inbox-hook.sh` — first turn of a new session only, appends the pending + PLAN menu to the user message. +- `zedra-agent-hooks.sh` — forwards lifecycle events to `zedra` when inside a + Zedra terminal; no-op otherwise. + +--- + +## 4. Why AGENTS.md "isn't being honored" — root causes, ranked + +### RC1 — `.hermes.md` shadows `AGENTS.md` in 4 repos (highest impact, fully deterministic) + +timer, slicer, sifter, acct-tools each carry `.hermes.md` AND `AGENTS.md`. The +`.hermes.md` files are ~900 B of boilerplate whose own text says "AGENTS.md has +the portable reference" — i.e., the author's intent was AGENTS.md = content, +.hermes.md = pointer. Hermes startup loads the pointer and **never loads the +content** (one-source rule, §3.1). In timer specifically, `AGENTS.md` carries +binding PLAN-timer guardrails and the correct wisp deploy form; none of that is +in the system prompt of a session rooted at `~/dev/timer`. + +### RC2 — Session cwd is home, so nothing project-level loads in CLI sessions + +state.db (last 40 CLI sessions): `cwd = /Users/patricksingletary` for all. +`_find_git_root(~/...)` = none, and home has no AGENTS.md/.hermes.md → empty +project-context section. The shared rules that are supposed to apply to "all +agents, all tasks" (`_shared/AGENTS.md` security + plan-intake sections) load +only when the session cwd is inside `~/dev/_shared`. That repo's rules reached +other sessions exclusively via subdirectory hints (§3.2) — best-effort, not +guaranteed. + +### RC3 — Model change (Aug 8→14: kimi-k3/v4-pro → v4-flash → v4-flash-0731) lowered tail-instruction adherence + +The prompt is assembled identically regardless of model, but the model classes +differ sharply in how faithfully they follow rules that sit at the end of a very +long system prompt (Hermes core + 100+ skill index + memory), especially across +long tool-heavy turns. Combined with RC1/RC2 (rules often arriving only as +tool-result hints), a flash model is the worst case for "rules that are barely +visible." Nothing in the harness re-orders or strengthens project context for +weaker models. + +### RC4 — Documented convention is inverted vs. the harness + +`_shared/AGENTS.md` line 3 and memory both assert "AGENTS.md primary, .hermes.md +secondary." Hermes semantics: `.hermes.md` wins when both exist at startup; +`AGENTS.md` only loads when no `.hermes.md` is in the walk. Anyone (human or +agent) reasoning from the stated convention predicts the wrong behavior, and +edits made to AGENTS.md in the 4 dual-file repos have no effect on Hermes +sessions. + +### RC5 — Startup and hint loaders disagree (same repo, two rule sets) + +Covered in §3.2/§4-RC1. A session rooted in `~/dev/timer` starts with the +`.hermes.md` pointer, but the first time it touches the tree it also receives +the full `AGENTS.md` as a hint — so mid-session the model suddenly has "new" +rules (or contradictory ones) it did not have at start. + +--- + +## 5. Model-change risk register + +| # | Risk | Mechanism | Severity | Notes / mitigation option | +|---|---|---|---|---| +| R1 | **Rule adherence drops on flash/cheap models** | Same prompt, weaker tail-following, esp. rules arriving as tool-result hints | HIGH (this incident) | Pin convention-heavy work to pro/strong model; keep rules early & terse | +| R2 | **`.hermes.md` shadowing hides AGENTS.md edits** | One-source startup rule | HIGH | Remove or merge the 4 dual-file repos; see D2 | +| R3 | **Context coverage is cwd-bound; CLI sessions run from home** | Startup loader only looks at session cwd | HIGH | Launch convention / terminal.cwd; ACP workspaces; see D3 | +| R4 | **Cross-project rules live only in `_shared` (cwd-gated) + memory (2,200-char cap, 98% full)** | No global rules slot other than SOUL.md | MEDIUM | SOUL.md is identity-only by design; add a global rules *skill* or widen memory; see D1 | +| R5 | **Skill doc for this exact topic is stale** | hermes-agent skill `project-context-files.md` predates chain-merge + hints + AGENTS.override.md | MEDIUM | Refresh skill reference (offered separately) | +| R6 | **Compression runs against per-model window** | threshold 0.5 of window; model switch mid-session = new window, cache break, possible immediate compaction | MEDIUM | 1.31M window makes this rare; short sessions never hit it | +| R7 | **Reasoning overrides: `high` default + `reasoning_effort: high`** | Cost/compliance tradeoff on flash; verbose reasoning may crowd instruction weight | MEDIUM | Re-check per model tier; not a correctness bug | +| R8 | **No fallback model configured** | 429/529/503 = hard failure, no failover (config has commented fallback block) | LOW | Enable fallback_model for resilience | +| R9 | **Model strategy docs drifted from reality** | MODEL_STRATEGY_GUIDE.md tier tables ≠ config.yaml (nous/deepseek); references dead scripts | LOW | Refresh or archive; it misleads future model picks | +| R10 | **Memory/user profile at 98% capacity** | Conventions can no longer grow there; new rules pushed to skills/repos where loading is conditional | MEDIUM | Consolidate stale entries; move procedural rules to skills | +| R11 | **ACP session context depends on Zed workspace root** | Zed sends cwd per session (`new_session(cwd)`); repo-rooted workspaces load that repo's single context file; multi-root/home leaves none | MEDIUM | Keep one repo per Zed workspace; confirm the workspace root carries the rules file you expect | + +--- + +## 6. Decision points (nothing below was executed — review first) + +**D1 — Where should cross-project rules (security, plan intake, no-emojis, +a/b/c style) live so every session sees them regardless of cwd or model?** +- a) A global rules skill (e.g. `hermes-conventions`) installed in the harness — loads by relevance, cwd-independent (recommended) +- b) `SOUL.md` expansion (always loaded, but it is the identity slot and shares budget with model identity) +- c) Keep as-is: memory (near-full) + `_shared/AGENTS.md` (cwd-gated) + hope for hints + +**D2 — The 4 dual-file repos (timer, slicer, sifter, acct-tools): which file wins?** +- a) Delete `.hermes.md`, let the full `AGENTS.md` load at startup (recommended — matches the "AGENTS.md primary" convention) +- b) Merge AGENTS.md content into `.hermes.md`, delete AGENTS.md (Hermes-native, loses Claude Code portability) +- c) Keep both but make `.hermes.md` a faithful superset and treat AGENTS.md as Claude-only + +**D3 — CLI sessions from home load no project context. Change the launch convention?** +- a) Set `terminal.cwd` (or launch `hermes` from the repo) whenever a session is repo-scoped (recommended) +- b) Add a home-level `.hermes.md`… not supported: `.hermes.md` walks stop at the git root, and home has none — file would be ignored outside home anyway +- c) Accept home = memory/SOUL-only sessions; do repo work in Zed ACP workspaces + +**D4 — Model policy going forward (documented where?)** +- a) Flash default for mechanical work; pro/strong model for planning, review, convention-heavy asks (recommended — matches MODEL_STRATEGY intent) +- b) Single default model, accept adherence variance +- c) (Pick a specific pinned model for rules-heavy sessions, e.g. deepseek-v4-pro) + +--- + +## 7. Appendix — evidence tables + +### 7.1 Model history (state.db, sessions table) + +| Period | Model(s) | Sessions | +|---|---|---| +| ≤ 2026-07-30 | poolside/laguna-xs-2.1:free, laguna-s-2.1:free | 19 | +| 2026-08-08 09:12 (last) | moonshotai/kimi-k3 | 5 | +| 2026-08-08 16:16 (last) | deepseek/deepseek-v4-pro | 32 | +| 2026-08-09 → 08-13 | deepseek/deepseek-v4-flash | 19 | +| 2026-08-14 → now | deepseek/deepseek-v4-flash-0731 | 40 | + +All CLI rows: `cwd=/Users/patricksingletary`, no git_repo_root. ACP rows store no +cwd column value; per-session cwd comes from the editor. + +### 7.2 Per-repo instruction files (all < 3 KB — no truncation involved) + +| Repo | AGENTS.md | .hermes.md | What Hermes loads at startup (cwd = repo root) | +|---|---|---|---| +| _shared | 1,983 B | — | AGENTS.md (shared conventions; public-repo rules) | +| zodiac | 2,716 B | — | AGENTS.md | +| timer | 1,705 B | 914 B | .hermes.md ONLY (AGENTS.md shadowed) | +| macos-sched-cleanup | 1,291 B | — | AGENTS.md | +| sifter | 1,236 B | 918 B | .hermes.md ONLY | +| web_themes | 1,245 B | — | AGENTS.md | +| slicer | 1,060 B | 965 B | .hermes.md ONLY | +| acct-tools | 832 B | 934 B | .hermes.md ONLY | +| ptharbor | 746 B | — | AGENTS.md | +| ATProtocol-Playground | 300 B | — | AGENTS.md | +| altifier | 287 B | — | AGENTS.md | +| verifier | 185 B | — | AGENTS.md | +| Interactive_Journal | 298 B | — | AGENTS.md | + +### 7.3 Key source citations (installed build 57f05e21) + +- `agent/prompt_builder.py` `build_context_files_prompt`: "Only ONE project + context type loads, first found wins: .hermes.md/HERMES.md (walk to git root) + → AGENTS.md chain (git root → cwd) → CLAUDE.md (cwd) → .cursorrules… SOUL.md + from HERMES_HOME is independent and always included." +- `agent/prompt_builder.py` `_agents_md_directory_chain` / `_load_agents_md`: + chain merge, per-dir `AGENTS.override.md` > `AGENTS.md` > `agents.md`, content + dedup, merged-chain budget. +- `agent/subdirectory_hints.py`: `_HINT_FILENAMES = [AGENTS.override.md, + AGENTS.md, agents.md, CLAUDE.md, claude.md, .cursorrules]` (no .hermes.md); + `_MAX_HINT_CHARS = 32_000`; hints appended to tool results; working dir + pre-marked loaded (startup handles it). +- `acp_adapter/session.py`: `new_session(cwd=…)` / `update_cwd(session_id, cwd)` + — cwd arrives from the ACP client (Zed) per session. +- `~/.config/zed/settings.json`: `agent_servers.hermes-agent` → + `/Users/patricksingletary/.local/bin/hermes acp`, `default_mode: dont_ask`, + tool permissions `default: confirm`. +- `~/.hermes/config.yaml`: `model.default deepseek/deepseek-v4-flash-0731`, + `model.provider nous`, `reasoning_overrides.default: high`, + `compression.threshold 0.5 / target_ratio 0.2 / protect_first_n 3 / + protect_last_n 20`, `memory.memory_char_limit 2200 / user_char_limit 1375`, + hooks wired to both hook scripts, `fallback_model` commented out. -- 2.51.2