This repository has no description
hm-trm PLAN.md
11 kB
Markdown
at main

DSEM Plan — Democratic Social-Ecological Model #

Living document. Started 2026-08-12. Status (2026-08-25): Phases 0–2 done, Phase 3 well under way. TRM reproduced at 84.79% (full test set); capacity-floor curve: D=512 84.8% / D=256 67.8% / D=128 38.8% / D=64 13.6%. Phase 3 mesoscale results so far: same-information diversity (symmetry standpoints, role embeddings) is null; distributing the evidence is the one mechanism that creates collective capability (solo D=64 at half evidence 0.0% → K=2 colony 10.9%); equal influence is load-bearing (confidence-weighted aggregation halves the colony to 4.7%) while its form is not (vote ≈ mean, boss ≈ quorum halting at K=2); K=2 beats K=4 in every cell of the K × view-fraction 2×2, denting the earlier coverage interpretation (see notes/log.md 2026-08-25). Now running: Batch 4, heterogeneous colonies (v1a+ — elder+juveniles vs untied vs solo-elder controls). Full detail in notes/log.md; design rationale in notes/colony-harness-design.md.

Standing lesson from the first null result #

Diversity must change what a unit knows, not merely how it is indexed. Any standpoint drawn from a symmetry of the task is epistemically empty by construction: the unit solves an equivalent problem and (with shared weights) returns the same answer. Every diversity mechanism we test from here should be checked against the question "does this decorrelate errors?" — measurable directly via scripts/analyze_colony.py (oracle-vs-mean-unit gap) without waiting for a full training run to disappoint us.

The idea #

Replace TRM's recursive self-correction with discursive correction among several tiny, differently situated models that improve a shared answer by interacting — aiming to operationalize social epistemology and ecological self-government (see notes/hrm-brain-claims-and-embodied-critique.md for the theoretical grounding, and refs/ for cached sources). Working hypothesis, after Farrell & Shalizi: a diverse, egalitarian, deliberating collective explores rugged solution landscapes better than a single recursive model of the same size — and this should show up as generalization.

Decisions taken (2026-08-12) #

  • Hardware: dev on the Mac laptop (MPS/CPU, tiny configs); full runs on the dedicated box — Ryzen AI Max+ 395 (Strix Halo APU: Radeon 8060S iGPU ~40 RDNA 3.5 CUs, gfx1151, 64 GB unified LPDDR5X, ~256 GB/s). SSH access from the laptop: TBD when Phase 0 starts.
  • Benchmarks: reproduce on Sudoku-Extreme first (cheapest, and where TRM's own ablations live); decide the wider benchmark set after reproduction.
  • Mechanism priorities: (1) situatedness / standpoint diversity, (2) collective self-government. Speech acts / constrained public language: later, not immediately.
  • Methodology: exploratory first — chase signal on scaled-down configs, backfill strict parameter/compute-matched controls once a mechanism shows promise.
  • Colony units (decided 2026-08-12): units are much-tinier TRMs that perform badly alone. Two regimes, tested in this order: mesoscale — K≈4–16 shrunken TRMs (D=64–128, ~85k–330k params each, whole-board view through distinct standpoints; the regime where votes, quorum, reputation, and domination are literal); then microscale — many ~5–50k-param units bound to loci (cell / row / column / box) with partial views, a learned constraint-propagation ecology (protist regime). Units are clonal by default: one shared tiny network, differentiation via situation + private latent + optional learned role embedding — keeps capacity in the small-data sweet spot and lets every unit learn from every situation. Prior-art note for microscale: adjacent to Recurrent Relational Networks (Palm et al. 2018, in TRM's own refs) and Neural Cellular Automata; novelty = TRM's small-data machinery (deep supervision, recursion, EMA, learned halting) + standpoints + self-government, in the 1k-example regime, vs their homogeneous synchronized message passing on ~180k puzzles.

Phase 0 — Environment & port (laptop + box) #

The official code (github.com/SamsungSAILMontreal/TinyRecursiveModels, now archived — vendor or fork it, don't depend on upstream) assumes Python 3.10 + CUDA 12.6 + PyTorch nightly + adam-atan2 (CUDA extension). The box is AMD, so:

  • Linux + ROCm PyTorch on the box (recent ROCm supports gfx1151; verify with a calibration run — if ROCm is rough, fall back to scaled-down runs while we consider a ~$1–2/h cloud L40S for headline runs, ≈$20–40 per full Sudoku training).
  • Make the training code device-agnostic (cuda/rocm/mps/cpu): replace adam-atan2 with a pure-PyTorch implementation; the attention-free MLP variant (the best Sudoku config) needs no FlashAttention; use SDPA where attention is needed.
  • Smoke test end-to-end on the laptop with a tiny config (few epochs, small batch).

Wall-clock expectations (from the TRM README): Sudoku-Extreme ≈18h on 1×L40S; Maze-Hard <24h on 4×L40S; ARC ≈3 days on 4×H100. The 8060S is plausibly 5–10× slower than an L40S for this workload → a full Sudoku run ≈3–7 days on the box. Fine for headline runs; iteration happens on scaled-down configs. ARC-AGI is out of local reach; cloud-only if we ever want it. Measure a steps/sec calibration early to replace these guesses with numbers.

Phase 1 — Reproduce TRM (Sudoku-Extreme) #

  1. Build the Sudoku-Extreme dataset (1k train examples × 1k augmentations, per the repo).
  2. Scaled-down reproduction on the box to validate the pipeline (hours, not days).
  3. Full run: TRM-MLP, T=3, n=6, 2 layers, EMA — target 87.4% ±3 (paper Table 1). Expect some seed variance; treat within-a-few-points as success.
  4. Optionally reproduce one or two ablation rows (e.g. T=2,n=2 → 73.7) to confirm the harness is sensitive enough to detect differences between variants — this is our instrument calibration, since DSEM claims will rest on deltas of a few points.
  5. Capacity-floor sweep: train solo TRMs at D=256/128/64 (≈1.3M/330k/85k params; FLOPs scale ~D², so D=128 is ~1/16 of a full run) and chart accuracy vs width. This curve is the x-axis of the whole project — DSEM's claim is that collectives of below-the-floor units climb back above it — and each point doubles as the K=1 baseline for Phase 3. The sweep costs less than one full reproduction run.

Deliverable: repro/ results + capacity-floor curve + a lab-notebook entry in notes/.

Phase 2 — Colony harness (refactor, no new science) #

Refactor training/inference into a colony harness where TRM is the K=1 degenerate case:

  • K agents, each with its own latent z_i and a standpoint transform ρ_i (for Sudoku: band/stack/digit permutations — symmetries of the puzzle, so each agent literally sees a different-but-equivalent problem and answers are mapped back through ρ_i⁻¹; the transform abstraction must also support view restriction — masking to a locus's neighborhood — so the same harness covers the microscale regime);
  • clonal weights with optional per-agent role embeddings (cheap heterogeneity without capacity blowup);
  • a shared, public answer y (the blackboard — stigmergic medium: agents act on it and perceive one another only through it);
  • a decision rule (per-cell vote / averaging in embedding space / learned combiner);
  • a halting policy (per-agent halt heads aggregated by quorum, replacing TRM's single head).

Free controls this gives us: K=1 (=TRM); K agents, no interaction, vote at the end (the Condorcet/ensemble control — note TRM's ARC protocol already does aggregation-without- deliberation via 1000-augmentation majority voting at test time, so this control is also the honest reading of prior work).

Phase 3 — DSEM v1 experiments (exploratory, scaled-down configs) #

All v1 experiments run in the mesoscale regime: agents are shrunken TRMs from the capacity-floor sweep (start at the widest clearly-below-floor width, likely D=128), so the question is always "do K weak units together beat one strong unit at matched cost?".

v1a — Situatedness. Clonal colony (shared weights, so parameter count stays ~one tiny unit's): K ∈ {2, 4, 8, 16} agents with distinct standpoints ρ_i, private z_i, shared y updated each round. Key contrasts: standpoint diversity vs seed/init diversity vs K=1 (both at unit width and at TRM's D=512); interaction-during-reasoning vs vote-at-the-end.

v1a+ — Heterogeneous capacities (Robin, 2026-08-13): the colony may do better with differently sized members than with near-identical clones — diversity of ability, not just standpoint (one larger + several smaller units, or a spread of widths from the capacity-floor sweep). Breaks pure weight-sharing (weights shared within a size class only) and requires care in parameter-matched comparisons (match total params and total compute at the colony level). Also the natural place to look for emergent division of labor: do big and small members take on different roles (proposer vs checker)?

v1b — Self-government. On the best v1a setup: collective halting (quorum vs single boss head); influence rules (equal vote vs confidence-weighted vs learned reputation — the last reintroduces hierarchy, which is the point of testing it); measure domination (does one agent capture the blackboard?) and crowding-out effects.

Diagnostics beyond accuracy (the epistemic-quality dashboard):

  • disagreement between agents over rounds, and whether disagreement predicts error (collective calibration);
  • the Mill test: how often an initially-minority correct proposal ends up as the collective answer through interaction (vs being crushed in the vote-only control);
  • robustness probes (train-distribution vs harder puzzles, corrupted givens) — the democratic hypothesis predicts robustness gains even where raw accuracy ties.

Later (v2a) — Microscale ecology. Once the harness and mesoscale results exist: many ~5–50k-param clonal units bound to loci with restricted views, acting on the shared board (stigmergic constraint propagation). Compare against the mesoscale colony at matched compute, and position against RRN/NCA prior art (see Decisions).

Later (v2b) — Speech acts: a constrained public message channel (small discrete token vocabulary) alongside or instead of blackboard deltas.

Phase 4 — Benchmarks & rigor backfill #

Once a mechanism shows signal: full-scale Sudoku-Extreme runs with parameter-matched AND compute-matched TRM baselines plus the no-communication ensemble; then extend to Maze-Hard (and revisit ARC-AGI/cloud). Report accuracy, params, forward passes, wall-clock, plus the epistemic diagnostics.

Working practices #

  • refs/ cached sources; notes/ lab notebook + research notes; PLAN.md this file, updated as decisions land.
  • Experiment tracking: local JSON/CSV + plots first; W&B optional later.