Lab notebook #
Newest entries at the top.
2026-08-25 — Batch 3 complete: K=2 wins every cell, and the coverage story takes a hit #
All three config-only runs finished (queue3, 2026-08-23 16:48 → 2026-08-25 09:53).
k2-vote final: 10.14% — discrete one-unit-one-vote vs averaging logits: a tie with mean aggregation (10.9%), well inside run noise. k2-boss final: 10.97% — one member holding sole authority over when deliberation stops, also a tie with quorum (10.9%). Two caveats on boss: its trajectory was unusually noisy (8.08 → 9.11 → 9.19 → 11.17 → 10.97 over the last five evals), and K=2 is the weakest possible test of concentrated authority — a boss there is half the electorate. If halting hierarchy hurts anywhere, it will be at K≥4; worth a rerun only if boss-like rules become load-bearing later.
The decision-rule scoreboard (D=64, half evidence, distributed) #
mean 10.9% ≈ vote 10.14% ≈ boss-halting 10.97% ≫ confidence-weighted 4.7%.
Whether influence is equal is worth a factor of two; how equality is implemented — logits or ballots, quorum or boss — moves nothing at K=2. Farrell & Shalizi's equality leg keeps looking like the operative ingredient, not any particular voting mechanism.
k2-evi25 final: 8.21% — after sitting at 9.3–9.9% for seven consecutive evals, it dipped on the last one. This end-of-run dip is now a pattern (k2-vote 10.92 peak → 10.14 final; k4-evi25 7.1 → 6.4); single final evals may systematically understate the plateau by ~1 point. Something to keep in mind before reading small final-vs-final gaps as real — or to fix by averaging the last three evals when reporting.
The K × view-fraction 2×2, complete (mean, quorum) #
| view per unit | K=2 | K=4 |
|---|---|---|
| 50% | 10.9% | 9.8% |
| 25% | 8.21% | 6.4% |
Monotone in both directions: K=2 beats K=4 in both rows, more per-member evidence helps at both K. But the off-diagonal comparison dents the 2026-08-21 "coverage, not crowd size" interpretation: at quarter evidence, the K=2 pair's union no longer covers the grid — it collectively holds roughly half the evidence (by design; see run_queue3.sh) — while K=4 at quarter evidence tiles the whole grid. The half-blind pair still wins, 8.21% vs 6.4%. Coverage cannot be the dominant variable. What survives from the 08-21 analysis is complementarity/indispensability (every K=2 member is load-bearing; K=4 members are partly redundant), plus possibly a coordination cost that grows with K. Pairs deliberate better.
Sharpened headline comparison: ~50% of the evidence concentrated in one head solves 0.0%; the same total evidence split across two heads that must talk solves 8.21%. Distribution itself — not the amount of information — is what creates the capability. Single seed, as ever.
Batch 4 launched — heterogeneous colonies (PLAN v1a+) #
run_queue4.sh started 2026-08-25 12:36 on halfmind: hetero-mixed-1elder [256,64,64,64],
then hetero-untied-k4 [64,64,64,64] (same members, no elder), then hetero-solo-d256
(the elder alone at half evidence). Design rationale in the script header: if mixed beats
both controls, the elder contributes something as a colony member it cannot contribute alone
— and param count alone can't explain it. Hetero code was smoke-tested on the laptop
(runs/Sudoku-smoke-ACT-torch/hetero-smoke) and the box's source tree verified identical
before launch.
2026-08-23 — Batch 2 complete: unequal influence halves the colony #
colony-d64-k4-evihalf-conf final: 4.7% — identical to the 9.8% run in every respect except that proposals are weighted by each unit's own confidence instead of averaged equally. Confidence-weighting destroys half the collective's value, the largest single effect any intervention has produced here.
Batch 2, all three #
| configuration | view/unit | K | aggregation | final |
|---|---|---|---|---|
| K=2 half, mean | 50% | 2 | mean | 10.9% |
| K=4 half, mean | 50% | 4 | mean | 9.8% |
| K=4 quarter, mean | 25% | 4 | mean | 6.4% |
| K=4 half, confidence | 50% | 4 | confidence | 4.7% |
Why unequal influence hurts: with distributed evidence a unit's confidence is not calibrated to its correctness — a member can be entirely certain about the cells it can see and confidently wrong about the rest. Weighting by confidence concentrates influence in whoever is most certain rather than whoever is most informed, and it performs far worse than one-unit-one-vote. This is a computational instance of the second leg of Farrell & Shalizi's argument (as cited in Stewardship): equality matters because it stops one party imposing its solution from an excess of power. Single seed — but a 5.1-point gap is well outside the ±1–2 point noise these runs show.
Batch 3 launched (~41h, config-only) #
Built on K=2, the best and cheapest configuration:
-vote: discrete one-unit-one-vote instead of averaging logits. Given that equal weighting beat confidence weighting, the form of equal weighting is now worth isolating.-boss: one member decides when the group stops deliberating. Quorum halting already costs ~2 points; does concentrating that authority cost more? Self-government, tested directly.k2-evi25: two units at quarter evidence — the most information-starved colony yet, and it completes the K × view_fraction 2×2.
Next after that: heterogeneous colonies (Robin's mixed-size idea, PLAN v1a+), which needs real code — per-size weight sets and a token-space blackboard so units of different widths can share the public answer at all.
2026-08-22 — Quarter-evidence: 6.4%. My prediction was wrong; here is the correction. #
colony-d64-k4-evi25 final: 6.4%. I pre-registered that four units at 25% evidence should match two at 50% "if coverage is what matters". It did not (6.4% vs 10.9%).
Where my framing was wrong: coverage was constant across every one of these runs —
view_enforce_coverage=True guarantees the union of member views is the whole grid. So the
prediction was never really about coverage; the variables that actually differed were
per-member visibility and K. Stating it as "coverage vs crowd size" was sloppy and
made a prediction that could not have been informative in the way I claimed.
The distributed-evidence family, complete #
| configuration | view per unit | units | final |
|---|---|---|---|
| solo, half evidence | 50% | 1 | 0.0% |
| colony, half evidence | 50% | 2 | 10.9% |
| colony, half evidence | 50% | 4 | 9.8% |
| colony, quarter evidence | 25% | 4 | 6.4% |
Three things fall out, and they are monotone and consistent:
- One → two members is the whole phenomenon. 0.0% → 10.9%. Capability from nothing.
- Past two, extra members do not help. K=4 is slightly worse than K=2 at equal per-member visibility, at double the compute — and the ablations explain why: at K=4 each clue is seen by ~2 units, so members are partly redundant (drop one → 1.82%), while at K=2 each member is indispensable (drop one → 0.00%).
- Per-member visibility matters independently of group coverage. Halving each member's view (50%→25%) at identical total coverage costs ~3.4 points. Blinder members can still contribute, but they contribute less — plausibly because a unit seeing ~20 cells cannot form a coherent enough proposal to be worth much to the others.
The most striking single comparison: four units at 25% evidence reach 6.4%, while one unit at 50% evidence — twice as sighted as any of them — reaches exactly 0.0%. A committee of the quarter-sighted beats a half-sighted individual, decisively.
Note evi25 was still climbing at its last evals (2.5 → 3.1 → 4.3 → 5.3 → 6.5 → 7.1 → 6.4) where the other runs had plateaued, so 6.4% may understate its asymptote; blinder members appear to need longer to learn to use the shared channel. Worth a longer run before treating the 3.4-point gap as settled.
colony-d64-k4-evihalf-conf (confidence-weighted aggregation vs mean) is now running, ~24h —
the last of batch 2 and the first test of a decision rule rather than an architecture.
2026-08-21 — K=2 beats K=4: it is coverage, not crowd size #
colony-d64-k2-evihalf final: 10.9% (12.83% when forced to use the full ACT budget) — better than the K=4 colony's 9.8%, at half the compute. Two units, each seeing ~50% of the grid, each individually scoring 0.0%.
Member ablation: K=2 has no redundancy at all #
| units silenced | K=2 | K=4 |
|---|---|---|
| none | 12.83% | 10.20% |
| one | 0.00% | 1.82% |
Silencing either member of the pair gives exactly zero. Silencing one of four left 1.82%, because with four units at half evidence each every clue is seen by ~2 units on average, so some of a silenced unit's evidence survives elsewhere.
Interpretation (predicted before the measurement, then confirmed): the effect is driven by coverage and complementarity, not by crowd size. Two units at half the grid each just barely cover it between them, making both indispensable. Four units at half each are heavily redundant — the extra members duplicate evidence, add compute, and contribute nothing. This also explains the otherwise-odd K=2 > K=4 ordering.
Prediction for colony-d64-k4-evi25 (running): four units at quarter evidence have the same
total coverage as two at half, sliced finer. If coverage is what matters, it should perform
comparably to K=2 despite far blinder members.
Free finding: collective halting costs ~2 points #
| K=2 | K=4 | |
|---|---|---|
| normal eval (quorum halting decides when to stop) | 10.9% | 9.8% |
| forced to use the full 16-step ACT budget | 12.83% | 10.20% |
The colony stops deliberating before it should — a jury that adjourns too soon. This is a
Phase 3 self-government result arriving for free, and it makes the halt_rule comparison
(quorum vs any vs boss) worth running on its own merits rather than as an afterthought.
Deliberation dynamics #
Pre-discussion disagreement: 65.95% (K=2) vs 92.76% (K=4) — as expected, since the metric counts cells where any unit dissents and four units dissent more readily than two. Both fall to ~2–3.5% after three rounds. Final per-unit rates [12.83, 12.76] versus colony 12.83: each member ends up knowing almost everything the group knows.
2026-08-20 — Roles ablation done (13.9%); full-evidence colonies all null #
colony-d64-k4-roles final: 13.9% — role embeddings, symmetry standpoints, full evidence. Slightly below the same colony without roles (15.4%), both indistinguishable from a lone unit (13.6%). The mid-run streak where roles led for three consecutive evals (9.7/12.5/14.0 vs solo 8.3/11.6/13.1) evaporated: it went on to trail at four of the next five evals. Another reminder that these narrow runs need finals, not trajectories.
The complete Phase 3 table so far #
| configuration | evidence per unit | final |
|---|---|---|
| solo D=64 | full | 13.6% |
| colony K=4, symmetry standpoints | full | 15.4% |
| colony K=4, symmetry + role embeddings | full | 13.9% |
| colony K=4, distributed evidence | ~50% each | 9.8% |
| solo D=64 | ~50% | 0.0% |
The pattern is now clear and consistent across three independent mechanisms: when every unit sees the same information, differentiating them — by viewpoint permutation, by learned role embedding, by anything tried so far — produces no collective benefit. Only distributing the evidence creates capability that does not exist in any member.
Note the apparent paradox in the table: the distributed-evidence colony's 9.8% looks worse than the full-evidence colonies' ~14–15%, but it is the only row where the colony is doing anything at all. The full-evidence rows are four copies of a unit that already scores 13.6% alone; the distributed row is four units that each score 0.0% alone.
run_queue2.sh launched automatically on completion: K=2 at half evidence (does the effect
need a crowd?), K=4 at quarter evidence (how blind can members get?), and K=4 with
confidence-weighted aggregation (can a confident minority prevail over the mean?). ~65h.
2026-08-20 — CORRECTED RESULT: every member is load-bearing #
Reran both diagnostics with the fixed instrumentation (gate 7 now pins it to forward()).
The corrected picture is stronger than the buggy one, and the retracted "consensus collapse"
was hiding a much better finding.
Member ablation — silencing units at inference (2,304 puzzles) #
| units speaking | solved | vs full colony |
|---|---|---|
| all 4 | 10.20% | — |
| 3 (drop 1) | 1.82% | −8.38 |
| 2 (drop 2) | 0.17% | −10.03 |
| 1 | 0.00% | −10.20 |
Every member is essential. Removing one destroys 82% of the capability; reducing to one gives exactly 0.00% — independently reproducing the separately-trained solo control's 0.0% by a completely different route. Two unrelated experiments converge on the same number, which is about as good as confirmation gets here.
Deliberation dynamics (corrected) #
| disagreement between units | |
|---|---|
| before discussion (blackboard empty) | 92.76% |
| after 3 rounds | 3.49% |
Units start almost totally disagreeing — they are, correctly, looking at different puzzles — and converge through the shared blackboard. Final per-unit solve rates [10.16, 10.07, 10.16, 10.07] versus colony 10.20%: by the end each member has absorbed nearly everything the group knows, which is what successful pooling looks like. Contrast the earlier buggy reading of 0.00% disagreement, which I wrongly framed as pathological uniformity.
Revised interpretation of convergence: low final disagreement here is the product of successful information sharing, not a failure mode. The consensus-collapse worry may still apply to some configurations — it is a real risk worth testing deliberately — but this run provides no evidence for it, and I should not have asserted it.
The claim, as it now stands #
A lone D=64 unit with ~50% of the clues solves 0.0% of Sudoku-Extreme puzzles, whether trained alone or extracted from the colony. Four such units, each equally blind but holding different clues, solve ~10% by revising a shared public answer. Removing any one of them collapses the group to near-zero. The capability exists only in the collective, every member is necessary to it, and it is not ensembling — averaging four zeros cannot produce ten.
2026-08-19 (late) — RETRACTION: the instrumentation was broken #
ColonyInner.forward_instrumented — the path used by analyze_colony.py and
colony_ablate.py — bypassed the view masks and gave every unit the full question.
When masking was refactored into _question_views, the edit matched a comment line that
existed only in forward(), so the instrumented path silently kept embedding the question
once and sharing it. For the distributed-evidence run (standpoints=none, diversity coming
only from the masks) this made all four units mathematically identical inside the analysis.
Retracted, all from measurements on colony-d64-k4-evihalf:
- "collective lift +8.76 points / mean unit alone 0.00%" — measured with full evidence.
- "consensus collapse, 0.00% disagreement, units learn to echo the blackboard" — the units were fed identical inputs by the measurement. There is no evidence for consensus collapse in this run. This was the more interesting of the two claims and it was an artefact.
- "members are REDUNDANT — silencing one costs nothing" — trivially true of identical inputs.
Not affected (these use forward(), which was always correct):
- Every training and eval number: TRM 84.79%, the capacity curve, colony-sym 15.4%, colony-evihalf 9.8%, solo-evihalf 0.0%. The headline comparison stands.
- The first colony's analysis (per-unit 16.41% ×4, oracle = mean). That run used
standpoints=sudokuwithview_fraction=1.0, andforward_instrumenteddid apply the standpoint permutation — so the "symmetry standpoints are epistemically empty" finding is intact.
Fix: forward_instrumented now calls _question_views, and gate 7 asserts the
instrumented path reproduces forward()'s output exactly at both full and partial evidence.
Reruns of the analysis and ablation are in flight.
Process lesson: an instrumentation path that duplicates model logic will drift from it, and when it drifts it produces confident, plausible, wrong findings — worse than a crash. Any diagnostic that re-implements a forward pass needs an equivalence test against the real one, written at the same time as the diagnostic, not after it misleads you.
2026-08-19 — CONTROL CONFIRMS COLLECTIVE INTELLIGENCE (solo-d64-evihalf) #
solo-d64-evihalf final: 0.0% exact accuracy, 39.2% cell accuracy — a fully trained lone D=64 unit seeing the same ~50% of the givens as a colony member (coverage enforcement off, so it genuinely lacks cells). Flat at 0.0% across every eval from step 19,530 to 65,100: this is not slow learning, it is incapacity.
The headline comparison #
| system | evidence per unit | collectively sees | final exact-acc |
|---|---|---|---|
| solo D=64, half evidence | ~50% | 50% | 0.0% |
| colony K=4 × D=64, half evidence each | ~50% | 100% | 9.8% |
| solo D=64, full evidence | 100% | 100% | 13.6% |
A single unit with half the evidence solves nothing at all. Four such units, each equally blind, solve ~10% of Sudoku-Extreme puzzles by talking to each other. The capability does not exist in any member; it exists only in the group. And this is emphatically not ensembling — averaging four models that each score 0% cannot produce 9.8%.
The colony recovers ~72% of what a unit with complete information achieves (9.8 vs 13.6), while no member ever holds complete information.
Why compute cannot explain it: the colony spends 4× the per-step compute of the control, but the control's deficit is information, not computation — with ~12 givens the puzzle is underdetermined (Sudoku needs ≥17 for a unique solution), so no amount of thinking recovers the answer. Its 39.2% cell accuracy is the signature of a model producing plausible-but-wrong grids: well above the ~11% of random digits, far below the ~70% the colony reaches.
Caveats, stated plainly #
- Single seed per configuration. The direction is unmistakable (0.0 vs 9.8 is not seed noise) but the magnitude is not yet nailed down.
- The colony still under-performs a full-information solo unit (9.8 vs 13.6), so pooling is lossy — consistent with the consensus-collapse finding of 2026-08-18 (units converge to 0.00% disagreement and stop contributing private knowledge).
- Not tested: whether a larger colony, or one with mechanisms that reward dissent, closes the remaining 3.8 points. That is the next question.
Where this leaves the thesis #
The project set out to ask whether differently-situated tiny models correcting one another could produce better outcomes than one model reasoning alone. Two findings, in order of solidity:
- Situatedness must be epistemic, not cosmetic. Standpoints drawn from a task's symmetry group produce zero error decorrelation and zero benefit (2026-08-18). Standpoints that distribute evidence produce capability from nothing.
- A shared channel breeds uniformity. Given a public blackboard and mean aggregation, units learn to echo consensus rather than contribute private knowledge — 0.00% final disagreement. This is the Stewardship essay's cognitive-normalisation worry, reproduced in miniature, and it is the current ceiling on colony performance.
2026-08-19 — DISTRIBUTED EVIDENCE: the colony solves what no member can #
colony-d64-k4-evihalf final: 9.8%. Each of the 4 units sees ~50% of the question's cells (~12 givens; Sudoku needs ≥17 for a unique solution), blackboard fully public, clonal weights.
| run | evidence per unit | final |
|---|---|---|
| solo D=64 | 100% | 13.6% |
| colony K=4, symmetry standpoints | 100% | 15.4% |
| colony K=4, half evidence | ~50% | 9.8% |
Raw accuracy is lower — of course it is, the units collectively have the same information but
each individually has half. The interesting measurement is scripts/analyze_colony.py:
BEFORE any discussion (round 1, blackboard empty): mean unit 0.00% any unit 0.00%
AFTER discussion: colony 8.76%
=> collective lift: +8.76 points
No unit solves a single puzzle from its own evidence; together they solve ~9%. Information is demonstrably pooled through the public blackboard.
Honest caveat: the "before" number is round 1 of 48, so 0% conflates missing information
with missing thinking time — a solo TRM also solves nothing after one round. The clean control
is a fully trained single unit with half the evidence (solo-d64-evihalf, K=1,
view_enforce_coverage=False so the lone unit really is missing cells), now running (~7h).
The theoretical expectation is that it lands far below 9.8%, because 12 givens leaves the
puzzle underdetermined — but expectation is not measurement.
Second finding — consensus collapse. Disagreement between units is 0.00% at the end of inference: they converge to identical outputs despite holding different evidence. They learn to read the blackboard rather than to contribute what only they know. This is exactly the uniformity pathology Robin's Stewardship essay describes (cognitive normalisation through a shared channel), arrived at from the opposite direction, and it is a plausible reason the colony does not exceed the full-information runs: once unanimous, extra members add nothing. Levers against it: aggregation that rewards minority-but-correct proposals (vote/confidence rather than mean), a bottlenecked token blackboard, role embeddings, heterogeneous capacities.
Process note: pkill -f <pattern> over ssh killed my own session for the second time
(pattern matched the ssh command line). Use stop_run.sh's PID file, or kill by PID resolved
on the box. Documented once, repeated anyway — hence this second, blunter note.
2026-08-18 — FIRST COLONY RESULT: null, with a precise cause #
colony-d64-k4 final: 15.4% vs solo D=64: 13.6% — a +1.8 point difference that is not a result. Mean of the last five evals: colony 12.4%, solo 11.6%; both runs swing ±4 points between evals. Single seed, no separation.
Curves (exact-acc %):
| step | 6510 | 13020 | 19530 | 26040 | 32550 | 39060 | 45570 | 52080 | 58590 | 65100 |
|---|---|---|---|---|---|---|---|---|---|---|
| colony K=4 | 1.7 | 8.3 | 11.4 | 11.2 | 14.6 | 14.6 | 13.1 | 9.3 | 9.7 | 15.4 |
| solo D=64 | 2.3 | 8.2 | 11.6 | 13.1 | 13.0 | 12.9 | 12.9 | 9.7 | 9.0 | 13.6 |
Why: the units are functionally identical (scripts/analyze_colony.py) #
At full ACT budget, 3,072 test puzzles:
| round | aggregate | best unit | mean unit | cell disagreement |
|---|---|---|---|---|
| 1 | 16.37 | 16.37 | 16.35 | 4.24 |
| 2 | 16.37 | 16.41 | 16.38 | 4.25 |
| 3 | 16.41 | 16.41 | 16.41 | 4.24 |
Per-unit solve rates, final round: 16.41 / 16.41 / 16.41 / 16.41. The oracle ("any unit right") equals the mean unit: not one puzzle in 3,072 is solved by one unit and missed by another. Errors are perfectly correlated, and perfectly correlated errors make aggregation mathematically incapable of helping — Condorcet and Hong–Page both require independence.
Root cause: I drew the standpoints from Sudoku's symmetry group (band/stack permutations, transpose). Those are precisely the transformations the task is invariant under, so each unit solves an equivalent problem in different coordinates and — sharing weights — returns the same answer mapped back. Diversity drawn from a task's symmetries is epistemically empty by construction. (Note the discussion machinery itself works: disagreement falls 19.5% → 8.3% → 4.2% across rounds, i.e. units genuinely converge through the blackboard. They converge on being wrong together.)
Note also: colony 16.4% here vs 15.4% at eval — the analysis forces all 16 ACT steps whereas training-time eval lets quorum halting stop early, so the colony halts slightly too eagerly. Minor, but the halt rule is costing ~1 point.
What this implies for the design #
Situatedness must mean different access to the problem, not different coordinates on the same access. Levers, in rough order of theoretical interest:
- Partial views / distributed evidence — each unit sees only part of the givens, the blackboard stays public. Shared public answer + privately held facts is the classic distributed-knowledge setup (Hutchins' navigation team); it forces integration and is the honest version of "situated". Needs a small code change (mask input per unit).
- Heterogeneous capacities (Robin's idea, PLAN v1a+) — different widths are genuinely different function classes, so errors should decorrelate.
- Independent (non-clonal) weights — the cheapest pure test of "does error decorrelation help at all", at K× the parameters.
- Role embeddings — break the unit symmetry with a learned per-unit bias so units can
specialize. Launched now as
colony-d64-k4-rolesbecause it needs no new code and attacks the diagnosis directly.
Ordering rationale: (4) runs today for free; (1) is the flagship and gets built next; (2) and (3) follow and are cheap variations once (1)'s machinery exists.
2026-08-17 — FIRST DSEM EXPERIMENT LAUNCHED (colony-d64-k4) #
Config: K=4 units at D=64, clonal weights (397,826 params total — same as one D=64 unit), fixed per-unit Sudoku standpoints, latent blackboard, mean aggregation, quorum halting, otherwise the identical protocol to every baseline (50k epochs, batch 768, EMA, subsampled evals). Running at 1.40 s/it compiled → ~25h ETA.
Why this configuration first: latent blackboard + mean aggregation is the minimal departure from TRM, so if it fails we know the problem is the colony idea rather than the token-blackboard bottleneck. Standpoints are the one thing we do change, because situatedness is the hypothesis under test.
The two comparisons it sets up:
- vs its own kind alone: solo D=64 = 13.6%, same parameters. Any gain is created purely by K units interacting.
- compute-matched vs the giant: colony 1.40 s/it vs D=512 TRM 2.40 s/it (colony is cheaper), D=512 = 84.8% with 12.6× the parameters.
Colony cost measurements (batch 768, uncompiled probe; scripts/time_colony.py):
| D | K=1 | K=2 | K=4 | K=8 |
|---|---|---|---|---|
| 64 | 0.64 | 1.24 | 2.40 | OOM |
| 128 | 1.21 | 2.29 | OOM | OOM |
Cost is linear in K. Memory is the binding constraint: the amdgpu driver exposes only
30.5 GiB of the box's 64 GB (default GTT limit), and the colony holds K× activations in the
backprop'd final round. K=8 at D=64 and K≥4 at D=128 do not fit at batch 768. If we need them:
amdgpu.gttsize=49152 kernel param (sudo + reboot) roughly doubles headroom; gradient
checkpointing is the code-side alternative.
Two GPU-only bugs found and fixed in colony.py before launch (CPU generator inside a cuda
device context; buffers pinned to CPU not following the ambient device). Added gate 5 to
scripts/check_colony.py — build under a device context, assert every buffer lands on the
accelerator. CPU-only gates could not have caught either.
2026-08-17 — CAPACITY-FLOOR SWEEP COMPLETE — the DSEM baseline curve #
| width | params | final exact-acc | s/step | run time |
|---|---|---|---|---|
| 512 | 5,034,506 | 84.8% (84.79 full test) | 2.40 | 44h15m |
| 256 | 1,486,602 | 67.8% | 1.10 | 20h21m |
| 128 | 695,690 | 38.8% | 0.66 | 12h13m |
| 64 | 398,538 | 13.6% | 0.39 | 7h17m |
This is the x-axis for every DSEM claim. Accuracy falls much faster than parameters: D=64 keeps 7.9% of the parameters but only 16% of the accuracy. A lone D=64 unit is comprehensively broken on Sudoku-Extreme — which is exactly the "situated neuron" regime Robin asked about, and it leaves enormous headroom for a colony to demonstrate value.
Curves (exact-acc %, subsampled evals every 6,510 steps):
| step | 6510 | 13020 | 19530 | 26040 | 32550 | 39060 | 45570 | 52080 | 58590 | 65100 |
|---|---|---|---|---|---|---|---|---|---|---|
| D=512 | 19.3 | 59.7 | 71.1 | 76.0 | 82.3 | 84.7 | 85.4 | 85.1 | 85.3 | 84.5 |
| D=256 | 12.1 | 18.3 | 46.7 | 53.6 | 59.3 | 62.9 | 63.8 | 64.3 | 65.8 | 67.8 |
| D=128 | 2.8 | — | 9.0 | 22.6 | — | 22.1 | 11.8 | — | 35.7 | 38.8 |
| D=64 | — | 8.2 | — | 13.1 | — | — | 12.9 | — | 9.0 | 13.6 |
Because these per-step evals exist for every width, colony runs can be compared against solo baselines at matched training steps without rerunning anything — useful if we ever shorten colony runs.
Every width shows a delayed sharp transition rather than smooth learning, and the narrow ones are unstable (D=128 dipped to 11.8% mid-run before recovering to 38.8%). Multiple seeds will be required before believing any colony-vs-solo difference of less than ~5 points.
2026-08-17 — Capacity-floor sweep: D=128 done (38.8%) #
D=128 final: 38.8% (cell 79.2%, q_halt 100%) in 12h13m at 0.66 s/it. Curve so far: 84.8 (D=512) → 67.8 (D=256) → 38.8 (D=128).
D=128 training was violently non-monotonic: 2.8 → 9.0 → 22.6 → 22.1 → 11.8 → 35.7 → 38.8. The dip at step 45,570 came with q_halt accuracy falling to 94.2% — a partial collapse of the kind the TRM paper says EMA is meant to prevent (we do use EMA; it softened but did not eliminate it). Narrow units are not just weaker, they are less stable. Two implications:
- For DSEM: colony members will be individually unreliable. That is arguably the point — an ecology of unstable units that stabilizes collectively would be a strong result — but it means colony runs need multiple seeds before any claim, since single-run noise at these widths is ±10 points mid-training.
- Methodological: I twice extrapolated mid-run and was wrong both times (called D=256's plateau ~3 evals early; read D=128's dip as a ceiling when it recovered to 38.8%). Report finals, not trajectories.
D=64 running at 0.39 s/it (~7h), 8.2% at step 13,020 — tracking close to D=128's early curve, consistent with the param-count finding that D=128 and D=64 differ by less than their names suggest (0.70M vs 0.40M, both dominated by the width-independent token-mixing MLP).
2026-08-16 — Capacity-floor sweep: D=256 done (67.8%) #
D=256 final: 67.8% exact-accuracy (subsampled eval; cell 88.5%, q_halt 99.8%) in 20h21m at 1.10 s/it. So halving the width costs 17 points (84.8 → 67.8) for 3.4× fewer params.
Learning curves compared (subsampled evals, exact-acc %):
| step | 6510 | 13020 | 19530 | 26040 | 32550 | 39060 | 45570 | 52080 | 58590 | 65100 |
|---|---|---|---|---|---|---|---|---|---|---|
| D=512 | 19.3 | 59.7 | 71.1 | 76.0 | 82.3 | 84.7 | 85.4 | 85.1 | 85.3 | 84.5 |
| D=256 | 12.1 | 18.3 | 46.7 | 53.6 | 59.3 | 62.9 | 63.8 | 64.3 | 65.8 | 67.8 |
Observation worth keeping: both widths show a sharp transition (D=512 between evals 1–2, D=256 between evals 2–3) rather than smooth improvement — narrower models reach the same qualitative shift later, not never. Mid-run I predicted D=256 would plateau at 64–68%; it kept creeping up to 67.8%, so "plateau" was slightly premature — late-training gains are small but real. Do not stop sweep runs early on apparent plateaus.
D=128 launched automatically, running at 0.66 s/it (~12h). Its first eval is 2.8% at step 6510 (vs 12.1% for D=256, 19.3% for D=512), consistent with a much lower landing point and an accelerating fall-off — i.e. the below-the-floor regime DSEM needs to beat.
2026-08-15 — Colony harness drafted and validated (Phase 2 first cut) #
Written while the full eval occupied the GPU: dsem/models/recursive_reasoning/colony.py,
dsem/config/arch/colony.yaml, scripts/check_colony.py. Design rationale in
notes/colony-harness-design.md. All four gates pass on laptop CPU and on the box:
- K=1 == TRM, bit-identical (max logit diff 0.00e+00 with weights transplanted). The colony is a strict generalization, so colony results are comparable to the 84.5% anchor.
- All 18 combinations of blackboard × aggregate × halt_rule run forward+backward with finite grads on every parameter.
- Clonal colonies keep parameters constant in K (397,826 at K=1,2,4,8) — diversity is free.
- Standpoints round-trip exactly; 8/8 distinct Sudoku views; agent 0 is the identity view.
Implementation notes: K units are folded into the batch axis (one forward for the colony); standpoints are fixed per-unit gather indices over the 81 puzzle positions (the prefix of puzzle-embedding slots is never permuted); the token blackboard re-embeds the public answer via a soft (differentiable) embedding each round.
Compute note for later matching: a K-unit colony costs ~K× the per-round forward of one unit, so K=8 at D=128 is roughly 2× a single D=512 TRM. Compute-matched comparisons must use measured s/step, not parameter counts.
Still to do: GPU smoke-train a K>1 colony; Phase 3 experiment matrix.
2026-08-15 — PHASE 1 REPRODUCTION COMPLETE (repro-mlp-3) #
RESULT: 84.79% exact-accuracy on the full 422,786-puzzle test set
(cell 94.29%, q_halt 99.93%, lm_loss 0.133) — results/repro-mlp-3-fulltest.json.
Paper/README claim: "around 87% (±2%)". We land 2.6 points below the central value, ~0.2
below the band's lower edge. Best in-training eval was 85.4% at step 45,570.
The 12k subsample predicted 84.54% vs 84.79% actual — a 0.24-point error, inside its ±0.4 s.e., which retroactively justifies the subsampled-eval protocol change (saved ~29h).
Learning curve (subsampled evals, EMA weights, every 5,000 epochs):
| step | 6510 | 13020 | 19530 | 26040 | 32550 | 39060 | 45570 | 52080 | 58590 | 65100 |
|---|---|---|---|---|---|---|---|---|---|---|
| exact% | 19.3 | 59.7 | 71.1 | 76.0 | 82.3 | 84.7 | 85.4 | 85.1 | 85.3 | 84.5 |
| cell% | 72.1 | 86.1 | 89.6 | 91.3 | 93.4 | 94.3 | 94.5 | 94.4 | 94.5 | 94.2 |
Late-training evals fluctuate ~±1 point, so the shortfall vs 87.4% is within ~1–2 seed-noise points. Candidate contributors, none verified: seed variance (paper reports a single run), our pure-PyTorch AdamATan2 vs the fused CUDA kernel, ROCm vs CUDA bf16 numerics. For DSEM this is sufficient: every variant will be measured on this same harness, so comparisons are internally consistent even if our absolute number sits ~2 points low.
- Wall-clock: 44h15m for 65,104 steps at 2.40 s/step + 10 × 5.1 min evals.
scripts/eval_checkpoint.pyvalidated: reproduces the training loop's own eval to the last digit (0.8453776 both ways) — the standalone measuring instrument is trustworthy.- Artifacts:
checkpoints/Sudoku-extreme-1k-aug-1000-ACT-torch/repro-mlp-3/step_65100(+ every 5k step),runs/.../repro-mlp-3/metrics.jsonl,results/repro-mlp-3-fulltest.json.
Parameter counts across widths (scripts/param_counts.py) #
| width | 512 | 384 | 256 | 192 | 128 | 96 | 64 | 32 |
|---|---|---|---|---|---|---|---|---|
| params | 5,034,506 | 2,670,730 | 1,486,602 | 894,538 | 695,690 | 448,810 | 398,538 | 348,266 |
D=512 → 5,034,506 params matches the paper's "5M" for Sudoku exactly — independent confirmation the port is architecturally faithful.
Design-relevant finding for DSEM: in the attention-free (mlp_t=True) variant, params do
not scale as D². The token-mixing SwiGLU acts on the sequence axis
(seq_len 81 + puzzle_emb_len 16 = 97) with an inner dim of 512, costing ~298k params that
are independent of hidden width. Hence the ~348k floor. Consequences:
- Capacity-floor sweep should use D=256/128/64 (below that, little is saved).
- Genuinely tiny microscale units need the token-mixing dim shrunk too (lower
expansionformlp_t, or a smallerpuzzle_emb_len), or the attention variant (params ~4D²/layer, which does shrink) — at the cost of ~13 points on Sudoku per the paper (87.4 vs 74.7). - A colony of K units at D=64 costs ~0.4M × K params: K=8 ≈ 3.2M, still under one TRM.
Capacity-floor sweep launched #
scripts/run_sweep.sh 256 128 64 — same protocol as the reproduction (50k epochs, batch 768,
subsampled evals), sequential, detached. D=256 measures 1.10 s/it vs 2.40 s/it at D=512,
i.e. 2.2× faster for 3.4× fewer params — wall-clock does not follow D² any more than params
do, so the sweep is ~20h + (less for smaller) ≈ 40h total. Expect the D=512 anchor 84.79% to
fall off progressively; the shape of that fall is the x-axis for every DSEM claim.
2026-08-13 — MISDIAGNOSIS CORRECTED: evals were never hanging, just slow #
The error. I called run 1's eval a GPU hang. It wasn't. scripts/eval_timing.py on run 1's
metrics shows eval #1 took 2.96h at a uniform 19.3 s/batch for 551 batches and then
training resumed instantly. What fooled me: print() to a redirected stdout is
block-buffered, so "Processing batch N" lines appear in ~8KB bursts (~275 lines) and the
count sits frozen between flushes. I read a frozen counter as a frozen process, and (in run 2)
zero lines as zero progress. Cost: ~8h of training discarded across two restarts, plus a
reboot Robin had to perform. PYTHONUNBUFFERED=1 is now set in the launcher so a counter that
looks stuck really is stuck.
The arithmetic I should have done first. An eval batch runs the full ACT loop to halt_max_steps=16 (eval always uses max steps), each step = T=3 × (n=6 z-updates + 1 y-update) = 21 network applications → 336 forward applications per batch, vs ~35 forward-equivalents for a training step at 2.4s. Predicted ~23 s/batch; measured 19.3. Entirely expected behaviour. Lesson: cost out the expected wall-time of a phase before calling it pathological.
Was the CWSR reboot needed? No evidence either way now — no hang was ever established.
cwsr_enable=0 is harmless and is what AMD recommends alongside the MES fix, so it stays, but
it is not credited with fixing anything.
Real cost structure (halfmind, paper config): train 65,104 steps × 2.40s = 43.4h; full
test-set eval = 2.96h each × 10 evals = 29.6h → 73h total. Fixed by subsampling in-training
evals: scripts/make_test_subset.py builds data/sudoku-testsub-12k (12,288 of 422,786
puzzles, seed 0, ±0.3% s.e. at p≈0.87) → 16 batches ≈ 5 min per eval, ~50 min total.
The authoritative full-test number comes from scripts/eval_checkpoint.py at the end.
repro-mlp-3 launched with data_paths_test=[data/sudoku-testsub-12k], everything else
identical to the README recipe. ETA ~44h.
New tooling (all rsync'd to the box): scripts/launch_run.sh (PID-file based, unbuffered),
scripts/stop_run.sh, scripts/eval_timing.py, scripts/make_test_subset.py,
scripts/eval_checkpoint.py.
Another footgun burned: pkill -f "python -m dsem.pretrain" from an ssh one-liner matches
the ssh command line itself and kills the session (and any sibling). Use stop_run.sh's PID file.
2026-08-13 — Run 1 post-mortem and repro-mlp-2 relaunch #
- Correction to the hang story: run 1's eval eventually recovered on its own — the stall
was a ~2.6h pathological slowdown (CWSR preemption storms, presumably), not a dead queue.
Eval #1 then completed: test exact-accuracy 20.2% at step 6,510 (10% of training),
cell acc 72.3%, q_halt_acc 99.96%; checkpoint
step_6510saved. Trajectory clearly on track vs the 87% target. - Decision: restart anyway. 9 remaining evals × ~3h stall risk ≈ +27h worst case, plus
permanent-wedge risk each time; and the box needs
cwsr_enable=0for the whole project. Robin applied the modprobe.d option + reboot; verifiedcwsr_enable=0active. - repro-mlp-2 launched (same exact config, fresh seed schedule identical — seed=0 as before; run 1 artifacts preserved under repro-mlp-1). 2.40s/it, ETA ~43.5h train + hopefully ~20min/eval now.
2026-08-13 — Run 1 GPU hang during eval #1; CWSR workaround #
-
repro-mlp-1 trained flawlessly for 4.3h (6,510 steps, train exact-acc up to ~29%), then hung mid-eval at batch 275/551: GPU 100% busy with a stuck queue, python main thread spin-polling at 99.8% CPU, HSA threads in kfd_wait_on_events, no kernel fault logged (unlike the MES fault, which was loud). Signature matches ROCm/ROCm#5590 — Strix Halo GPU hang under sustained compute, fixed/mitigated by
amdgpu.cwsr_enable=0(per #5724's guidance, both the good MES firmware AND cwsr_enable=0 are needed on this platform). -
Fix:
/etc/modprobe.d/amdgpu-cwsr.confwithoptions amdgpu cwsr_enable=0+ reboot. -
Cost: eval #1 never completed → no checkpoint → restart training from scratch (−4.3h).
-
Watchlist for run 2: if it hangs again in eval, next lever is subsampled in-training evals (full test set only at the end) to shrink the sustained-load window — protocol-neutral for training since eval is purely observational (EMA copy, carry untouched).
-
Calibration (paper config, batch 768, MLP variant, torch.compile on): 0.415 steps/s steady → 2.40s/step; loss visibly falling within 130 steps (2.55→2.49). The box is only ~2.4× slower than the paper's L40S.
-
repro-mlp-1 launched (nohup-detached on halfmind): exact README recipe —
epochs=50000 eval_interval=5000, lr 1e-4, wd 1.0, MLP variant, T=3/n=6, ema=True, batch 768 → 65,104 steps, ETA ~43.5h. Target: ~87% ±2 exact accuracy (paper Table 1 / README). Log:runs/repro-mlp-1.out; metrics:runs/Sudoku-extreme-1k-aug-1000-ACT-torch/repro-mlp-1/metrics.jsonl; checkpoints every eval undercheckpoints/. -
Open question: wall-cost of each full-test-set eval (423k puzzles × up to 16 ACT steps); first eval lands ~4.3h in — measure it, and only intervene if egregious.
-
Next after completion: capacity-floor sweep (D=256/128/64) — task #7.
2026-08-12 — halfmind (Ryzen AI Max+ 395 box) setup #
- Access:
ssh halfmindfrom the laptop. Working dir on the box:~/Code/hm-trm(rsync'd from laptop, excluding .venv/data/runs/checkpoints; data synced separately). - Box: Ubuntu 26.04, kernel 7.0.0-27-generic, 61GB RAM, ROCm 7.2.4 at /opt/rocm, amdgpu 6.19.4 DKMS module built for this kernel, all gc_11_5_1 (Strix Halo) firmware present.
- Driver gotcha: amdgpu was not loaded at boot (no journal entries at all — coldplug miss;
the iGPU enumerates as a secondary "Display controller").
/dev/kfdabsent untilsudo modprobe amdgpu. Persist withecho amdgpu | sudo tee /etc/modules-load.d/amdgpu.conf. - Python gotcha: system Python is 3.14 → hydra 1.3.2 crashes at import
(
ValueError: badly formed help string, Python 3.14's stricter argparse). Fixed by using uv-managed Python 3.13.15 (matches laptop). No python3.14-venv apt package installed either, so uv (userspace, no sudo) is the path:uv venv --python 3.13 .venv. - Wheels: official PyTorch stable index has rocm7.0 wheels through cp314;
installed
torch 2.10.0+rocm7.0(bundles its own ROCm 7.0 libs; system 7.2.4 fine). If gfx1151 kernels turn out missing at runtime, fallbacks:HSA_OVERRIDE_GFX_VERSION=11.0.0or AMD's TheRock gfx1151 nightly wheel index. - CPU smoke passed on the box (DSEM_DEVICE=cpu, same tiny config as laptop): train → eval → checkpoint, exit 0. GPU smoke + calibration pending the driver load.
- Secure Boot saga:
modprobe amdgpu→ "Key was rejected by service" — the AMD DKMS module (6.19.4) is signed with an unenrolled MOK key. Fix:dkms uninstallfor the running kernel restores the archived Canonical-signed in-tree module; no reboot needed for the load itself. (Alternative would have beenmokutil --import+ console access at boot.) - Kernel 7.0 GPU bug: with in-tree amdgpu on 7.0.0-27, KFD initializes and rocminfo works,
but any ROCm compute-queue creation faults: official torch rocm7.0 wheel (has gfx1151
kernels) segfaults at first alloc; TheRock gfx1151 nightly (2.11.0+rocm7.13) hangs; journal
shows
GCVM_L2_PROTECTION_FAULTfrom client CPF (command processor fetching unmapped queue memory). HSA_OVERRIDE_GFX_VERSION / HSA_XNACK / HSA_ENABLE_SDMA made no difference.amdgpu.user_queue=-1(auto) on 7.0 — suspected user-mode-queue path. Booting the installed 6.17.0-40 kernel did NOT fix it (same CPF fault) — kernel version was not the variable. - Also ruled out:
amd_iommu=force_isolationon the cmdline (removed, 6.17 pinned as permanent GRUB default in the same edit; backup at /etc/default/grub.bak) — no change. HSA_XNACK / HSA_ENABLE_SDMA / HSA_USE_SVM / GPU_MAX_HW_QUEUES — no change. BIOS UMA carve-out is 512MB ("auto") + 31GB GTT — that is the recommended Linux config, not a bug. - ROOT CAUSE: regressed MES scheduler firmware for gfx1151 in current linux-firmware (Ubuntu 26.04's 20260319 snapshot AND AMD's ROCm 7.2-era amdgpu-dkms-firmware both carry it). Known upstream: ROCm/ROCm#5724 ("MES 0x83" hang/memory-fault, open since Nov 2025); Arch forum confirms downgrade fixes the exact GCVM_L2_PROTECTION_FAULT/CPF signature. Fix applied: gc_11_5_1_mes1.bin + gc_11_5_1_mes_2.bin from linux-firmware tag 20251111 (git.kernel.org) into /lib/firmware/updates/amdgpu/ (wins the fw search path); originals backed up at /var/backups/mes-orig/. CAUTION: a linux-firmware or amdgpu-dkms-firmware package upgrade may overwrite these — if GPU faults return after an update, re-apply.
- GPU verified after fix: alloc OK, bf16 2048³ matmul 9.5 TFLOPS, fp64 OK, 32.8GB visible. Running stack: kernel 6.17.0-40 (pinned), in-tree amdgpu, torch 2.11.0+rocm7.13.0a20260424 (TheRock gfx1151 nightly — keep this exact wheel for reproducibility; stable rocm7.0 wheel untested post-fix).
2026-08-12 — Phase 0 complete: vendored, ported, smoke-tested #
Done today
- Cached all sources in
refs/(TRM paper + blog, HRM paper, Brette, stewardship essay); wrote the theory report (notes/hrm-brain-claims-and-embodied-critique.md) andPLAN.md. - Vendored
SamsungSAILMontreal/TinyRecursiveModelsat commitc011037(upstream archived) intovendor/TinyRecursiveModels/, kept pristine; provenance invendor/PROVENANCE.md. - Port:
scripts/port_from_vendor.pygeneratesdsem/from the vendor tree — rerunnable, every patch asserts an exact match count. Changes: imports underdsem.*(incl. the two runtime dynamic-import prefixes inutils/functions.pyand the evaluator prefix — these bit us once);torch.deviceselection cuda/rocm→mps→cpu (dsem/device.py, override withDSEM_DEVICE); pure-PyTorchAdamATan2transcribed from the upstream CUDA kernel (decoupled wd, lerp moments,atan2(m, sqrt(v̂)), no 4/π constant) and verified against an analytic first step; wandb replaced bydsem/run_logger.pyJSONL shim (real wandb viaWANDB=1);torch.compileonly on cuda; nccl→gloo fallback; no float64 in stablemax loss on MPS (fp32 there, fp64 elsewhere);pin_memorycuda-only. - Data: smoke dataset
data/sudoku-smoke(100 puzzles × 11 variants, test truncated to 256). Fulldata/sudoku-extreme-1k-aug-1000build kicked off (for the box). - Smoke run (MPS, fp32, MLP variant, T=3/n=6, batch 32, 20 epochs = 62 steps,
halt_max_steps=4): trained end-to-end, ACT halted sequences during training (q_halt_accuracy 1.0, one early halt observed), EMA eval swap worked, 2 eval passes over 256 puzzles ran the full ACT loop, checkpoint saved, exit 0. lm_loss 2.567→2.556 — consistent with 62 steps inside a 2000-step LR warmup; no correctness signal expected at this scale.
Gotchas recorded for the box (ROCm) setup
requirements.txtupstream is CUDA-specific (torch cu126 nightly, adam-atan2 ext, triton); our venv needs only: torch, numpy, einops, tqdm, coolname, pydantic, argdantic, omegaconf, hydra-core, huggingface_hub, pyyaml. numba only if running ARC evaluators.- Eval on the full test set is 423k puzzles ≈ 551 batches × up to 16 ACT steps — budget for
it (paper protocol evals every 5000 epochs; keep
eval_intervalhigh). - Upstream augmentation (
shuffle_sudoku) uses unseeded global numpy RNG — dataset builds are not bit-reproducible; build once and rsync everywhere. epochs % eval_interval == 0is asserted; total optimizer steps =epochs × total_groups / global_batch_size(one sampled variant per puzzle per epoch).
Next (needs the box): ROCm venv + DSEM_DEVICE/bf16 calibration run (steps/sec at paper
config, batch 768), then scaled-down validation, then the full reproduction run
(epochs=50000 eval_interval=5000, MLP variant, ema=True; target 87%±2 exact accuracy) and
the capacity-floor sweep (D=256/128/64) from PLAN Phase 1.