diff --git a/design-notes/frame-lanes-for-concurrent-matching.md b/design-notes/frame-lanes-for-concurrent-matching.md new file mode 100644 index 0000000..fdc4d3c --- /dev/null +++ b/design-notes/frame-lanes-for-concurrent-matching.md @@ -0,0 +1,497 @@ +# Frame Lanes: Concurrent Matching Within a Single Activation + +Design note extending `pe-redesign-frames-and-pipeline.md` with per-frame +matching lanes. Addresses the problem of multiple simultaneous pending operands +for the same dyadic instruction within a single activation — required for loops +and recursion. + +## Companion Documents + +- `pe-redesign-frames-and-pipeline.md` — base architecture (frames, pipeline, + approaches A/B/C) + +--- + +## Problem Statement + +The current frame model maps each `activation_id` to exactly one `frame_id`. +Within that frame, each dyadic instruction gets one matching slot indexed by +`offset % matchable_offsets`. This means at most one pending operand per +instruction per activation at any time. + +This fails for loops. Consider a counted loop where a dyadic ADD instruction +receives feedback from its own INC output: + +``` +iteration 1: L operand arrives at offset 3 → stored at presence[frame][3] +iteration 2: L operand arrives at offset 3 → collision! presence[frame][3] + is already set, so the hardware thinks this is a MATCH + (pairing two L operands from different iterations) +``` + +Without disambiguation, the second iteration's L operand is incorrectly paired +with the first iteration's L operand instead of waiting for its own R partner. + +The original design solved this with a 2-bit generation counter per context +slot, but that consumed token payload bits (6 bits total for ctx+gen) and +limited the instruction offset field. The frame redesign dropped the generation +field to widen offset to 8 bits and simplify the token format. + +We need a mechanism that provides generation-like disambiguation without +re-adding a token field. + +--- + +## Proposed Solution: Matching Lanes + +### Core Idea + +Split the tag store mapping from one-to-one into many-to-one. Multiple +`activation_id` values can map to the **same physical frame** but with different +**lane indices**. Lanes share the frame's constants and destinations (written +once at setup) but provide independent matching slots. + +``` +tag_store[act_id] → (frame_id, lane) + +Constants/dests: frames[frame_id][slot] — shared across all lanes +Match data: match_data[frame_id][offset][lane] — per-lane +Presence: presence[frame_id][offset][lane] — per-lane (1 bit) +Port: port_store[frame_id][offset][lane] — per-lane (1 bit) +``` + +### Frame Control Token Extensions + +The frame control token format (prefix `011+00`) has 3 spare bits in flit 1 +and a 16-bit payload in flit 2: + +``` +Frame control flit 1: [0][1][1][PE:2][00][op:1][act_id:3][spare:3] = 16 bits +Frame control flit 2: [payload:16] +``` + +The `op` field currently encodes ALLOC (0) and FREE (1). We split these into +four operations using 1 spare bit: + +``` +op spare[2] operation +── ──────── ───────── +0 0 ALLOC_NEW allocate fresh frame, assign lane 0 +0 1 ALLOC_SHARED share existing frame, assign next free lane +1 0 FREE_FRAME release lane AND return frame to free list +1 1 FREE_LANE release lane only, frame stays allocated +``` + +**ALLOC_SHARED** uses flit 2 to carry the parent activation_id whose frame +should be shared: + +``` +ALLOC_SHARED flit 2: [parent_act_id:3][spare:13] +``` + +The PE looks up `tag_store[parent_act_id] → (frame_id, _)`, picks the next +free lane from `lane_free[frame_id]`, and records +`tag_store[act_id] = (frame_id, lane)`. + +**FREE_LANE** removes the act_id → (frame, lane) mapping, clears that lane's +presence/port bits across all matchable offsets, and marks the lane as free. +The frame remains allocated (constants/dests preserved). + +**FREE_FRAME** does the same as FREE_LANE, then additionally returns the +frame to the free list. Should only be issued when all lanes for that frame +are free (or the PE can force-clear remaining lanes). + +### Lifecycle for a Loop + +``` +1. ALLOC_NEW(act_id=0) → frame 2, lane 0 +2. Setup: write constants/dests to frame 2 +3. Iteration 1 seed tokens use act_id=0 + +4. Before iteration 2: + ALLOC_SHARED(act_id=1, parent=0) → frame 2, lane 1 + Iteration 2 seed tokens use act_id=1 + +5. When iteration 1 completes: + FREE_LANE(act_id=0) → lane 0 freed, frame 2 stays + +6. Before iteration 3: + ALLOC_SHARED(act_id=2, parent=1) → frame 2, lane 0 (recycled) + Iteration 3 seed tokens use act_id=2 + +7. When all iterations done: + FREE_FRAME(act_id=last) → frame 2 returned to free list +``` + +Constants are written once in step 2. Each iteration gets its own matching +lanes via a different act_id sharing the same frame. + +### ABA Safety + +The ABA concern: a stale token from iteration 1 (act_id=0) arrives after +act_id=0 has been freed and re-allocated for iteration 3 (act_id=2). + +This is safe because: +- FREE_LANE removes act_id=0 from the tag store entirely +- When act_id=2 is allocated for iteration 3, it uses act_id=2 (not act_id=0) +- Any stale token with act_id=0 hits "act_id not in tag store" → rejected + +With 3-bit act_id (8 values) and at most 4 lanes per frame, there are 4 IDs +of ABA distance between allocation and re-use of the same act_id value. Given +that stale tokens drain within single-digit cycles, this is sufficient. + +--- + +## Hardware Impact by Approach + +### Lane Count + +L = number of lanes per frame. Practical values: 2 (1 bit) or 4 (2 bits). + +With L=4 and 4 frames: 16 possible (frame, lane) pairs. With 8 matchable +offsets: 128 match slots total. This provides 4 simultaneous pending operands +per instruction per frame — enough for most loop depths. Deeply nested +recursion beyond L would require frame splitting across PEs (the assembler +already supports this). + +### Approach C: 74LS670 Lookup (Recommended v0) + +**Tag store changes:** + +Currently: 2× 670, act_id → {valid:1, frame_id:2, spare:1} + +With lanes: 2× 670, act_id → {valid:1, frame_id:2, lane:1} for L=2 +(the spare bit becomes the lane index). For L=4, we need +{valid:1, frame_id:2, lane:2} = 5 bits, which exceeds one 670's 4-bit width. + +Options: +- **L=2 (1-bit lane):** fits in existing 670 layout. Zero additional chips for + tag store. The spare bit becomes the lane bit. +- **L=4 (2-bit lane):** need a third 670 to hold the extra lane bit (and + valid moves there too). +1 chip. + +**Presence/port metadata changes:** + +Currently: 4× 670 indexed by frame_id, each word holds presence+port for +2 offsets across 4 frames. Layout: + +``` +670 chip N (offsets 2N, 2N+1): + word[frame_id] = {pres_2N:1, port_2N:1, pres_2N+1:1, port_2N+1:1} +``` + +With lanes, the index becomes `[frame_id:2][lane]` instead of just +`[frame_id:2]`. The 670 has 4 words, so: + +- **L=2:** index is `[frame_id:2][lane:1]` = 3 bits, but the 670 only has + 2-bit addressing (4 words). We need to double the 670 count: 8× 670 for + presence+port, with lane as chip-select. **+4 chips.** + + Alternatively, re-pack: each 670 word holds presence+port for 1 offset + across 2 lanes: `{pres_L0:1, port_L0:1, pres_L1:1, port_L1:1}`. Then + we need 8 offsets × 1 chip each = 8× 670, indexed by frame_id (2 bits), + with offset selecting the chip. This is the same +4 chips but cleaner. + +- **L=4:** index is `[frame_id:2][lane:2]` = 4 bits. The 670 has 4 words + (2-bit address), so we'd need 4 670s per offset-pair, one per frame_id. + That's 4 × 4 = 16 670s for presence+port. Impractical. At L=4, the + presence/port metadata should move to SRAM or use a different register + approach. + +**Match operand data:** + +Currently in frame SRAM at `[1][frame_id:2][match_slot:3]` (match slots are +the low 8 offsets within the frame). With lanes, the address becomes +`[1][frame_id:2][match_slot:3][lane]`. + +- **L=2:** address is `[1][frame_id:2][match_slot:3][lane:1]` = 7 bits within + the frame region. 128 entries × 16 bits = 256 bytes. Well within SRAM + capacity. No additional chips. + +- **L=4:** address is `[1][frame_id:2][match_slot:3][lane:2]` = 8 bits. + 256 entries × 16 bits = 512 bytes. Still fits in SRAM. + +**Lane free tracking:** + +Per frame, track which lanes are free. For L=2: 1 flip-flop per frame × 4 +frames = 4 bits. For L=4: a 2-bit counter or 4-bit bitmask per frame = 8–16 +bits. Either fits in a single 74LS174 (hex D flip-flop) or similar. **+1 chip.** + +**Approach C summary (L=2):** + +| Component | Before | After | Delta | +|----------------------------|--------|--------|-------| +| act_id → (frame_id, lane) | 2× 670 | 2× 670 | 0 | +| Presence + port metadata | 4× 670 | 8× 670 | +4 | +| Bit select mux | 1–2 | 1–2 | 0 | +| Lane free tracking | 0 | 1 | +1 | +| Frame SRAM | 2 | 2 | 0 | +| **Total delta** | | | **+5 chips** | + +**Approach C summary (L=4):** + +Presence/port at L=4 exceeds practical 670 count. Two options: + +(a) Move presence/port to SRAM. Pack all 4 lanes' presence+port for one +(frame, offset) into a single 16-bit word: +`{pres0:1, port0:1, pres1:1, port1:1, ..., pres3:1, port3:1, spare:8}`. +Read in 1 SRAM cycle, same SRAM chip as frame data. Adds 1 cycle to matching +(read presence word before reading/writing match data). **+0 chips, +1 cycle.** + +(b) Use 74LS189 register files instead of 670s for presence/port. 189s are +16-word × 4-bit, addressed by `[frame_id:2][offset:2]` = 4 bits. Two 189s +(8 bits) hold presence+port for 4 lanes at one (frame, low-2-offset) combo. +With offset[2] as chip-select, that's 4× 189. **+4 chips (replacing 4× 670 +with 4× 189)**, net change depends on baseline. + +Option (a) is simpler and fits the v0 "minimal chips" philosophy. The extra +SRAM cycle for presence is identical to Approach A's tag read — it just +applies to the lane dimension instead. + +### Approach A: Set-Associative Tags in Frame SRAM + +Approach A already uses SRAM for tag storage. Lanes change the tag word format. + +Currently, each tag word packs 4-way set-associative entries: + +``` +{way0_valid:1, way0_act:3, way1_valid:1, way1_act:3, ...} = 16 bits +``` + +With lanes, the tag word already supports the concept — each way IS effectively +a lane. The act_id comparison finds the matching way, and the way index IS the +lane. The only change: ALLOC_SHARED must write the same frame region for +multiple act_ids. + +Actually, Approach A's set-associative structure already provides something +very close to lanes. The ways in the tag word serve the same purpose — multiple +act_ids can have pending operands at the same offset, disambiguated by act_id +comparison. The 4-way associativity gives 4 simultaneous pending matches per +offset across ALL activations. + +**Key difference:** in Approach A, the ways are shared across all activations +at that offset (global pool). The lane model gives per-frame isolation. Under +Approach A, if two different functions both have a pending operand at offset 3, +they consume 2 of the 4 ways. Under the lane model, each frame has its own L +lanes — no cross-activation contention. + +**Approach A with lanes:** the tag word becomes: + +``` +{way0_valid:1, way0_act:3, way0_lane:1, way1_valid:1, way1_act:3, way1_lane:1, ...} +``` + +This doesn't fit in 16 bits for 4 ways with L=2 (5 bits × 4 = 20 bits). Would +require wider tag words (32-bit SRAM or 2 reads per tag lookup), or reducing +to 2 ways. + +Alternatively, since act_id already implies frame_id (via the tag store), +Approach A doesn't benefit from lanes in the same way. The set-associative +structure already provides the disambiguation — adding lanes on top is +redundant. **Approach A doesn't need lanes; its ways serve the same purpose.** + +The real question for Approach A is: does 4-way associativity (global, shared) +provide enough concurrent matching depth? For loops: yes, as long as no more +than 4 iterations have pending operands at the same offset simultaneously. +For mixed workloads with multiple activations: depends on access patterns. + +### Approach B: Full Register-File Match Pool + +Original Approach B: 8-entry global pool with `{valid:1, act_id:3, offset:6, +port:1, data:16}` per entry, fully associative. + +**With lanes, the pool needs a lane field:** `{valid:1, act_id:3, offset:3, +lane:1, port:1, data:16}` for L=2. The comparator now matches on +`(act_id, offset, lane)` — but lane is derived from act_id via the tag store, +not carried in the token. So the comparator actually still matches on +`(act_id, offset)` as before. + +Wait — that's the key insight. Since the token carries act_id (not frame_id + +lane), and different iterations use different act_ids, the existing Approach B +pool already disambiguates correctly without any lane concept at all: + +- Iteration 1 (act_id=0): L operand stored as `{act_id=0, offset=3, ...}` +- Iteration 2 (act_id=1): L operand stored as `{act_id=1, offset=3, ...}` +- These don't match because act_id differs. + +**Approach B already handles concurrent matching across iterations, provided +each iteration uses a distinct act_id.** The only addition is the +ALLOC_SHARED/FREE_LANE mechanism to allow multiple act_ids to share one frame. +No changes to the match pool hardware at all. + +The constraint: the global pool has 8 entries total. With 4 iterations × 2 +pending operands each = 8 entries consumed. A tight but functional limit. + +**B+670 variants:** same analysis. The 670s resolve act_id → frame_id for +constant/dest access. The match pool (whether fully indexed or semi-CAM) uses +act_id directly and already disambiguates. **Zero additional match hardware +for lanes.** + +### Approach B+670 Indexed (Dedicated Register Slots) + +Currently: `[frame_id:2][offset:2:0]` = 5-bit address, 32 entries dedicated. +One entry per (frame, offset) pair. + +With ALLOC_SHARED: multiple act_ids map to the same frame_id, but they get +different lanes. The match data must be indexed by `[frame_id:2][offset:3] +[lane]` instead of just `[frame_id:2][offset:3]`. + +- **L=2:** 6-bit address, 64 entries. 8× 189 chips (up from 8). Actually, + the original B+670 indexed already uses 8× 189 for 32 entries of 16-bit + data. Doubling to 64 entries means 16× 189. That's a lot. Alternatively, + use SRAM for the doubled range: `[frame_id:2][offset:3][lane:1]` = 6 bits + within the match region. 64 entries × 16 bits = 128 bytes. Easily fits in + the shared SRAM chip. But then we lose the "zero SRAM cycles for matching" + advantage. + + Better option: keep register file, use the 670 presence bits to encode lane. + The 670 already stores `{presence, port}` per (frame, offset). With L=2, + expand to `{presence_L0, port_L0, presence_L1, port_L1}`. This is exactly + the same as the Approach C lane expansion above: 8× 670 for + presence+port. **The match data register file doubles, the presence 670s + double. +8 register chips, +4 670 chips = +12 chips.** Steep. + + For B+670 indexed, the more practical approach at L>1 is to fall back to + SRAM for match data and keep the 670s only for act_id resolution and + presence tracking. This effectively converts B+670 indexed into Approach C + with lanes — SRAM match data, 670 metadata. + +### B+670 Semi-CAM (Associative Within Frame) + +Currently: per-frame associative pool with W ways. Tag stores +`{valid:1, offset:3}` per way. Comparators search offset within frame. + +With lanes: each entry's tag becomes `{valid:1, offset:3, lane:1}` for L=2. +The comparator matches on `(offset, lane)` where lane comes from the 670 +lookup. **+1 bit per comparator.** For 3-bit offset + 1-bit lane = 4-bit +compare, each 74LS85 (4-bit comparator) handles one entry exactly. + +The pool's way count (W) determines how many simultaneous pending matches +per frame. With L=2 and W=4: 4 pending matches, shared across 2 lanes. +Each lane can use up to W entries (the pool is shared within the frame, +not partitioned per lane). This is actually better than strict per-lane +isolation — if lane 0 has 3 pending and lane 1 has 1, they use 4 entries +total without wasting any. + +**Semi-CAM hardware delta for L=2:** + +| Component | Before (W=2) | After (W=2, L=2) | Delta | +|-------------------|--------------|-------------------|-------| +| Tag registers | 2 chips | 2 chips | 0 | +| Comparators | 2 chips | 2 chips | 0 | +| Data registers | 4 chips | 4 chips | 0 | +| **Total** | | | **0** | + +The only change is the tag width grows by 1 bit (offset:3 → offset:3 + +lane:1 = 4 bits), which fits in the same comparator. **Zero additional chips +for the semi-CAM itself.** + +The ALLOC_SHARED / FREE_LANE logic: +1 chip (lane free tracking). +The 670 tag store: +0 chips (lane bit fits in spare bit). + +**This makes B+670 semi-CAM the most natural fit for lanes.** The +associative pool already handles variable-occupancy matching; adding a +lane bit to the tag is free in hardware. + +--- + +## Approach Comparison with Lanes + +| Property | A | C (L=2) | C (L=4) | B+670 semi W=2 | B+670 semi W=4 | +|---------------------------|--------------|-------------|--------------|----------------|----------------| +| Needs lanes at all? | no (ways) | yes | yes | yes | yes | +| Extra chips for lanes | 0 | +5 | +1 (SRAM) | +1 | +1 | +| Pending matches/frame | 4 (shared) | 2 per lane | 4 per lane | W (shared) | W (shared) | +| Extra SRAM cycles | 0 | 0 | +1 (pres) | 0 | 0 | +| Cross-activation contention | yes (global) | no | no | no | no | +| Implementation complexity | none | moderate | moderate | minimal | minimal | + +**Winner for lanes: B+670 semi-CAM.** Zero additional match hardware, lanes +come free via the existing associative tag. The 670 tag store absorbs the +lane bit in its spare capacity. Only cost is 1 chip for lane free tracking +and the ALLOC_SHARED/FREE_LANE control logic. + +**Runner-up: Approach C with L=2.** +5 chips (all 670s for doubled +presence/port). Simple, well-understood, but the 670 count is getting high +(10 670s per PE). + +--- + +## SRAM Address Map Update + +With L=2 lanes, the frame SRAM match region doubles: + +``` +v0 address space with lanes (L=2): + + IRAM region: [0][offset:8] instruction templates + capacity: 256 instructions (512 bytes) + + Frame region: [1][frame_id:2][slot:6] per-activation storage + capacity: 4 frames × 64 slots = 256 entries (512 bytes) + (constants, destinations, accumulators — shared across lanes) + + Match region: (Approach C / SRAM-based) + [1][1][frame_id:2][offset:3][lane:1] match operand data + capacity: 4 × 8 × 2 = 64 entries (128 bytes) + (carved from frame region address space, or separate region) +``` + +Total: 512 + 512 + 128 = 1152 bytes. Still well under 32Kx8 capacity. + +For B+670 semi-CAM: match data lives in register files, not SRAM. The SRAM +address map is unchanged (frame region only holds shared constants/dests). + +--- + +## Assembler Impact + +The assembler must: + +1. **Detect loops and recursion** that require concurrent matching. Static + analysis of feedback arcs in the dataflow graph. + +2. **Allocate activation IDs per iteration.** The loop prologue emits + ALLOC_SHARED for each new iteration's act_id before injecting seed tokens. + The loop epilogue emits FREE_LANE when an iteration completes. + +3. **Track lane depth.** If a loop's concurrency exceeds L (or W for + semi-CAM), the assembler must either: + - Insert synchronisation barriers (drain iteration N before starting N+L) + - Split the loop body across PEs to reduce per-PE concurrency + - Report a warning (analogous to matchable_offsets exceedance, AC5.8) + +4. **Generate setup tokens.** ALLOC_SHARED tokens carry the parent act_id + in their payload. The codegen pass already generates frame control tokens; + this extends the format. + +--- + +## Open Questions + +1. **L=2 vs L=4 for v0.** L=2 is cheaper (+5 670s for Approach C, +0 for + semi-CAM) and handles 2-deep loop pipelining. L=4 handles deeper nesting + but costs more in metadata storage. Recommendation: L=2 for v0, upgradable. + +2. **Loop iteration management.** Who manages the ALLOC_SHARED / FREE_LANE + sequence? Options: + - **Compiler-generated:** the assembler statically emits alloc/free tokens + as part of the loop control flow. Simple, but inflexible. + - **PE-internal:** a loop counter mechanism in the PE automatically + rotates lanes. More complex hardware, but simpler programs. + - **Hybrid:** compiler generates the control flow, PE provides the + lane allocation hardware. (Recommended.) + +3. **Semi-CAM way count vs lane count.** With B+670 semi-CAM, W (ways per + frame) and L (lanes per frame) interact. W=4 with L=2 gives 4 pending + matches shared across 2 lanes — 2 pending per lane on average, more if + one lane is quiet. Is W=2 sufficient? Depends on the number of dyadic + instructions with simultaneously pending operands. + +4. **Interaction with SC arc execution.** Strongly-connected arc blocks + execute sequential instructions within a single activation. Lanes are + orthogonal — SC arcs don't need concurrent matching (they're sequential). + But the frame_id latch for SC arcs must also latch the lane. Trivial + addition. + + diff --git a/docs/design-plans/2026-03-07-frame-lanes.md b/docs/design-plans/2026-03-07-frame-lanes.md new file mode 100644 index 0000000..732fcc3 --- /dev/null +++ b/docs/design-plans/2026-03-07-frame-lanes.md @@ -0,0 +1,270 @@ +# Frame Matching Lanes Design + +## Summary + +Extend the PE's frame-based matching to support multiple simultaneous pending +operands per instruction within a single activation. Multiple `activation_id` +values share one physical frame (constants/destinations) while maintaining +independent matching state per lane. Required for loop pipelining and recursion. +Changes span token types, PE internals, codegen, monitor, and tests. Assembler +macro expansion for automatic loop pipelining is out of scope. + +## Definition of Done + +The PE emulator supports matching lanes — multiple activation IDs sharing one +physical frame with independent match/presence/port storage per lane. FrameOp +gains ALLOC_SHARED and FREE_LANE. FREE_FRAME auto-detects last lane. +ALLOC_REMOTE is data-driven (frame constant flag for shared vs new). Existing +tests pass with the updated tag_store tuple API. New tests demonstrate +shared-frame matching, lane exhaustion rejection, smart free behaviour, and +data-driven ALLOC_REMOTE. Assembler macro expansion for automatic loop +pipelining is explicitly out of scope. + +## Acceptance Criteria + +### AC1: Tag Store Tuple API + +- **frame-lanes.AC1.1:** `tag_store` maps `act_id → (frame_id, lane)` where + `lane` is an `int` in range `[0, lane_count)`. +- **frame-lanes.AC1.2:** `PEConfig.initial_tag_store` type is + `dict[int, tuple[int, int]]`. PE constructor initialises tag_store from it. +- **frame-lanes.AC1.3:** `PEConfig.lane_count` field exists with default 4. + Controls third dimension of match arrays. +- **frame-lanes.AC1.4:** All existing tests pass with updated tuple API. + +### AC2: Separate Match Data Storage + +- **frame-lanes.AC2.1:** Match operand data lives in + `match_data[frame_id][offset][lane]`, separate from `frames[frame_id][slot]`. +- **frame-lanes.AC2.2:** `presence[frame_id][offset][lane]` is a 3D bool + array. `port_store[frame_id][offset][lane]` likewise. +- **frame-lanes.AC2.3:** `_match_frame()` uses `(frame_id, match_slot, lane)` + to read/write match data, presence, and port. +- **frame-lanes.AC2.4:** `frames[frame_id][slot]` remains shared across all + lanes. Constants and destinations are NOT per-lane. + +### AC3: FrameOp Extensions + +- **frame-lanes.AC3.1:** `FrameOp.ALLOC_SHARED` added. When received, + PE looks up `parent_act_id` (from payload), finds parent's `frame_id`, + assigns next free lane from that frame's lane pool, records + `tag_store[act_id] = (frame_id, lane)`. Clears only that lane's + presence/port bits. +- **frame-lanes.AC3.2:** `FrameOp.FREE_LANE` added. Removes tag_store entry, + clears that lane's presence/port/match_data across all matchable offsets. + Does NOT return frame to free list. +- **frame-lanes.AC3.3:** `FrameOp.FREE` (existing) becomes smart: removes + tag_store entry, clears lane data. If no other tag_store entries reference + the same frame_id, returns frame to free list and clears frame slots. If + other entries exist, behaves like FREE_LANE. +- **frame-lanes.AC3.4:** `FrameOp.ALLOC` (existing) unchanged — allocates + fresh frame, assigns lane 0. +- **frame-lanes.AC3.5:** `FrameAllocated` event gains `lane: int` field. + `FrameFreed` event gains `lane: int` and `frame_freed: bool` fields. +- **frame-lanes.AC3.6:** When all lanes for a frame are occupied and + ALLOC_SHARED is received, PE emits `TokenRejected` with reason + "no free lanes" and drops the token. + +### AC4: ALLOC_REMOTE Data-Driven + +- **frame-lanes.AC4.1:** ALLOC_REMOTE reads `fref+2` from frame. If value + is non-zero, emits `FrameControlToken` with `op=ALLOC_SHARED` and + `payload=parent_act_id`. If zero, emits `op=ALLOC` as before. +- **frame-lanes.AC4.2:** No new opcodes. Behaviour is entirely data-driven + from frame constants. + +### AC5: FREE_FRAME Instruction + +- **frame-lanes.AC5.1:** `FREE_FRAME` opcode uses the smart FREE behaviour + from AC3.3. Frees the executing token's activation lane; returns frame to + free list only if last lane. + +### AC6: Monitor and Snapshot Updates + +- **frame-lanes.AC6.1:** `PESnapshot.tag_store` type becomes + `dict[int, tuple[int, int]]`. +- **frame-lanes.AC6.2:** `PESnapshot` gains `match_data`, `lane_count` fields + reflecting the separated match storage. +- **frame-lanes.AC6.3:** Monitor REPL `pe` command displays lane info in + tag_store output. +- **frame-lanes.AC6.4:** Monitor graph JSON serialises lane info correctly. + +### AC7: Codegen Updates + +- **frame-lanes.AC7.1:** `codegen.py` generates `initial_tag_store` with + `(frame_id, lane)` tuples. Existing single-activation code uses lane 0. +- **frame-lanes.AC7.2:** No codegen changes needed for ALLOC_SHARED (manual + construction only for now). + +### AC8: Test Coverage + +- **frame-lanes.AC8.1:** Test: two act_ids sharing a frame via ALLOC_SHARED + have independent matching — L operand for act_id 0 does not interfere with + L operand for act_id 1 at the same offset. +- **frame-lanes.AC8.2:** Test: ALLOC_SHARED with all lanes occupied emits + TokenRejected. +- **frame-lanes.AC8.3:** Test: FREE on a shared frame frees only the lane; + other lanes' data is preserved. FREE on last lane frees the frame. +- **frame-lanes.AC8.4:** Test: ALLOC_REMOTE emits ALLOC_SHARED when + `fref+2` is non-zero. +- **frame-lanes.AC8.5:** Test: ALLOC_REMOTE emits ALLOC when `fref+2` is + zero (backwards compatible). +- **frame-lanes.AC8.6:** Test: full loop pipelining scenario — two + iterations of a dyadic instruction running concurrently on different + lanes, both producing correct results. + +## Architecture + +### Current Model + +``` +tag_store[act_id] → frame_id (1:1 mapping) +frames[frame_id][slot] (constants + dests + match data mixed) +presence[frame_id][offset] (1 pending operand per instruction) +port_store[frame_id][offset] (port of pending operand) +``` + +### New Model + +``` +tag_store[act_id] → (frame_id, lane) (many:1, multiple act_ids per frame) +frames[frame_id][slot] (constants + dests ONLY, shared) +match_data[frame_id][offset][lane] (per-lane operand storage) +presence[frame_id][offset][lane] (per-lane presence bits) +port_store[frame_id][offset][lane] (per-lane port metadata) +lane_free[frame_id] → set[int] (available lanes per frame) +``` + +### Frame Control Token Payload Convention + +``` +ALLOC: payload ignored (or return routing) +ALLOC_SHARED: payload = parent_act_id (low 3 bits) +FREE: payload ignored +FREE_LANE: payload ignored +``` + +### ALLOC_REMOTE Frame Slot Convention + +``` +fref+0: target_pe (int) +fref+1: target_act_id (int) +fref+2: parent_act_id (0 = ALLOC_NEW, non-zero = ALLOC_SHARED) +``` + +### Smart FREE Behaviour + +When FREE or FREE_FRAME executes for an act_id: +1. Look up `(frame_id, lane)` from tag_store +2. Remove tag_store entry for act_id +3. Clear match_data/presence/port_store for that lane across all offsets +4. Return lane to `lane_free[frame_id]` +5. Scan tag_store: does any other entry reference frame_id? + - No → return frame to free_frames, clear all frame slots + - Yes → frame stays allocated, constants/dests preserved + +### Lifecycle Example: Loop Pipelining + +``` +1. ALLOC(act_id=0) → frame 2, lane 0 +2. Setup: write constants/dests to frame 2 +3. Iteration 1 seeds use act_id=0 + +4. ALLOC_SHARED(act_id=1, parent=0) → frame 2, lane 1 +5. Iteration 2 seeds use act_id=1 + (act_id=0 and act_id=1 match independently at same offsets) + +6. Iteration 1 completes: FREE(act_id=0) → lane 0 freed, frame stays +7. ALLOC_SHARED(act_id=2, parent=1) → frame 2, lane 0 (recycled) +8. Iteration 3 seeds use act_id=2 + +9. All done: FREE(act_id=last) → last lane, frame returned +``` + +## Existing Patterns + +- **Frame control handling:** `_handle_frame_control()` in `emu/pe.py` already + dispatches on `FrameOp` enum values. Adding ALLOC_SHARED/FREE_LANE follows + the same pattern. +- **Token rejection:** `TokenRejected` event already exists and is emitted for + invalid act_ids. Lane exhaustion follows the same pattern. +- **Smart free (precedent):** The existing FREE already validates act_id + presence in tag_store before freeing. The smart-free extension adds a + scan step after removal. +- **Data-driven opcode behaviour:** ALLOC_REMOTE already reads frame slots + to determine target. Reading an additional slot for shared-vs-new is the + same pattern. +- **Codegen initial_tag_store:** Already generates `act_id → frame_id` + mappings. Extending to tuples is mechanical. + +## Implementation Phases + +### Phase 1: Foundation Types and Tag Store API (2 tasks) + +Update FrameOp enum, PEConfig, and tag_store type across the codebase. +All existing tests adapted to tuple API. No new behaviour yet. + +### Phase 2: Separated Match Storage (2 tasks) + +Extract match_data from frames into its own 3D array. Update _match_frame +to use lane dimension (always lane 0 for now). Update presence/port_store +to 3D. Verify matching still works identically. + +### Phase 3: ALLOC_SHARED, FREE_LANE, Smart FREE (3 tasks) + +Implement new FrameOp handlers. Add lane_free tracking. Implement smart +FREE behaviour. Add events with lane fields. Write tests for all new ops. + +### Phase 4: ALLOC_REMOTE Data-Driven and FREE_FRAME Update (2 tasks) + +Update ALLOC_REMOTE to read fref+2 for shared-vs-new. Update FREE_FRAME +opcode to use smart free. Write tests. + +### Phase 5: Monitor, Snapshot, and Codegen Updates (2 tasks) + +Update PESnapshot, capture(), REPL formatting, graph JSON. Update codegen +initial_tag_store to emit tuples. Verify monitor displays lane info. + +### Phase 6: Integration Tests (1 task) + +Full loop pipelining scenario test. Two concurrent iterations on shared +frame, both producing correct results. E2E verification. + +## Additional Considerations + +### ABA Safety + +With 3-bit act_id (8 values) and at most 4 lanes per frame, there are 4 IDs +of ABA distance between allocation and re-use of the same act_id value. +FREE removes the act_id from tag_store entirely, so stale tokens with freed +act_ids hit rejection. Re-allocation uses a different act_id value. + +### Hardware Mapping + +This design maps cleanly to Approach C (670 lookup) with L=2 at +5 chips, +or to B+670 semi-CAM with zero additional match hardware (lane bit fits in +existing comparator width). See `design-notes/frame-lanes-for-concurrent- +matching.md` for full hardware analysis. + +### Future Work + +- **Assembler loop macro:** `#loop_counted` and `#loop_while` could auto- + generate ALLOC_SHARED/FREE_LANE control flow with act_id rotation. +- **Lane depth analysis:** static analysis in the allocator to warn when + loop concurrency exceeds lane_count. +- **SC arc interaction:** frame_id latch for strongly-connected arc execution + must also latch the lane. Trivial addition when SC arcs are implemented. + +## Glossary + +- **Lane:** An independent matching slot within a shared frame. Multiple + act_ids can map to the same frame_id with different lane indices, + providing concurrent matching without duplicating constants/destinations. +- **Lane pool:** The set of available lanes per frame, tracked by + `lane_free[frame_id]`. Initially all lanes are free; ALLOC assigns lane 0, + ALLOC_SHARED assigns the next free lane. +- **Smart free:** FREE behaviour that auto-detects whether the freed lane + is the last one using that frame. If last, returns frame to free list. + If not, preserves frame for remaining lanes. +- **Parent act_id:** The activation ID whose frame should be shared during + ALLOC_SHARED. Used to look up the target frame_id.