# CCU and Inter-Cluster Routing (Design Sketch) Exploratory design document for scaling beyond a single cluster (8 PEs + 8 SMs). Not required for v0/v0.5. Documents the architectural path from single-cluster flat addressing to multi-cluster hierarchical routing, without requiring changes to the token format, PE design, or bus protocol. Companion to `bus-interconnect-design.md` (intra-cluster bus) and `pe-redesign-frames-and-pipeline.md` (PE architecture and token format). --- ## Context: Single-Cluster Scaling Limits With split AN/CN/DN buses and 3-bit node addressing (bit[15] repurposed as address high bit on each split bus), a single cluster supports up to 8 PEs and 8 SMs with flat addressing. Beyond 8 PEs per bus, contention degrades throughput and physical bus length creates signal integrity concerns. The next step is clustering: multiple independent bus segments, each serving a subset of nodes, connected by an inter-cluster network. --- ## Cluster Architecture A cluster is a self-contained bus domain: ``` cluster = 7 PEs + 1 CCU + up to 8 SMs sharing local CN, AN, DN buses ``` The CCU (Cluster Control Unit) occupies one PE address on the local bus (e.g., PE_id = 7). It appears as a normal bus node from the perspective of other PEs and SMs in the cluster. Internally, it bridges the local bus domain to the inter-cluster network. ``` ┌─────────────────────────────────────────────┐ │ Cluster 0 │ │ │ │ PE0 PE1 PE2 PE3 PE4 PE5 PE6 [CCU] │ │ │ │ │ │ │ │ │ │ │ │ ═╪════╪════╪════╪════╪════╪════╪══════╪═ │ │ │ CN bus (local) │ │ │ │ ═╪════╪════╪════╪════╪════╪════╪══════╪═ │ │ │ │ │ SM0 SM1 SM2 SM3 SM4 SM5 SM6 SM7 │ │ │ │ │ └─────────────────────────────────────────┼────┘ │ inter-cluster link │ ┌─────────────────────────────────────────┼────┐ │ Cluster 1 │ │ │ │ │ │ PE0 PE1 PE2 PE3 PE4 PE5 PE6 [CCU] │ │ ... │ └──────────────────────────────────────────────┘ ``` 7 PEs per cluster is the usable compute capacity. the 8th PE address is consumed by the CCU. with 4 clusters, that's 28 PEs + 4 CCUs. with 16 clusters, 112 PEs + 16 CCUs. --- ## The CCU Is a PE The CCU is architecturally a standard PE with two differences: 1. **an additional output/input port** connected to the inter-cluster link (instead of, or in addition to, the standard local bus ports). 2. **specialised IRAM contents** loaded at boot: instructions for matching routing tokens with data tokens, constructing inter-cluster packets, and injecting received packets onto the local bus. The CCU has the same pipeline, the same frame model, the same matching mechanism, the same ALU, and the same bus interface as any other PE. It runs dataflow instructions. The compiler generates the CCU's code alongside the regular PE code. This means the CCU is **programmable**. different programs can load different CCU behaviour: - **simple forwarding**: match routing + data, construct inter-cluster packet, emit. minimal IRAM, no computation. - **load balancing**: inspect token metadata, rewrite destination based on cluster load metrics, forward. requires a few extra instructions and frame slots for the load table. - **traffic shaping**: buffer and reorder inter-cluster traffic to reduce link contention. uses frame storage as a small reorder buffer. - **aggregation**: combine multiple small tokens into a bulk transfer across the inter-cluster link. reduces per-token link overhead. - **monitoring**: count inter-cluster traffic by destination, report utilisation metrics. diagnostic mode. If none of this is needed, the CCU can be stripped to the minimum components required for routing — just enough pipeline to receive a token, do a table lookup, and emit a reformatted packet. At that point it's essentially a hardware router that happens to use the PE instruction set as its microcode format. --- ## Inter-Cluster Token Flow Inter-cluster communication requires two tokens at the source PE, because a single 2-flit token carries 16 bits of data but the CCU needs both the data (16 bits) AND the final destination routing (16 bits) = 32 bits total. ### Sending (Source Cluster) The sending PE emits two tokens to the CCU (local PE_id = 7): ``` token 1 (routing info): flit 1: [PE=7, offset=CCU_ROUTE, act_id=A] flit 2: final_dest_flit1 (pre-formed flit 1 for the destination cluster's local bus: destination PE_id, offset, act_id — all relative to the destination cluster's address space) token 2 (data): flit 1: [PE=7, offset=CCU_DATA, act_id=A] flit 2: actual_data (the value being forwarded) ``` Both tokens carry the same `act_id`, which the CCU uses for dyadic matching — token 1 and token 2 are the two operands of a dyadic instruction in the CCU's IRAM. When both arrive, the CCU has the complete information to construct the inter-cluster packet. The sending PE's instruction for this uses mode 2 (fanout) or two separate single-dest instructions. The frame holds the gateway destination (`[PE=7, offset=CCU_ROUTE, ...]`) and the final destination as a constant. The compiler generates this automatically for any cross-cluster edge in the dataflow graph. ### CCU Processing The CCU's IRAM contains a dyadic instruction at `CCU_ROUTE`/`CCU_DATA` offset pair. When both tokens match: ``` left operand = final_dest_flit1 (from token 1) right operand = actual_data (from token 2) CCU instruction: FORWARD (a specialised opcode, or a generic PASS with CHANGE_TAG mode that uses the left operand as the output flit 1) ``` The CCU emits the inter-cluster packet on its inter-cluster output port: ``` flit 0: [dest_cluster:4][src_cluster:4][flags:8] inter-cluster header flit 1: final_dest_flit1 (from left operand) destination-local routing flit 2: actual_data (from right operand) payload ``` Flit 0 is the inter-cluster routing header. The CCU derives `dest_cluster` from the final_dest_flit1 (the destination PE_id maps to a cluster via a small lookup table in the CCU's frame or a dedicated routing register). ### Receiving (Destination Cluster) The destination cluster's CCU receives the 3-flit inter-cluster packet on its inter-cluster input port. It strips flit 0 (consumed for routing) and injects flit 1 + flit 2 onto the local bus as a normal 2-flit token. The destination PE sees a standard compute token arrive — it has no idea the token originated in another cluster. The flit 1 routing fields already contain the correct local PE_id, offset, and act_id for the destination PE. The compiler ensured this at compile time. ### Overhead - **bus bandwidth**: 2 tokens on the source cluster's CN bus (vs 1 for intra-cluster). the routing token is pure overhead. - **latency**: CCU matching + inter-cluster link traversal + destination CCU injection. roughly 10-20 cycles depending on link speed and contention, vs 2 cycles for intra-cluster. - **CCU pipeline occupancy**: the CCU processes one inter-cluster transfer per dyadic match cycle. throughput limited by CCU pipeline speed — same as any PE. for high cross-cluster traffic, this is the bottleneck. The compiler's placement algorithm has a strong incentive to keep communicating instructions within the same cluster. Inter-cluster edges are the minority case in a well-partitioned program. --- ## CCU Routing Table The CCU needs to map destination global addresses to cluster IDs. For small-scale systems (4-16 clusters), this is a trivial lookup: ``` routing table: global_PE_id → dest_cluster_id 32 entries × 4-bit cluster_id = 16 bytes fits in a single 74LS670 pair or a handful of frame slots ``` The table is loaded during bootstrap, same as any other frame data. It can be updated at runtime (via PE-local write tokens) to support dynamic remapping or migration. For larger systems, the routing table can be hierarchical: first level maps to a cluster group, second level (at the cluster group's gateway) maps to the specific cluster within the group. This is standard prefix- based hierarchical routing — each level consumes some address bits and forwards based on the prefix. --- ## Inter-Cluster Link Options The inter-cluster link is the physical connection between CCUs. Its protocol is independent of the intra-cluster bus — the CCU translates between them. Several options: **Point-to-point links.** direct connection between each pair of CCUs. 4 clusters = 6 bidirectional links (full mesh). each link is a 16-bit bus + control signals. maximum bandwidth but wiring grows as N×(N-1)/2. practical up to ~4-8 clusters. **Ring.** CCUs connected in a ring. each CCU has a "left" and "right" link. packets forwarded around the ring until they reach the destination CCU. minimal wiring (N links for N clusters). latency grows with ring size. the ring uses the same flit-based protocol as the intra-cluster bus, just with inter-cluster headers. **Second-level shared bus.** all CCUs share an inter-cluster bus with its own arbiter. simplest to build, same protocol as intra-cluster. contention at large scale, but fine for 4-8 clusters. **Hierarchical.** groups of clusters share a bus at one level, groups of groups at the next. prefix-based routing at each level. this is the eventual scaling path if the system grows to 16+ clusters. For early multi-cluster experiments (2-4 clusters), point-to-point or a second-level shared bus is simplest. the CCU's inter-cluster port is just another set of 373 latches driving/receiving on whichever link topology is wired up. --- ## Interaction with SM Traffic SMs are cluster-local. an SM in cluster 0 is only reachable from cluster 0's PEs (via the local AN bus). cross-cluster SM access goes through the CCU: ``` PE 3 (cluster 0) wants to read SM 2 (cluster 1): 1. PE 3 sends routing + data tokens to CCU (local PE 7) 2. CCU forwards to cluster 1's CCU via inter-cluster link 3. cluster 1's CCU receives, injects as an SM token on local AN bus 4. SM 2 processes the read, emits response on local DN bus 5. cluster 1's CCU captures the response (it's addressed to a non-local PE — the return routing flit 1 carries cluster 0's PE 3 as the destination, which doesn't exist locally) 6. cluster 1's CCU forwards response to cluster 0's CCU 7. cluster 0's CCU injects response onto local DN bus → PE 3 receives ``` this is a 4-hop round trip (PE→CCU→link→CCU→SM→CCU→link→CCU→PE). slow, but correct. the compiler should place data structures on SMs in the same cluster as the PEs that access them most frequently. for the response path (steps 5-7), the destination CCU needs to recognise "this DN token is addressed to a non-local PE" and forward it. the simplest mechanism: the CCU monitors the DN bus and captures any token whose destination PE_id doesn't match any local PE. this requires the CCU to know which PE_ids are local — a 7-bit "local PE mask" register. cost: one comparator + mask check. alternatively, the SM's return routing could explicitly target the CCU (PE_id = 7) with a second-level routing token. the originating PE packs the return routing to go "SM → local CCU → inter-cluster → source CCU → source PE." this is set up at compile time in the frame constants. more complex routing setup, but no CCU bus snooping needed. --- ## What Changes at the PE Level: Nothing The PE architecture, token format, instruction encoding, frame model, and bus interface are completely unchanged for multi-cluster operation. The only difference a PE sees is that some of its destinations route to PE 7 (the CCU) instead of directly to another PE or SM. The compiler handles this transparently — the PE doesn't know or care whether its destination is local or remote. The CCU IS a PE. it uses the same board, the same chips, the same IRAM format. it just has different code loaded and a different output connector. --- ## Build Progression ``` phase PEs SMs clusters inter-cluster notes ────── ──── ──── ──────── ────────────── ────────────────── v0 1-2 1 1 none single shared bus v0.5 4 4 1 none shared or split bus v1 7+1 8 1 none split bus, CCU slot reserved v2 14+2 16 2 point-to-point first inter-cluster link v3 28+4 32 4 mesh or ring multi-cluster v4+ 56+8 64 8+ hierarchical prefix routing between CCU groups ``` The CCU slot (PE 7) can be reserved from v1 onward even if no inter-cluster link exists yet. the address is simply unused. when the second cluster is added, a PE board with inter-cluster firmware loads into the PE 7 slot. --- ## Open Questions 1. **CCU instruction set.** does the CCU need any opcodes beyond what a normal PE has? FORWARD might be a CHANGE_TAG mode 4 instruction with the inter-cluster port as the output. if so, zero new opcodes — just a different output latch wired to the inter-cluster link instead of (or in addition to) the CN bus. 2. **Inter-cluster link width.** 16-bit (same as intra-cluster) is simplest. narrower (8-bit serialised) saves wiring at the cost of latency. wider (32-bit) reduces per-token link cycles but doubles the connector. 16-bit is the natural starting point. 3. **CCU FIFO depth.** the CCU processes inter-cluster traffic at PE pipeline speed. if cross-cluster traffic bursts exceed the CCU's processing rate, packets queue. the CCU's input FIFO (or bus reg + BUSY) is the bottleneck. a deeper FIFO (NOS IDT7201 or SRAM-based) at the CCU's inter-cluster input would absorb bursts. 4. **Multiple inter-cluster links per CCU.** for mesh topologies (4+ clusters), each CCU needs links to multiple neighbours. the CCU would need multiple inter-cluster output ports (one per neighbour), selected by the routing table. each port is a 373 pair + control. or: a single link to an inter-cluster switch/router (separate from the CCU). 5. **SM placement across clusters.** should every cluster have its own SMs, or should some clusters be compute-only (PEs + CCU) with SMs concentrated in dedicated memory clusters? the architecture supports both — it's a compiler placement decision.