A collaborative coding-agent orchestrator for atproto radl.app
README.md

Radial daemon #

radiald dispatches artifact requests into isolated containers. The container has open egress by design, but never receives Radial/atproto credentials: the local sidecar is its only protocol write path. Operators are responsible for blocking cloud metadata endpoints in their deployment environment.

Claims #

A request with an assignee dispatches straight to that agent. A request with no assignee is open to whoever claims it. The daemon writes a claim record, waits until it has ingested its own claim and seen it win the tie-break (earliest createdAt, then lowest URI — every reader computes this identically), and only then launches a turn, renewing the lease while it runs. If another operator's claim wins, the ledger row reads lost and any container already started is killed. That confirmation wait is what bounds duplicated work between two operators to one ingestion cycle; it costs an open request one cycle of latency and an assigned request nothing. run.claims has the timing knobs, radiald claim reset <request-uri> clears a row's local state (the claim generations already written are permanent records and are kept), and docs/operators.md covers what to change when.

Nothing releases a claim early: a lease cannot be shortened (readers would not adopt the rewrite) and Radial deletes no records, so "release" means "stop renewing" and a losing claim lingers until it lapses. This is harmless — a loser can never displace the winner in the tie-break — but it is why the default lease is short.

Auto-review uses this open-request path. The daemon chooses one eligible local review-capable agent identity to author the single review request, but omits assignee; any capable agent in the space, including one operated elsewhere, may claim and execute it. Authorship and execution are therefore separate, at the intended cost of the normal claim-ingestion confirmation cycle. Fold-ranked reviewers stagger fallback authorship. If simultaneous requests still land, a deterministic shared predicate keeps later agent-authored requests out of claim and dispatch selection; after mutual ingestion, the losing author retracts its request and abandons any local turn already in flight. Human-authored and assigned review requests remain deliberate and are never deduplicated.

Forges #

The forge is a registry, selected per project by the host of its gitUrl (run.forges). GitHub is observation-only: turns push and open their own pull requests with gh. Tangled is an atproto forge — turns push over ssh with a per-agent key the daemon mounts read-only at /run/radial-forge, and the daemon itself opens the pull request (an sh.tangled.repo.pull record), because tangled ships no user-facing CLI. Pull state is read from Bobbin, tangled's public read-only appview (https://api.tangled.org, configurable as api), with a direct PDS read behind it — as the whole answer when the appview is unavailable, and as corroboration whenever it reports open, which is also what its index says when it has not yet covered a merge. A project there is identified by its repository's own DID (not its owner's), read from the owner's sh.tangled.repo record. docs/radial-json.md has the configuration; docs/adr-tangled-forge.md has the reasoning and what remains unverified without a live knot.

A tangled pull is a patch, not a ref, and that one difference sets the rule the two forges do not share: a round is built with git format-patch, which OMITS merge commits and still exits 0 with an empty stderr. So the tangled implementation prompt tells an agent to rebase onto the base (GitHub's still merges, because nothing there is ever turned into a patch), and #formatPatch proves the round is whole — no merges in the range, one patch per commit — and refuses rather than publishing one that cannot apply. A round that lost a commit is not a visible failure downstream: the appview reports it as a conflict, against a branch that may merge as a fast-forward. --binary is passed for the same reason, so a file git considers binary — including a text file that has picked up a NUL byte — reaches the round with a body rather than as Binary files a/x and b/x differ.

GitHub credentials #

Implementation turns act directly as the operator on GitHub. radiald init reuses gh auth or starts gh auth login --hostname github.com --git-protocol https --web. A repository-scoped fine-grained PAT may instead be supplied as GH_TOKEN (legacy GITHUB_TOKEN is accepted). Tokens are never stored in Radial configuration or protocol records. Keep GH_TOKEN set for the daemon; it is injected only into implementation turns on GitHub projects and is also used for observation polling. A tangled turn holds its ssh push key and no GH_TOKEN at all.

Model credentials #

Each agent profile names the harness it runs on — claude (Claude Code), pi (the pi coding agent, which supports 15+ providers), or codex (the Codex CLI) — and each harness declares the provider environment variables it understands. For claude that is a dedicated, spend-capped ANTHROPIC_API_KEY or a CLAUDE_CODE_OAUTH_TOKEN (minted from a Claude subscription with claude setup-token — see docs/running-an-agent.md for the trade-offs); for pi it is whichever variable the chosen provider documents (OPENAI_API_KEY, GEMINI_API_KEY, OPENROUTER_API_KEY, …); for codex it is CODEX_API_KEY, a Business/Enterprise CODEX_ACCESS_TOKEN, or the run.codexAuth.mode = "chatgpt-session" flow initialized by radiald codex login (not OPENAI_API_KEY, which current codex releases no longer accept as authentication for codex exec). All are extensible with run.modelEnv. The daemon forwards a variable only when it is set in its own environment, prints the names (never the values) at startup, and refuses to start when a loaded profile's harness has no credential at all. Forwarded credentials reach turn containers only; check containers are deliberately secret-free (env: {}) and may use normal network access for dependency installation.

An unknown harness is refused, at radiald init and again at radiald run startup — never silently downgraded to another one. See docs/radial-json.md.

Direct PR workflow #

An implementation turn configures gh git authentication, works on its reserved radial/impl-* branch, commits with the agent/DID attribution trailer, pushes, and creates or updates its PR. Its PR body includes the deterministic Radial artifact URI. It then calls:

radial artifact submit --title "<short title>" --branch "$RADIAL_BRANCH" --commit "$(git rev-parse HEAD)" --pr "$(gh pr view --json url -q .url)" --body-file summary.md

The daemon synchronously verifies a lowercase full SHA and a canonical same-repository PR URL, then writes the typed artifact record. Turns must not merge. Forge reconnaissance is intentionally allowed: gh pr view, gh api, and arbitrary web fetches work in a turn.

Networking #

With run.network omitted Docker uses its ordinary bridge. A custom network is supported, but host is rejected. This is a credential boundary, not an egress allowlist: protect the daemon socket and cloud metadata service at the operator layer.

Turn transport defaults to auto: a bind-mounted Unix socket on Linux, and TCP on macOS/Windows where VM-backed Docker cannot connect to a bind-mounted host Unix socket. The TCP path binds an ephemeral, token-gated server on all host interfaces so Docker can reach it; firewall it, never expose it to a LAN, and do not use it for production deployment. Explicit unix and tcp values override platform detection.

Turn types #

Three kinds of turn run in the same container shape and differ only in their terminal record:

request type terminal record checkout forge token
a registry type (plan, implementation, adr, …) artifact read-only, except implementation implementation only
review review verdict on the named subject read-only, pinned to the subject's commit when it links one yes (PR reconnaissance)
answer message — a reply in the goal's thread read-only, project default branch no

The daemon stamps terminal records and blocked-turn questions with the effective model at the instant each write begins. A concrete model reported by streamed harness output wins over the configured invocation model; an unknown harness default is omitted. This provenance is daemon-owned and is not accepted from the container-side RPC.

review and answer are built-ins of this layer, not registry data: type create refuses both names, and neither ever becomes a row in a goal's unit list.

An answer turn costs a full container and a shallow clone for what is often one paragraph. That is accepted for v1 — it is what makes a reply attributable and sandboxed like any other turn — and a checkout-less answer turn is the obvious optimization once the feature sees real use. Answer turns share run.concurrency with everything else; they are logged distinctly (turn <uri> (answer)), and if they start starving implementation turns the fix is a separate cap rather than a priority scheme.

Every turn's bundle.md carries a ## Community comments section when the goal has blessings on it — comments from people who are not members of the space, which an active member vouched for with a blessComment record. The daemon does nothing special for these: it ingests blessings through ordinary member polling like every other record, and the section is built from bundle.guestComments, which buildTurnBundle() fills from GoalView.blessedComments and from nothing else. An unblessed guest comment is not in any store the daemon holds — nothing polls a non-member's repo — so it cannot reach a bundle by any path. The section is separate from the thread and names both DIDs, so a turn can tell which lines came from inside the space; the actual bound on what an injected one can do is unchanged and write-side (design §13).

The daemon never authors an answer request. Auto-review is the one daemon-authored hop there is; every answer request is written by a human, which is what keeps a reply from commissioning the next reply (design §10, and §14's rejected alternative C). The property is asserted over the source — test/auto-review.test.mjs, "the daemon writes no other request": exactly one create(COLLECTIONS.artifactRequest) exists under src/, and it builds an open review.

Private spaces #

A private space carries its records in signed envelopes instead of writing them to a PDS (design §18). Nothing in dispatch, claims, ledgers, checks or auto-review is aware of it: private mode is two substitutions and no branches above them.

  • PrivateSpaceRuntime (private-space.ts) replaces SpaceSyncRuntime behind the SpaceRuntime interface runDaemon holds: envelopes off a PrivateBus, the device directory off members' public repos (that poller lists the three directory collections and nothing else), and the same unchanged materialize(). Its blobs.db holds the bytes records name — too large for an envelope, pulled from peers by the CID they hash to, and what is still missing is derived from the corpus rather than tracked (design §18, ADR §15).

  • privateActorRegistry replaces each actor's client with one whose three Radial write methods seal envelopes. createForeign/putForeign/uploadBlob stay on the PDS deliberately: a sh.tangled.repo.pull record is a public forge record, and a space being private says nothing about where its project's code review happens. An actor with no device key on this machine is dropped from the registry rather than left holding a PDS client.

  • reconcileCapabilities (private-space.ts) is the one thing the daemon must do here that it never does in a public space: publish its own profiles' agent records into the replica. A private fold reads agent only out of an envelope — the device gate, and correctly, since a capability anybody with repo custody could mint is not one the space granted — so radiald init's PDS write is invisible here and index.agents would stay empty forever. It writes the effective types for this space (actorTypesFor) resolved into artifactTypes with no scopes field, keyed by the profile name like every other agent record, and only when the record differs: the comparison sorts and ignores createdAt, because an unstable one would seal an envelope every tick that every peer would then keep. openPrivateSpaces runs it before the first fold (so the coverage advisory never fires on a space the daemon is about to cover itself) and refreshCapabilities() retries it on the address-refresh interval, for the replica that had not yet read this daemon's own device record back and so refused its own first envelope.

packages/daemon/test/private-dispatch.test.mjs runs the whole path over MemoryPrivateBus: two daemons catch up off the bus, one wins the claim by the existing deterministic tie-break, its turn's artifact is sealed into an envelope the other folds, and every replica's indexDigest() agrees — with a PDS client that throws on every method, so nothing can have reached one.

Keys live in <data directory>/devices.json (0600), one per identity:

radiald device list                        # what this machine holds, and each key's state
radiald device publish                     # every identity's public `device` record
radiald device rotate --did did:plc:…      # new key, old one retired — the fold does not notice
radiald device retire <deviceKeyId>        # withdraw its address: nobody dials or serves it again

The endpoint (private-transport.ts, private-run.ts) #

radiald run serves a private space over a real connection (docs/adr-private-mode-iroh.md §20).

  • One endpoint for the daemon, not one per space. PrivateEndpoint binds it through @radial/transport-iroh — imported dynamically, so a daemon with no private space never loads a native module — and routes each inbound frame to the attached space its topic names. A deviceAddress is a fact about a device and names one endpoint id; an endpoint per space would make two spaces rewrite one record forever.
  • Its identity persists in <data directory>/private/endpoint.key (0600, atomic rename).
  • The address is published at startup, per identity, under the device key's own rkey — and only when the endpoint or its relays changed, so an ordinary restart writes nothing to any PDS.
  • …and kept current while it runs (ADR §21). Every run.privateAddressRefreshMs (60 s default) openPrivateSpaces().refreshAddresses() compares where the endpoint is reachable now against what this process last published and republishes only on a change, so a machine that changed network is dialable again without a restart and one that did not writes nothing. The relay list is sorted before comparison: its order is the transport's and means nothing.
  • Gossip wakes the run loop (ADR §21). A private space has no PDS to poll, so onArrival is the private half of Jetstream's onRecord — it interrupts the loop's sleep, which is what makes PrivateSpaceRuntime.streaming a true statement when run.jetstream uses it to drop the tick rate to the backfill interval. A notification, never a delivery: the envelope is admitted by the cycle either way.
  • Inbound connections are authorized before they are adopted, per attached space, against the same directory accept() authorizes each frame against. An unnamed endpoint is closed immediately; a fresh replica learns routes only from its transfer ticket (ADR §31).
  • Startup order is load-bearing: device key → endpoint → address → replica → ticket check. Each is the next one's precondition; private-run.ts is that sequence and says why at each step.

The replica lives in <data directory>/private/<founder-did>~<rkey>/ — the same directory the radial CLI uses, with the same space.json (the layout is in @radial/core/node so the two cannot disagree). The daemon must be completely stopped before any CLI access to that replica, not merely between syncs. For a CLI grant or ticket mint, use the same data directory: start radiald long enough to publish its address and catch up, stop it completely, run the CLI command, then restart radiald before the recipient syncs. The CLI and daemon must never overlap on the replica.

packages/daemon/test/private-transport.test.mjs drives all of it — two daemons, shared public repos, and a loopback transport carrying real encoded frames — through the same startup path production uses.

Troubleshooting a native startup crash #

The one found so far is fixed. radiald run died intermittently with SIGSEGV, SIGBUS, or a SIGABRT out of the allocator, after private transport: binding native iroh endpoint and before the line that reports the bound endpoint — and started cleanly on a rerun. @number0/iroh@1.1.0 borrows the RelayMode argument of its async Endpoint.bind and keeps using it after the call returns, so the argument was collectable while the bind was still reading it; packages/transport-iroh now pins every native object it hands to a native async call (transport-iroh's README and ADR §32). A build from before that fix crashed 9 starts in 12 on Linux/arm64 under node --test packages/transport-iroh/test/argument-lifetime.test.mjs, and 0 in 12 after. There is no version to upgrade to: 1.1.0 is the current release, and the defect is in how the binding lends the argument, not in a fixable pin.

Anything else that dies this way is a new fault, and the procedure below is how to characterise it.

pnpm's ERR_PNPM_RECURSIVE_EXEC_FIRST_FAIL message only reports that radiald died. If it names SIGSEGV, collect the last private transport: startup phase and a native backtrace: on Linux use coredumpctl info and coredumpctl gdb (or the platform equivalent). Core dumps may contain credentials and private records; keep the dump secret and share only a redacted, symbolized backtrace. Record the Radial commit, Node and pnpm versions, OS, architecture, libc, container use, whether run.privateSpaces is non-empty, and the relay mode. Redact DIDs, paths, keys, and tokens. Read the last phase as a lower bound: under pnpm the daemon's stdout is a pipe, Node writes to a pipe asynchronously, and a crash drops whatever it had not flushed. The phase you see is one the process reached, not provably the last one.

The isolated probe exercises import, bind, online/address discovery, and close in fresh child processes. A crash in one iteration is reported by signal without losing the remaining results, and a start that neither exits nor dies is killed and reported too (--start-timeout-ms, default 120000):

pnpm --filter @radial/transport-iroh probe -- --iterations 100 --relays disabled
pnpm --filter @radial/transport-iroh probe -- --iterations 100 --relays default
pnpm --filter @radial/transport-iroh probe -- --iterations 100 --relays disabled \
  --key-file /path/to/copied/endpoint.key --timeout-ms 10000

Run both a fresh identity and a copy of the affected installation's <dataDir>/private/endpoint.key. Never delete, rotate, or probe with the live key or data directory: that changes the published identity and confounds the comparison. Use a copied temporary config and data directory for destructive experiments. Compare enough fresh starts to exceed the observed failure rate, first with relays disabled and then with default relays. An external supervisor may restart a crashed daemon as a temporary mitigation, but JavaScript cannot catch or safely retry a native segmentation fault.