# Remote Development Server — Build Spec **Status:** Reviewed and audited; all open questions resolved. Ready for handoff. **Audience:** An agent executing the build, with a human available to approve interactive steps. > **Amending this spec? Read §13 first.** It encodes seven review techniques, two checks that > only work once an implementation exists, and eighteen standing rules — each derived from a defect > that actually occurred here. Several of those defects were introduced *by fixes to earlier > ones*, so amendments carry the same risk as the original. --- ## 1. Objective Stand up a single remote Linux box that serves as the operator's primary development environment. All work happens on this server. The local machine is a thin client running a terminal. ### Success criteria 1. Operator can connect from a laptop and resume the exact session state from the previous connection, including running processes. 2. Session survives laptop sleep, network changes, and disconnection. 3. No *fetchable* secret is stored at rest on the server. The five items in §7.3 are the accepted, documented exception — do not design around eliminating them. 4. The machine can be recreated from backup, verified by an end-to-end drill, not by assertion. 5. Typing in `vim` is comfortable in-region and tolerable long-haul — *tolerable* deliberately, since mosh's predictive echo was removed (§10) and the round trip is felt at that distance. 6. JavaScript/TypeScript and Rust toolchains are first-class; adding a third language later is a one-line change, not a redesign. 7. Provisioning is scripted and idempotent, and the scripts are proven to reproduce the box by periodic rebuild from a stock image — not assumed to. 8. No script destroys or overwrites an existing configuration file, on either machine. 9. Every scheduled job is observable by absence, not only by failure — a timer that stops firing raises an alert, delivered by push and email. ### Non-goals - Multi-user access. Single operator. - Production hosting. This box runs no public service. - Unattended or headless agent execution. Explicitly out of scope — see §9. - Mobile targets (iOS, Android). Out of scope — see §10. - High availability. Recovery is measured in tens of minutes, not seconds. --- ## 2. Design constraints What drove the decisions below. Preserve these if revising. | Factor | Value | | --- | --- | | Latency profile | Same region as the server most of the time; occasional long-haul sessions at several times the round trip | | Editor | `vim` (not neovim) | | Languages | JavaScript/TypeScript and Rust; polyglot-ready | | Package manager | `pnpm`, used heavily, TypeScript monorepos | | Git forge | Tangled, managed knot on `tangled.org` | | Secrets manager | Proton Pass | | Rebuild frequency | Rare | The latency profile is the one that constrains §4.1 and §10: the region is chosen for the common case, and the editor must stay usable — not merely functional — on the long-haul case. --- ## 3. Architecture Four layers with distinct lifecycles. Do not merge them. | Layer | Contents | Lifecycle | | --- | --- | --- | | **Host OS** | Packages, dotfiles, toolchains | Scripted, drift-checked monthly, snapshotted | | **Durable state** | Repos, uncommitted work, pnpm store, `~/.claude`, shell history | Backed up daily, never rebuilt | | **Secrets** | Vault contents | Fetched on demand, never at rest | | **Residue** | Five items that cannot be fetched on demand | §7.3 | The load-bearing consequence: everything fetchable is fetched, so a restored host carries only the §7.3 residue and is re-supplied by authenticating. Whole-disk Hetzner backups *do* capture that residue; §7.3 explains why that is acceptable rather than something to engineer away. --- ## 4. Infrastructure ### 4.1 Server - **Provider:** Hetzner Cloud - **Instance:** CX53 — 16 vCPU, 32 GB RAM, 320 GB NVMe (~€29.49/month; verify in console, pricing moved twice in 2026) - **Region:** Germany over Helsinki for Asia routing. Falkenstein (`fsn1`) is the default; Hetzner also runs cx53 capacity out of a second German site within a couple of milliseconds of it, so when the default has no capacity the choice falls back to whichever currently can place one, rather than to latency. See §12.1 on `HCLOUD_LOCATION`. - **OS:** Ubuntu 24.04 LTS - **Timezone: `Europe/Berlin`**, set explicitly via `timedatectl`. Not cosmetic: every systemd timer and every healthchecks `tz` field (§6.3) must agree with it. A host left on UTC against Berlin-scheduled checks is off by an hour, and off *variably* across DST — producing DOWN alerts twice a year that look like real failures. `verify.sh` asserts it. Do not downsize. The RAM sizing is deliberate on three counts: `tsserver` on a monorepo occupies 2–6 GB alone, Rust link steps are memory-hungry, and the operator runs **several project sessions concurrently** (§4.3), each potentially holding its own language server and watchers. Four active TypeScript projects can exhaust 32 GB unaided. Unallocated RAM serves as page cache for the pnpm store and cargo registry, which dominates repeat build times. **Disk is the resource to watch**, not CPU. Rust `target/` directories reach tens of GB across a few projects, and electron-builder caches a separate Electron binary per platform and arch. See §5.6 for required hygiene. Expect 80–150 GB in steady state; 320 GB is adequate but not lavish. Alert at 80% usage. ### 4.2 Network - Install Tailscale; join the operator's tailnet. - **Close port 22 to the public internet.** SSH is reachable over the tailnet only. - Enable MagicDNS so the host resolves as `$FLIT_SERVER`. - UFW: deny all inbound except the Tailscale interface. - **Disable node key expiry for this machine** — see below. This is not optional. #### Node key expiry is a scheduled lockout Tailscale nodes created from an auth key inherit key expiry, **180 days by default**. When it lapses the machine drops off the tailnet, and since public SSH is closed, the operator is locked out of their only development machine — roughly six months after build, with no warning, quite possibly while travelling. Two requirements: 1. **Disable key expiry** for this node in the Tailscale admin console, and assert it in `bootstrap.sh` phase 3 rather than trusting it was done. Use the **persistent** auth key here — not the ephemeral one the drills use (§7.1), which would make the production node vanish from the tailnet the moment it went briefly offline. 2. **`verify.sh` asserts remaining validity** from `tailscale status --json`. FAIL if expiry is enabled and under 30 days out. A machine that will become unreachable is drift, and it is the one kind the operator cannot fix after the fact. #### Break-glass: a console password Hetzner Cloud provides a browser console that bypasses SSH entirely — so a firewall or Tailscale lockout *is* recoverable, contrary to what this section previously claimed. But only if the console itself can be logged into. `bootstrap.sh` sets a strong password for `$FLIT_USER` and escrows it to Proton Pass alongside the backup credentials (§6.2). `PasswordAuthentication no` (§4.4) keeps it useless over SSH; it exists solely for the browser console. Without it the console opens to a prompt nothing can satisfy, and a recoverable mistake becomes a rebuild. Cheap insurance against the highest-consequence failure in the build. #### Hard ordering constraint Join the tailnet and confirm connectivity *before* closing public SSH. The guard has two halves, and **the important half runs on the laptop**: - **`bootstrap.sh` (laptop) must prove inbound reachability** — open a fresh SSH connection to the MagicDNS name over the tailnet and run a marker command. Only on success does it invoke `40-firewall.sh`, passing `--tailnet-verified`. - **`40-firewall.sh` (server) refuses to run without that flag**, and independently re-checks `tailscale status` before applying rules. A server-side check alone is insufficient and was the earlier mistake here. `tailscale status` reporting connected proves the *server* reached the tailnet; it says nothing about whether the *operator* can reach the server. Tailnet ACLs, unpropagated MagicDNS, or a laptop not on the tailnet all break inbound while server-side status looks perfectly healthy. If the inbound check fails, phase 3 stops with public SSH still open. That is the correct outcome: an instance briefly reachable on port 22 is a far smaller problem than one reachable by nobody. ### 4.3 Access One transport, many sessions. **Sessions are per project, not one shared session.** The operator runs several terminals concurrently, each attached to an independent tmux session with its own windows, working directory, and running processes. Do not configure a single `main` session. ```bash ssh $FLIT_ALIAS -t tmux new-session -A -s api -c '~/workspace/api' ``` `new-session -A` attaches if the session exists and creates it otherwise, so one command covers both cases. A laptop-side helper wraps this — see `README.md` §2. The operator's local `~/.ssh/config` needs this block: ``` Host $FLIT_ALIAS HostName $FLIT_SERVER ServerAliveInterval 30 ServerAliveCountMax 3 ControlMaster auto ControlPath ~/.ssh/cm-%r@%h:%p ControlPersist 10m ForwardAgent yes ``` `HostName` is explicit rather than left to alias-equals-hostname: `$FLIT_ALIAS` (default `flit`) is what the operator types, `$FLIT_SERVER` (default `dev`) is the Hetzner server name and the MagicDNS hostname it resolves to, and the two are allowed to differ (§4.4). `ForwardAgent` is what keeps git credentials off the server entirely (§7.2). > **This section states requirements, not manual steps.** `laptop/install.sh` installs the SSH > block and the shell helpers; `25-terminal.sh` sets `window-size latest` server-side. See §5.7, > and do not also apply them by hand. **SSH only. mosh is not installed** — see §10 for why, and do not reintroduce it. Resilience comes from tmux plus the keepalive and multiplexing settings above: a dropped connection costs a reconnect, never state. What SSH does not provide is predictive local echo, so **typing in `vim` from Asia will feel the round trip** — roughly 180 ms per keystroke in insert mode. That is the accepted cost (§10), and it is why success criterion 5 says *tolerable*. A UTF-8 locale is still set system-wide (`en_US.UTF-8` via `locale-gen` in `10-packages.sh`, asserted by `verify.sh`). It was originally a mosh requirement; it stays because vim, tmux, and `ripgrep` all mishandle unicode filenames and box-drawing characters without it. `loginctl enable-linger $FLIT_USER` so tmux survives logout. ### 4.4 User account Hetzner's Ubuntu images provide `root`. All work runs as an unprivileged user. - **Username: `$FLIT_USER`** (default `dev` — set once in `lib/common.sh`, sourced by both machines), created during bootstrap phase 1 via cloud-init `user_data`, not by a later script — the account must exist before phase 2 rsyncs anything into its home directory. - Member of `sudo`, passwordless sudo for provisioning only. - The operator's public key is installed at creation; `PermitRootLogin no` and `PasswordAuthentication no` in sshd. - **A strong password is set for `$FLIT_USER`**, escrowed to Proton Pass **and read back before the server is created** — escrowing after creation would risk a machine existing whose break-glass password lives nowhere. It cannot be used over SSH — `PasswordAuthentication no` — and exists solely for the Hetzner browser console when Tailscale or UFW has made the machine otherwise unreachable (§4.2). Without it the console is a login prompt nothing can satisfy. - `loginctl enable-linger $FLIT_USER`, so tmux sessions survive logout. - `~/workspace` created at provisioning, owned by `$FLIT_USER`. Every project session, the pnpm store invariant (§5.5), and the backup set (§6.2) assume it exists. - **Passwordless sudo is permanent**, not revoked after provisioning. `bootstrap.sh` is applied by hand whenever `verify.sh` reports drift (§5.3), and its server-side `server/lib/` scripts need sudo on an ongoing basis. On a single-operator box reachable only over the tailnet this is the honest trade; do not add a revocation step that the next `bootstrap.sh` run would have to undo. `$FLIT_USER` (a rename of the earlier `DEV_USER`, same value and role) is one of five configuration inputs `lib/common.sh` defines and both machines source, each defaulting to the value this project has always used so a fresh clone works unconfigured on either side: | Variable | Default | Derives | | --- | --- | --- | | `FLIT_NAME` | `flit` | `/etc/$FLIT_NAME` and `/var/lib/$FLIT_NAME` on the server, the server repository directory `~/$FLIT_NAME`, the run-log directory, the vault name, the Hetzner ssh-key name `$FLIT_NAME-operator`, the Storage Box name `$FLIT_NAME-backup`, systemd unit names, and healthchecks slugs (`lib/jobs.sh`) | | `FLIT_SERVER` | `dev` | The Hetzner server name and the ssh `HostName` | | `FLIT_ALIAS` | `flit` | The ssh `Host` alias the operator types, and the two shell helper function names (§4.3) | | `FLIT_USER` | `dev` | The operator account on the server | | `FLIT_LAPTOP_HOME` | none | The laptop's home directory, pushed at run time; there is no committed default, since one would name a path that exists on nobody's laptop but the one it was written on | **Changing these values does not rename an already-provisioned server.** They are inputs consulted when a machine is built, not a description read back from one: the unix account, the Hetzner server and the tailnet hostname all keep whatever they were created as. What a new value changes is the project's own naming and what the operator types. `ssh $FLIT_ALIAS` resolves to `HostName $FLIT_SERVER`, so the alias and the server may differ, and after a rename they usually do — which is why `FLIT_ALIAS` exists separately at all. `README.md`'s "This deployment" section records the values a given deployment actually holds. Every `~` and `$USER` elsewhere in this spec means `/home/$FLIT_USER` and `$FLIT_USER`. Nothing in the design runs as root after provisioning. ### 4.5 Getting a project onto the box `laptop/push-project.sh [remote-name]` copies a working directory from the laptop to `~/workspace/`. It is the same push model as everything else here (§5.1): the laptop holds the authoritative copy, the server never fetches one, and there is no `--pull`. **Why a copy and not a clone.** Cloning the repository on the server is one command and reproduces the committed half of a project. The half it drops is what makes a checkout a *working* directory: a gitignored `.env` or `.dev.vars`, a `handoff.md`, `.claude/settings.local.json`, an edit made and not yet committed. Those are precisely the files a project cannot be resumed without, and precisely the ones no clone can produce. **What is excluded, and why that list is short.** Anything the server can rebuild from a file that did transfer — `node_modules` from a lockfile, `target/` from `Cargo.toml`, `dist/`, `.venv`, `.wrangler`, build caches. This is not only a bytes argument: a `node_modules` built on an arm64 laptop contains compiled bindings that do not run on the amd64 server, so copying it produces a directory that is worse than an absent one. `.worktrees` and `.git/modules` are excluded for a different reason — both hold absolute laptop paths that name nothing on the server. A project needing more adds patterns to a `.flit-push-exclude` file in its own root; the committed list stays general. **A name is not a fact, so git overrules the general list.** `build` in that list means "output some command regenerates", but nothing stops a repository from tracking a source file at `apps/publisher/build/docs.ts`, and one does. Dropping it makes the server's checkout permanently dirty with a deletion no later push can heal, because the exclusion that caused it also blocks the file that would repair it — and every push from then on stops to ask about overwriting changes nobody made. So `git_protect_list()` (`lib/common.sh`) names every tracked file, and every ancestor directory of one, ahead of the general patterns; rsync takes the first rule that matches. Ancestors are not optional: rsync prunes an excluded directory during the walk and never looks inside it, so protecting the file without protecting `apps/publisher/build/` protects nothing. The three tiers, in the order the filters are assembled: a project's own `.flit-push-exclude`, then git's tracked files, then the general list. **Explicit intent beats git, and git beats a guess about a name.** A project's file outranks the protection deliberately — it is written by someone looking at that project who has decided a path does not belong on the server, and that decision covers a tracked file as readily as an untracked one, since a repository may track a data set the server has its own copy of. That is the one tier that can still strand a tracked file, so after the transfer the run asks the server for `git ls-files --deleted` and reports what is missing. git on the far side is the check rather than a second pass over the patterns, because it reports the state that resulted instead of predicting it. **Session transcripts move with the project.** Claude Code keys them by the project's absolute path with the slashes replaced by dashes, so the laptop's key and the server's are different strings and the copy is re-keyed on the way. This transfer never deletes, whatever `--delete` was given for the project: the server accumulates its own sessions for the same project, and they are different sessions rather than stale copies of the laptop's. The files are per-session and append-only, so the two histories merge. **The one destructive case is guarded.** The server is a development machine, and an uncommitted edit on it is real work. A push that would overwrite it lists the server's dirty files and asks before proceeding; `--force` answers in advance, `--print-only` transfers nothing. `--delete` is off by default, so a file existing only on the server survives a push — the mirror is opt-in because the destructive reading of "sync" is the wrong default in the one direction that cannot be undone. **A clean tree is not proof the laptop is still authoritative.** The dirty-tree guard catches an uncommitted edit on the server, but an agent working there can commit its own work, leaving a clean tree that is ahead of the laptop rather than behind it — indistinguishable to the guard, since both read as clean. Because the push is one-way and rsync does not merge, overwriting that copy with the laptop's would lose the work outright rather than merge with it. A project whose authoritative copy is temporarily the server's carries a `.flit-no-push` file in its root; `push-project.sh` refuses before transferring anything and prints the file's first five lines as the reason. `--force` overrides the dirty-tree guard by design, so it does not also lift this one, and `--print-only` refuses too — a dry run against a project on hold is not a case worth reporting. Removing the file is what lifts the hold. **chook is a project like any other.** Its source lives under its own remote, so it reaches `~/workspace/chook` through `push-project.sh` and carries its `.git` the way every other project here does, rather than through `bootstrap.sh`'s `phase_2_deliver()` as a flat `git archive HEAD` with no `.git` at all — a form that left work done on the box with nowhere to commit. It is the one project that must be pushed to a freshly provisioned box before an agent runs there: the delivered `settings.json` wires every hook to `chook.mjs` under `~/workspace/chook`, and each one fails until that checkout exists. `chook.toml` is unaffected — still part of the `~/.claude` payload, still rendered and linked to `~/.config/chook.toml` (§5.8); only the chook source moved. --- ## 5. Provisioning **No Docker.** See §10 for why, and do not reintroduce it. ### 5.1 Model: push, not pull The laptop provisions the server. The server never fetches its own configuration. A pull model has a bootstrap cycle: the `server/lib/` scripts live on the operator's knot, cloning them needs SSH authentication, and configuring SSH authentication is part of what `45-ssh-config.sh` does. Breaking that cycle requires either a public repository or a bootstrap token on the box — both of which add something this design is trying to avoid. Push has no such cycle. `bootstrap.sh` runs on the laptop: provisions the instance, rsyncs the repository over SSH, and runs the `server/lib/` scripts remotely. No credential is needed on the server to obtain its own configuration, and the Hetzner API token never leaves the operator's machine. `bootstrap.sh` itself runs under `pass-cli run` (§7.1), so the Hetzner API token, Tailscale auth key, Storage Box credentials, and healthchecks read-write API key are all vault-resolved. One secrets mechanism, not two. The **ntfy token is not in this list** — it lives at healthchecks.io as an integration and never reaches either machine (§6.3). `pass-cli` must therefore be installed on the laptop **before** `bootstrap.sh` runs. It is a prerequisite, not something `laptop/install.sh` provides — that script runs inside bootstrap, which is already too late. #### Infrastructure state **No Terraform.** Provisioning is driven by the `hcloud` CLI: `bootstrap.sh` phase 1 creates the production instance, and both drills create and destroy their own the same way. There is no state file. The record of what exists is the Hetzner API, with the instance's identity in its labels. Two rules made this a decision rather than an omission, and having no state satisfies both: - **State must not live on the laptop.** Local state means a lost laptop orphans the record of what exists, leaving instances nobody can cleanly destroy. - **Drills must never share state with production.** A drill's destroy step could otherwise remove the production server's entry, and concurrent runs corrupt state outright. The drills are deliberately outside any managed infrastructure — they are disposable by definition. Terraform would manage the instance only; everything that makes the box a development environment lives in `server/lib/` and is outside its scope. Reproducibility is proven by `drill-rebuild.sh` rebuilding from a stock image (§8.2), which is stronger evidence than a state file. `infra/README.md` records the reasoning and the trigger for revisiting it — a second machine, even a temporary one for a migration, at which point a remote backend must be chosen alongside it. ### 5.2 Repository layout A dedicated `$FLIT_NAME` repository on `tangled.org` (default name `flit`). Dotfiles remain a separate repository — they are used on the laptop too — and are **delivered by rsync in phase 2**, not cloned on the server (§5.4). ``` flit/ ├── bootstrap.sh # RUNS ON LAPTOP — provision, push, execute ├── infra/ │ └── README.md # why there is no Terraform here (§5.1) ├── laptop/ │ ├── install.sh # configures the operator's macOS machine │ ├── prepare-vault.sh # collects the web-UI credentials (§12.2) │ ├── healthchecks.sh # creates/upserts checks from lib/jobs.sh — needs the RW key (§6.3) │ ├── provision-storagebox.sh # creates the Storage Box, escrows its credentials — phase 9 (§6.2) │ ├── backup-credentials.sh # generates + escrows the restic password and Storage Box key (§6.2) │ ├── prune-tailnet.sh # sweeps stale drill nodes — needs the operator vault (§8) │ └── push-project.sh # moves a working directory to ~/workspace/ (§4.5) ├── server/ │ ├── lib/ # each run over ssh by bootstrap.sh's run_server_lib(); the numbers │ │ # name the files, not the run order (§13.2) │ │ ├── 05-flit-env.sh # writes /etc/$FLIT_NAME/env — phase 4 (§7.3) │ │ ├── 10-packages.sh │ │ ├── 20-toolchains.sh │ │ ├── 22-claude-config.sh │ │ ├── 25-terminal.sh # tmux.conf, vim, env vars (§5.7) │ │ ├── 30-tailscale.sh │ │ ├── 40-firewall.sh # must run after 30, see §4.2 │ │ ├── 45-ssh-config.sh # server-side Tangled host block (§7.2) │ │ ├── 50-caches.sh │ │ ├── 60-backup.sh │ │ └── 70-timers.sh # installs timers + job wrapper from lib/jobs.sh (§6.3) — NO healthchecks API │ ├── bin/ # the jobs the timers run │ │ ├── job-wrapper.sh # /start, exit code, log body (§6.3) │ │ ├── backup.sh # → $FLIT_NAME-backup-daily (§6.2) │ │ ├── restic-check.sh # → $FLIT_NAME-restic-check-weekly (§6.2) │ │ ├── check-disk.sh # → $FLIT_NAME-disk-daily │ │ ├── check-node-pin.sh # → $FLIT_NAME-node-pin-weekly (§5.4) │ │ ├── check-project-config.sh # → $FLIT_NAME-project-config-weekly │ │ └── reset-tailscale-clone.sh # clone identity guard, Before=tailscaled (§6.1) │ └── files/ # the only files pushed verbatim rather than generated │ ├── mise.toml # pinned Node → ~/.config/mise/config.toml (§5.4) │ ├── cloud-init.sh # rendered into user-data at server create (§4.2) │ └── restore-cloud-init.sh # the same, for a host booted from a backup (§8.1) ├── fixtures/ # minimal pnpm workspace + Rust crate for drill assertions (§8) ├── lib/ │ ├── jobs.sh # the one committed table of scheduled jobs (§6.3) │ ├── common.sh # THE VARIABLE SET (§4.4) and every shared primitive │ ├── drill-common.sh # what both drills share — instance creation, assertions (§8) │ ├── common.test.sh # the contract for cfg_apply() and the rest (§5.8) │ └── reset-tailscale-clone.test.sh # the clone guard, exercised via GUARD_ROOT (§6.1) ├── verify.sh # read-only assertions; nonzero exit on drift ├── drill-restore.sh # §8.1 — from backup; can the work be recovered? ├── drill-rebuild.sh # §8.2 — from stock Ubuntu; are the scripts complete? ├── .githooks/pre-commit # rejects a staged .env holding a literal (§7.1) ├── local.conf.example # HCLOUD_SSH_PUBKEY, HCLOUD_LOCATION — copy to the │ # gitignored local.conf (§12.1) └── .env # pass:// references for every read item in §7.1's # vault inventory; safe to commit (values are references) ``` ### 5.3 Idempotency and drift detection Every `server/lib/` script **must be idempotent.** Each is invoked separately by `run_server_lib()`, so each checks current state before acting on its own; a second run of any of them on a converged machine is a no-op. This is required so that re-running one is always safe, not so that it can be run automatically. > **Guard on observed state, never on a sentinel file.** `[ -f /var/lib/$FLIT_NAME/.done ] && > exit 0` satisfies "a second run is a no-op" perfectly while guaranteeing the machine never > converges again — the cheapest way to pass this requirement without meeting it. Every guard > below asks the system what is true, not what a previous run claimed. Installers that are *not* naturally idempotent must be guarded on a state check, not on their own error handling. The known ones here: | Step | Guard | | --- | --- | | `rustup` | `command -v rustup` — the installer errors on an existing installation and would fail the whole run | | `mise` | `command -v mise`; runtime installs are idempotent once it exists | | Claude Code native installer | `command -v claude` | | Storage Box key upload | Skip if the key already authenticates (§6.2) | | Credential escrow | Skip if vault readback already matches (§6.2) | | `restic init` | Skip if `restic cat config` succeeds (§6.2) | | healthchecks checks | `unique: ["slug"]` upsert (§6.3) | **Detection is automated. Remediation is attended.** - **Weekly systemd timer runs `verify.sh` only** — read-only assertions, nonzero exit on divergence, alert to the operator. > **Weekly, not monthly, because detection must not lag the remedy.** Hetzner's automatic > Backups keep 7 rolling daily slots (§6.1). Drift found on day 28 of a monthly cycle is three > weeks past the last rollback point that could have undone it, and the offending change is > buried under a month of others. Weekly costs nothing — the run is read-only. > > Even weekly exceeds the 7-day window, so be clear about what covers what: **Hetzner Backups > cover breakage noticed immediately; restic and the rebuild drill cover everything else.** > Do not treat the 7 slots as a safety net for anything `verify.sh` finds. It checks that expected packages and toolchains are present, timers are **loaded** (unit files valid and installed), and the §5.5 filesystem invariant holds. It asserts timers are *enabled and scheduled* only when `--no-external-state` is not in effect: drills deliberately install the backup timer without enabling it (§12.1), so asserting "firing" unconditionally would make §8.2 unsatisfiable. It also asserts the §5.7 terminal settings — `set-clipboard` is `external` and not `on`, `allow-passthrough` is on, and `CLAUDE_CODE_DISABLE_ALTERNATE_SCREEN` is exported. All assertions query effective state rather than file contents, per §5.8. #### How `verify.sh` must report Three requirements. Each exists because violating it produces alerts the operator learns to ignore, which costs the entire drift-detection layer. **1. Run every assertion. Never `set -e`.** Each check runs independently, records its own result, and the script reports all of them. Exiting at the first failure means three drifts take three months to surface at a monthly cadence. **2. Three outcomes, not two: PASS, FAIL, and SKIP.** A check that *cannot run* is not a check that failed. | Situation | Outcome | | --- | --- | | Assertion evaluated, condition false | **FAIL** — real drift, alert | | `releases.electronjs.org` unreachable, so the Node pin cannot be compared | **SKIP** — not drift | | `pass-cli` token expired, so the vault cannot be queried | **SKIP** — see below | Only FAIL pings `/1`. A run of PASS and SKIP pings `/0` — **subject to the floor below.** > **A run that skips everything must not report healthy.** As written above, a `verify.sh` that > SKIPped every assertion would ping `/0` and be indistinguishable from a green box. That is the > cheapest possible wrong implementation, and it silently disables the entire drift layer. > > Therefore: **every assertion declares whether it is skippable.** Only one is — the vault escrow > check (needs an unexpired token). Everything else must produce PASS or FAIL, including both > Tangled assertions (§7.2): the box carries its own key now, so neither key mode nor push has a > missing-agent case left to skip for. If a non-skippable assertion cannot run, that is a > **FAIL**, not a SKIP. > > The Node pin comparison is **not** on this list: it left `verify.sh` entirely and lives in > `bin/check-node-pin.sh` with its own check (§5.4). One assertion, one script, one cadence. > > `verify.sh` also emits its PASS/FAIL/SKIP counts in the ping body, so the dashboard shows what > was actually evaluated rather than only the exit code. **3. The escrow assertion must not cry wolf.** Checking that the backup credentials are still in Proton Pass (§6.2) needs a working `pass-cli` token — which expires every 90 days by design (§7.3). If an expired token reported "backup credentials missing from vault," the highest-urgency alert in the system would fire every quarter for a routine, harmless reason, and would be disregarded by the second occurrence. Distinguish *cannot query* (SKIP, with a distinct "token expired" message) from *queried and absent* (FAIL, urgent). > **The same SKIP fires from `bootstrap.sh` on a token that is perfectly valid, every time.** > `vault_session_ok()` (`lib/common.sh`) needs either a live cached `pass-cli` session on the > server or `PASS_CLI_TOKEN` in the process environment. `bootstrap.sh`'s `flit_remote_env()` > forwards FLIT_NAME, FLIT_USER, FLIT_SERVER, FLIT_LAPTOP_HOME and FLIT_IN_DRILL to > `phase_11_verify()`'s remote run — deliberately not the token, which reaches the server only > through `laptop/backup-credentials.sh` escrowing it into `secrets.env` (§7.3) — so that path > reaches neither and SKIPs regardless of the token's actual state. The scheduled run gets > the token from the timer unit's `EnvironmentFile=-/etc/$FLIT_NAME/secrets.env` line > (`70-timers.sh`, §7.3) and PASSes on the same token. A SKIP reported by `bootstrap.sh` is > therefore not evidence that the token is missing or expired — only the weekly timer, or an > on-demand run under the unit's own environment (`systemd-run --property=EnvironmentFile=...` > against the escrow check's unit), actually redeems the token and settles that question. Do not > "fix" the SKIP by adding the token to `flit_remote_env()`: that would put an unattended-timer > credential into an interactive laptop-to-server channel for a check whose real answer is > already available, weekly, without it. Assertions must also be independent: a missing `pnpm` should report as its own FAIL, not cascade into a confusing error inside the §5.5 filesystem check. - **Remediation is `bootstrap.sh --target $FLIT_SERVER` from the laptop** when `verify.sh` reports something, or when the repository changes. > **Not a hand-run of the `server/lib/` scripts on the server.** Check creation lives in > `laptop/healthchecks.sh` (§6.3), because it needs the healthchecks read-write API key — which > stays in the vault and never reaches the server (§7.1). A hand-run of `70-timers.sh` on the box > installs timers and job scripts but cannot reconcile the checks themselves. > > Remediation therefore goes through the same push path as the original build: one code path, and > the credential stays where §7.1 says it stays. The `server/lib/` scripts remain directly > runnable for diagnosis of the server half. > **Do not put the `server/lib/` scripts on an unattended timer.** Bash idempotency is easy to get subtly > wrong — `apt install` is naturally idempotent, `echo >> file` is not. A monthly unattended > re-run of a script containing one non-idempotent line can break a working machine at 3am on > the operator's only box. That is a worse outcome than the drift it would prevent. **What this does not catch.** `verify.sh` detects drift *away* from the scripts — a deleted file, an edited config, a stopped timer. It does not detect additions: a package installed by hand at 2am to unblock a build and never committed to any `server/lib/` script. Asserting the absence of unexpected packages is impractical, so this class of drift accumulates silently and is caught only by the rebuild drill in §8.2. That drill, not this timer, is what verifies the scripts are complete. ### 5.4 Packages and toolchains **Base:** `tmux`, `vim`, `git`, `ripgrep`, `fd-find`, `curl`, `jq`, `ufw`, `restic`, `build-essential`, `pkg-config`, `libssl-dev` **Runtime version management: `mise`.** Chosen over `fnm` because the box is polyglot and `fnm` is Node-only; adding Python or Go later is then a one-line change to `.mise.toml`. Manage Node through `mise`; enable `pnpm` via `corepack`. **Node version: pinned to the version bundled by the current Electron stable release.** Keeping the dev runtime identical to the one the app ships against avoids a class of bug that only appears in packaged builds. The pin is a **literal version string** committed to the repository at `server/files/mise.toml`, which `20-toolchains.sh` installs to **`~/.config/mise/config.toml`** — mise's global config. > A `.mise.toml` at the `$FLIT_NAME` repo root would pin Node *for that repository only*. mise > resolves configuration per directory, so the global default must go in `~/.config/mise/`. > Per-project `.mise.toml` files then override it normally. Do not have `20-toolchains.sh` resolve the version at install time — that reintroduces exactly the unattended drift §5.3 argues against, and would silently change the runtime under a working project. Instead, **`check-node-pin.sh` runs weekly** (`$FLIT_NAME-node-pin-weekly`, §6.3) and compares the pinned version against the newest stable semver entry in `https://releases.electronjs.org/releases.json`. The feed has used both `deps.node` and `node` for the bundled Node version, so the parser accepts either rather than binding the dead-man to one revision of an external schema. The operator bumps `server/files/mise.toml` deliberately and re-runs `bootstrap.sh --target $FLIT_SERVER`. > **`verify.sh` does not duplicate this check.** One assertion, one implementation, one cadence. > An earlier draft had `verify.sh` doing the comparison monthly *and* a separate weekly check — > which would have meant whichever an agent implemented, the other check went DOWN with nothing > ever pinging it. > > The general rule: **each check in §6.3 maps to exactly one script**, and no assertion lives in > two places at two cadences. **Divergence is advisory; staleness is a failure.** `check-node-pin.sh` pings `/0` when it finds a newer Electron stable, putting the version difference in the ping body — visible in the dashboard, no alert. It pings `/1` only when the pin is **two majors or ~16 weeks behind**. The reason is temporal. Electron ships a major every 8 weeks, so a check that FAILed on any divergence would sit DOWN for most of the year. Two things break when it does: healthchecks alerts on the *flip* to down and then goes quiet, so an already-down check cannot signal anything new — and while it is down, "running and reporting divergence" becomes indistinguishable from "stopped running entirely," which is the exact property §6.3 exists to provide. A permanently red check is not a check. Expect this to fire a few times a year: Electron ships a major every 8 weeks and moves to even-numbered Node versions as they reach Active LTS. Projects that need a different runtime override it with their own per-project `.mise.toml`; the pin above is the global default, not a constraint on every repo. **Rust: `rustup`, not `mise`.** `rustup` is the standard path and handles cross-compilation target management properly, which matters for the deferred macOS work (§11). Also install `sccache` and `cargo-sweep` — see §5.6. **Electron (Linux targets):** electron-builder's Linux build dependencies. Verify these before the first build rather than discovering them mid-task. **Headless Electron testing:** `xvfb`. Tests run under `xvfb-run`; no display forwarding is required and none should be configured. **Other:** Tailscale; Proton Pass CLI (`pass-cli`); Claude Code via the **native installer** (npm and homebrew are legacy paths — fetch the current command from rather than hardcoding it here); dotfiles symlinked from the copy delivered in phase 2. > **Dotfiles are pushed from the laptop, never cloned on the server.** Cloning them from > `tangled.org` would need SSH authentication on the box, and configuring that authentication is > itself part of what phase 7 does (§7.2) — the same bootstrap cycle §5.1 describes for the > `$FLIT_NAME` repository, and it still applies here: phase 2 delivers dotfiles before phase 7 > installs the box's own Tangled key, so there is nothing to authenticate with yet at delivery > time. In a drill the gap is permanent rather than an ordering problem — `--no-external-state` > skips installing the key entirely (§7.2), so no later phase closes it. `bootstrap.sh` rsyncs the > dotfiles repo alongside `$FLIT_NAME` in phase 2. The server holds a copy, not a working clone; > the operator commits from the laptop. ### 5.5 pnpm configuration One non-obvious requirement: > The pnpm content-addressable store and every `node_modules` directory **must sit on the > same filesystem.** pnpm hardlinks from store to `node_modules`; across a filesystem > boundary the hardlink fails and pnpm silently falls back to copying. There is no error — > only slower installs and multiplied disk usage. Keep `~/workspace` and the pnpm store on the same volume. Set `store-dir` explicitly in `~/.npmrc` rather than relying on the default. This is an assertion, not a manual check — it belongs in `verify.sh`: ```bash test "$(stat -c %d "$(pnpm config get store-dir)")" = "$(stat -c %d ~/workspace)" ``` `%d` is the device number — the standard way to compare filesystems. On the current single-disk design this always passes; it is cheap insurance for the day a volume is added, or for the dedicated-hardware move in §11 where two NVMe devices are involved. ### 5.6 Build cache hygiene Required, not optional — this is what keeps §4.1's disk budget honest. - **`sccache`** as `RUSTC_WRAPPER`, so shared dependencies compile once across projects. - **Shared `CARGO_TARGET_DIR`** rather than a `target/` directory per project. - **Monthly systemd timer:** `cargo sweep --time 30` and `pnpm store prune`. Both stores grow monotonically otherwise. - **Daily disk check at 80%** — the `$FLIT_NAME-disk-daily` check in §6.3. > **The prune timer deliberately has no healthchecks check of its own**, which is the one > exception to §6.3's rule that everything scheduled is dead-man-covered. Its failure mode is > gradual disk growth, and `$FLIT_NAME-disk-daily` detects that within a day. A ninth check would fire > on a job whose only symptom is already monitored. Noted so the omission reads as a decision > rather than an oversight. ### 5.7 Terminal ergonomics **Required, not cosmetic.** Without these the box is unpleasant to use by the second day, so they are scripted, not left to the operator. Rationale for each is in `README.md` §4–5. **Server — `25-terminal.sh` writes `~/.tmux.conf`:** ``` set -g allow-passthrough on set -g set-clipboard external set -g window-size latest set -g mouse on ``` `set-clipboard external` over `on` is a security decision, not a default. Under `on`, any process inside tmux can set the system clipboard — the operator runs Claude Code against untrusted repositories, and clipboard hijacking followed by a paste into a shell is a real attack path. `external` confines clipboard-setting to tmux itself, at the cost of copying via copy-mode rather than vim's yank register. **Do not "fix" this by setting `on`.** **Sessions survive a reboot.** `25-terminal.sh` clones `tmux-resurrect` and `tmux-continuum` at pinned tags into `~/.tmux/plugins` (no plugin manager) and adds `@continuum-restore on` to the managed block. Continuum saves every 15 minutes and restores when a tmux server starts, so `25-terminal.sh` also installs `$FLIT_NAME-tmux.service`, which starts one at boot through a login shell (systemd's own environment lacks the `PATH` a restored `claude` needs). Sessions, windows, panes and working directories return; running programs do not, except `claude`, which `@resurrect-processes` relaunches as `claude --continue` — the most recent conversation in that directory, so several panes in one directory resume the same one. `verify.sh` checks the setting and both plugin files. **Server — shell profile:** ```bash export CLAUDE_CODE_DISABLE_ALTERNATE_SCREEN=1 ``` Forces Claude Code's classic renderer, keeping the conversation in native scrollback where tmux copy-mode reaches it. Must be set **on the server**, where Claude Code runs; setting it on the laptop is a no-op and the obvious mistake in this architecture. > **This decision was re-justified when mosh was removed (§10).** The earlier argument was that > mosh already handled the bandwidth problem, making the fullscreen renderer's viewport-diffing > redundant. Without mosh that is no longer true: on a slow link, fullscreen genuinely sends > fewer bytes. > > Classic still wins, on a different basis. `set-clipboard external` (above) makes tmux copy-mode > the *only* path from the server to the system clipboard, and copy-mode needs the conversation > in native scrollback. Fullscreen would move it into the alternate screen buffer and its mouse > capture would fight copy-mode directly. Losing the clipboard costs more, every day, than the > bandwidth costs on occasional slow links. **Laptop → server — `push_terminfo()`, called from `phase_6_terminal()` right after `25-terminal.sh`:** pushes the operator's terminal description to the server. Ubuntu 24.04 ships ncurses 6.4, whose terminfo database predates several terminals in current use — `xterm-ghostty` arrived in ncurses 6.5, and `ncurses-term` does not help, since the entry does not exist at 6.4. ssh forwards `TERM` verbatim, so tmux on the server exits with `missing or unsuitable terminal` the first time it meets a description it has never indexed — exactly what happened on the first real use of the box. Only the laptop knows which terminal the operator sits at, so the push runs in that direction: `infocmp -x "$TERM"` piped to `tic -x -` over `remote`. `tic` runs without sudo, writing `~/.terminfo` — the database tmux consults — rather than the system database, which both skips a root-owned path under the operator's home and needs no privilege escalation to do it. Deliberately not gated by `--skip-laptop`: it writes to the server and only reads from the laptop. It skips quietly when `TERM` is unset, `dumb`, or `unknown`, when the laptop has no `infocmp`, or when the description already resolves on the server; on failure it defers (§5.8) with the exact by-hand command. **Laptop — `laptop/install.sh` writes:** - `~/.ssh/config` — the `Host $FLIT_ALIAS` block from §4.3, including `ForwardAgent` - `~/.config/ghostty/config` — `clipboard-read`, `clipboard-write`, `copy-on-select` - Shell rc — the `$FLIT_ALIAS()` and `${FLIT_ALIAS}s()` helpers from §4.3 (default `flit()` and `flits()`). **Detect the operator's shell from `$SHELL` and install the matching syntax.** The helpers as written are POSIX/zsh function syntax and will not parse under fish, which needs `function flit … end`. If the shell is unrecognised, defer under §5.8 and print rather than guessing. Both halves of the OSC 52 chain must agree; configuring one without the other fails silently, which is why they are installed by scripts rather than documented as steps. All writes in this section are subject to §5.8 — existing configuration is never overwritten, and anything that cannot be applied surgically is printed for manual application instead. **Not scriptable:** dev servers must bind `0.0.0.0` rather than `127.0.0.1` to be reachable over the tailnet. This is per-project configuration. `verify.sh` cannot check it; the README documents it as a symptom instead. ### 5.8 Configuration file safety **Rule: never destroy or overwrite an existing configuration file.** This binds every script that touches a config file — `laptop/install.sh`, `25-terminal.sh`, `45-ssh-config.sh`, and anything added later — on both the laptop and the server. #### Managed blocks Where a script must edit a file it does not own, it edits **only** between delimiters: ``` # >>> $FLIT_NAME managed >>> ... # <<< $FLIT_NAME managed <<< ``` `lib/common.sh` also recognises the markers this project wrote before the rename to flit — `# >>> dev-server managed >>>` / `# <<< dev-server managed <<<` — and deletes a legacy block in the same rewrite that installs the current one, so a block `cfg_apply()` wrote under the old text is replaced rather than orphaned alongside a second, new one. - **File absent → create it**, along with any missing parent directories. Creating a file that did not exist destroys nothing, so it is not a violation of this section's rule. This is the common case on a fresh Mac, which ships with no `~/.zshrc`, no `~/.ssh/`, and no `~/.config/ghostty/` until Ghostty has been launched once. Create `~/.ssh` mode `0700` and `~/.ssh/config` mode `0600`, or OpenSSH will refuse to use them. - Block absent, file present → insert it (see placement note below). - Block present → replace its contents only. - Anything outside the markers → never read for meaning, never modified, never reordered. Ghostty reads `~/.config/ghostty/config` but also honours a macOS-native path under `~/Library/Application Support/`. **Detect which the operator actually has** before writing; if both exist, defer under the rule below rather than guessing which one wins. Prefer indirection over inline content where the format supports it, so the managed block stays a single line: `Include ~/.ssh/config.d/$FLIT_NAME` for SSH, `config-file = …` for Ghostty, `source-file` for tmux, `source …` for shell rc. The payload then lives in a file the scripts fully own, and the operator's file gains one line. > **Placement is format-specific, and getting it backwards fails silently.** > > - **SSH: prepend.** OpenSSH uses the *first* value obtained for each keyword. An `Include` > appended below an existing `Host *` block is silently ignored, and `ssh $FLIT_ALIAS` quietly loses > `ForwardAgent` — breaking git push with no error. The managed block goes at the **top** of > `~/.ssh/config`. > - **Ghostty and shell rc: append.** Later values win, so the block must come last. > - **tmux: append**, same reason. > > Do not apply a single append-everywhere rule. Two of these formats want opposite ends. Take a timestamped backup before any write, including successful ones. #### When surgical application is impossible Apply nothing and **print the exact config to the terminal** for manual application. Emit the target path, the reason, and the literal text to paste — not a description of it. **Check effective state before deciding anything.** If the desired setting is already true — `tmux show -gv set-clipboard` already returns `external`, `ssh -G $FLIT_ALIAS` already reports `forwardagent yes` — there is nothing to apply and nothing to defer. Report it satisfied and move on, whether it got that way from a managed block or from the operator's own hand-editing. This matters most on repeat runs. Symlinked dotfiles are the expected case (below), so without this check every remediation run would re-print the same deferral for a config applied manually months earlier. Deferrals the operator has already acted on are noise, and noise in the remediation path is how drift alerts get ignored. > **This check has no equivalent in `deliver_rendered()`, the sibling contract > `server/lib/22-claude-config.sh` writes by hand for `settings.json` and `chook.toml`.** JSON > and TOML carry no comment syntax, so neither can hold the `# >>> $FLIT_NAME managed >>>` > markers a managed block scopes a comparison to; `deliver_rendered()` compares the destination's > whole content against the rendered file instead — canonicalised for JSON, byte for byte for > TOML. A managed block is what lets `cfg_apply()`'s effective-state check ignore everything > outside it; a whole-content comparison has no such scope, so any byte written by anything other > than the delivery reads as a real difference and defers. Claude Code itself writes a `"theme"` > key into `~/.claude/settings.json` on first run, into `settings.json` rather than the > `settings.local.json` that `deliver_rendered()` never touches and that exists for exactly this > box-local state. No operator edit is involved, and > because the comparison is whole-content rather than block-scoped, every later > `deliver_rendered()` call defers again — forever, not once — until the key is moved by hand into > `settings.local.json`, the one file the script never touches. The fix is that move, not a change > to the comparison: with no delimiter to scope around, nothing in `deliver_rendered()` can tell a > third party's addition apart from a real conflict, so keeping box-local state out of the > delivered file is the only thing that keeps this contract honest. Bail out to printing when the state is **not** already satisfied and: | Condition | Why | | --- | --- | | A key is already set to a conflicting value outside the managed block | Silently winning an override is worse than asking | | The file is a symlink into a dotfiles repository | Writing would modify a git working tree the operator manages | | The file exists but does not parse as expected | Unknown structure; assumptions unsafe | | Permissions or ownership prevent a safe write | Do not escalate | **The symlinked-dotfiles case is the expected one, not the edge case.** The operator keeps dotfiles in a repository (§5.4). Scripts must detect symlinks before writing and print rather than following them into the repo. A script that prints instead of writing has succeeded at its real job. It continues through its remaining steps and summarizes everything deferred at the end — one consolidated list, not a failure per invocation. **Exit code 3 means "completed, with deferrals."** Not 0, which would let deferrals pass unnoticed; not 1, which `bootstrap.sh` treats as a phase failure that halts the build (§12.1). Since symlinked dotfiles are the *expected* case, exit 1 would abort the normal run at phase 6. `bootstrap.sh` collects exit-3 deferrals, continues, and prints them together at the end. Support `--print-only` on every such script, so the operator can review the full set of changes before any file is touched. **For `bootstrap.sh`, `--print-only` means: mutate nothing, anywhere.** Not "skip config writes but perform the rest" — that reading would provision a CX53 during what the operator believed was a dry run. No instance is created, no package installed, no API called with a side effect. The output is deliberately asymmetric, and that is worth understanding rather than treating as a defect: | Target | What `--print-only` can report | | --- | --- | | **Anything local to the laptop** — `~/.ssh/config`, Ghostty, shell rc, and the healthchecks reconciliation in phase 8 | **Exact.** These are inspectable now: it reads the real files and the real check list, and prints precisely what it would write or defer, and why. | | **Anything on the server** — every phase that provisions, installs, or configures the box | **A plan only.** No server exists to inspect, so it lists intended actions without asserting current state. | This is a **laptop/server** split, not a phase-number split. Several phases have both halves — phase 6 configures Ghostty locally *and* tmux remotely — and work has moved between them during review. Any description keyed to phase numbers goes stale the next time something moves. `--print-only` exits **0**. It is an inspection, not a run, so the exit-3 deferral convention does not apply — nothing was deferred because nothing was attempted. #### Consequence for `verify.sh` Because configuration may have been applied by hand, in the operator's own layout, `verify.sh` must assert **effective state, not file contents** wherever effective state can be queried: ```bash tmux show -gv set-clipboard # not: grep ~/.tmux.conf timedatectl show -p Timezone --value # not: grep /etc/timezone stat -c %d ~/workspace # device number, for the §5.5 invariant ``` Grepping for a managed block would report drift on a correctly configured machine, which trains the operator to ignore the alerts. **Three constraints on how it runs**, each of which silently produces wrong answers otherwise: - **Tailscale key expiry** is asserted every run (§4.2). FAIL if expiry is enabled and under 30 days out — a machine about to become unreachable is the one form of drift that cannot be fixed after the fact. - **`verify.sh` runs as `$FLIT_USER`**, never as root. Almost everything it inspects is user-scoped — `~/.npmrc`, `~/workspace`, tmux settings, the shell profile. Run as root it reads root's configuration and reports confidently on the wrong user. The systemd timer must set `User=$FLIT_USER`. - **Environment variables have no queryable effective state.** A systemd timer does not source shell profiles, so testing `$CLAUDE_CODE_DISABLE_ALTERNATE_SCREEN` directly is always false in a timer run. Assert it through a login shell instead: `bash -lc 'echo $CLAUDE_CODE_…'`. This is the one documented exception to "query the tool" — there is no tool to query. - **Laptop-side settings are not server-side assertions.** `ssh -G $FLIT_ALIAS` describes the *operator's machine*; `verify.sh` runs on the server, which has no `Host $FLIT_ALIAS` block. Laptop configuration is checked by `laptop/install.sh --print-only`, not by the weekly timer. Do not add `ForwardAgent` or Ghostty assertions to `verify.sh`. --- ## 6. Backups and alerting Two backup mechanisms with different failure modes, both required, plus the notification channel they and everything else report through. ### 6.1 Whole-machine rollback Enable **Hetzner automatic Backups**: a 20% flat surcharge on the server price with seven slots on a rolling schedule (~€5.90/month at current pricing). This is the whole-machine undo button for a broken OS. Manual Hetzner snapshots are full compressed images with no deduplication and no automatic retention — a snapshot taken one month still bills months later unless deleted. Do not use manual snapshots as a substitute for the rolling backup. #### The golden snapshot is temporary — deleted automatically Phase 12 takes one, and phase 13's restore drill uses it because no automatic backup exists yet (§12.1). It is used exactly once: the first quarterly `drill-restore.sh` run made without `--from-snapshot` restores from a real automatic backup instead, which is proof that backup now exists, and `drill-restore.sh` deletes the golden snapshot at the end of that same passing run — the deletion-side half of the condition `phase_12_snapshot()` checks on the creation side. `phase_12_snapshot()` skips entirely once `hcloud image list --type backup` reports a backup created from `$FLIT_SERVER`, which is the creation-side half of that condition: an automatic backup existing is exactly the state in which the substitute is no longer needed. Without that condition the phase created another image on every full run, so a deletion elsewhere was undone by the next `bootstrap.sh`. Deletion is best-effort and never fails an already-passing drill — a failed `hcloud image delete` warns instead, since the cost of a leftover snapshot is a monthly charge, not a broken recovery. A months-old golden image is worse than no golden image, because it invites restoring from it and diverges further from reality as the machine changes. Once an automatic backup exists, **the scripts are the golden image**. That is §10's entire argument against Docker, and a stale snapshot lingering would quietly contradict it. #### A snapshot clone inherits the imaged server's tailnet identity `hcloud server create-image` images the whole disk, including `/var/lib/tailscale/tailscaled.state`. A server later created from that image boots holding the imaged server's node key, so the control plane treats it as the same node, renames it after the clone's hostname, and the imaged server drops off the tailnet. This occurred: a `drill-restore.sh` run — which boots the golden snapshot per §8.1 — took production off the tailnet. Because `server/lib/40-firewall.sh` applies a host-level `ufw default deny incoming` with `allow in on tailscale0`, and no Hetzner cloud firewall exists, the tailnet is the only inbound path (§4.2). Recovery required the Hetzner browser console, using the escrowed `$FLIT_USER` console password (§4.2), and `tailscale up --force-reauth`. The fix is `server/bin/reset-tailscale-clone.sh`, installed by `server/lib/30-tailscale.sh` as a systemd oneshot ordered `Before=tailscaled.service` and pulled into the same boot transaction via `WantedBy=tailscaled.service`. It compares `/sys/class/dmi/id/product_serial` — which the hypervisor writes before cloud-init or the network exist — against the serial recorded the last time this host joined the tailnet, and removes `/var/lib/tailscale` on a mismatch, so the clone registers as a new node instead of stealing the imaged server's. It exits 0 on every path so it can never block boot. `verify.sh` asserts the unit is loaded and enabled. ### 6.2 Granular file recovery **restic → Hetzner Storage Box over SFTP**, daily via systemd timer. Client-side encrypted, so the backup target never sees plaintext. > **Deliberately chosen over Object Storage.** Storage Box moved off Hetzner's legacy Robot > console into Hetzner Console and the Hetzner API — Robot's support for it ended on > 29 July 2025 — and `hcloud` (1.67+) exposes it as `hcloud storage-box`. Robot is therefore > not a fallback for anything here; where the CLI and API have a gap, the box itself is the > remaining interface. > `laptop/provision-storagebox.sh` creates it in phase 9, before `laptop/backup-credentials.sh` > runs, and escrows its username, hostname, and password — plus the derived repository URL — the > same way phase 9 escrows everything else (§7.3). There is no manual ordering step. Kept over > Object Storage for a fixed monthly price and **no egress charge on restore**, which matters > precisely on the day a restore happens under pressure. > > restic addresses it as `sftp://uXXXXXX@uXXXXXX.your-storagebox.de:23/backups`, on the > **non-standard SSH port 23**. Authenticate with a dedicated SSH keypair generated on the > server for this purpose and uploaded to the Storage Box — not the operator's Tangled key, and > not a password. > > Object Storage would also have been fully API-provisionable, but restic's S3 backend uses > `AWS_ACCESS_KEY_ID`/`AWS_SECRET_ACCESS_KEY` as variable names for *any* S3-compatible > endpoint. No Amazon service is involved either way; the naming is historical. SFTP avoids the > question entirely. Back up: - `~/workspace` (repos, including uncommitted work) - `~/.claude` - Shell history and vim undo directories Exclude: - `node_modules`, pnpm store (reconstructible from lockfiles) - `CARGO_TARGET_DIR`, `sccache` cache (reconstructible; large) Retention: `--keep-daily 7 --keep-weekly 4 --keep-monthly 6`. #### A zero exit is not a successful backup `restic backup` exits 0 when it backs up nothing, or backs up the wrong path. A job pointed at an empty directory would ping `/0` daily, forever, and the first sign of trouble would be a restore that produces nothing. The §8.1 drill catches it — but only quarterly, so up to 90 days of backups could be worthless before anyone knows. The daily job must therefore assert **content, not exit status**: - The new snapshot contains a known sentinel path under `~/workspace`. Search a materialised listing, never `restic ls | grep -q`: `grep -q` exits at the first match, the still-writing restic takes SIGPIPE, and `pipefail` reports 141 for the pipeline. The assertion then fails once the listing outgrows the pipe buffer, and fails as *sentinel missing* — the alarm for backing up the wrong path, raised by a backup that was correct. Materialising the listing also separates a genuine `restic ls` failure from a real miss, which discarding its stderr had conflated. - Its size is within an expected band of the previous snapshot — a sudden collapse to near-zero is a failure even though every command succeeded. **On the first backup there is no previous snapshot**; assert only the sentinel path and record the baseline. Do not let "no previous snapshot" silently satisfy the size check, which would make the very first backup the one least verified. **`restic check --read-data-subset=1/4` weekly** — one job, one check (`$FLIT_NAME-restic-check-weekly`), full repository coverage every four weeks. Plain `check` validates repository structure only and will pass over pack files whose contents are corrupt, so some fraction of the data must actually be read. A *separate monthly* read-data job was the obvious alternative and is wrong twice over: it would be a scheduled job with no healthchecks check of its own, and it would put one concern on two cadences — both violations of the standing rules in §13. Rotating a quarter of the repository each week gives the same coverage inside the job that already exists. #### Initialising the repository Two steps that are easy to omit and both fail on the first run. **1. Upload the SSH key to the Storage Box.** Hetzner accepts it via `ssh-copy-id` against the non-standard port, using the `-s` flag for Storage Box's restricted shell: ```bash ssh-copy-id -s -p 23 -i ~/.ssh/storagebox_ed25519.pub uXXXXXX@uXXXXXX.your-storagebox.de ``` This requires the **Storage Box password**, which is in Proton Pass alongside the username and hostname because `laptop/provision-storagebox.sh` escrowed all three moments earlier in the same phase, when it created the box. Resolve it via `pass-cli` here — it is needed only during provisioning, while the operator is present, so it does not become §7.3 residue. **2. Escrow both credentials to Proton Pass *at generation time*, before anything else.** > **This is the single most important step in the build.** The restic repository password and the > Storage Box private key are generated on the server. If the only copies live on that server, > then losing it means the backups protecting against exactly that loss are encrypted with a key > that died with the machine. The backup set is then permanently unrecoverable. > > **The escrow requires a different credential path from the rest of the build.** Proton Pass > access tokens are **read-only** — they cannot create or edit items — and §7.3 deliberately > keeps the token that way. So escrow cannot use it. > > **Generation therefore also happens on the laptop, not the server.** This is not a stylistic > choice: `60-backup.sh` runs on the box, which holds only the read-only token, so a script there > could generate these credentials but never escrow them. Generating on the server and escrowing > afterwards would also leave a window in which the only copy of the key protecting the backups > lives on the machine it protects against. > > Writing goes through the **operator's own authenticated `pass-cli` session** on the laptop. > `bootstrap.sh` is laptop-run and interactive by default, so this is available; it must verify > an authenticated session exists *before* phase 9 generates anything, rather than discovering > the problem after creating credentials it cannot store. > > **Do not "fix" this by widening the access token to read-write.** That would turn the §7.3 > residue — a token sitting on the server — into one that can rewrite the entire vault, which is > a far worse trade than an interactive step at build time. > > `bootstrap.sh` writes each item immediately after generating it and reads it back to confirm, > before `restic init` runs. **The readback must be a fresh `pass-cli` query compared against the > generated value** — not an echo of the variable already in memory, which would pass trivially > while proving nothing. This is the one place where a convincing-looking implementation costs > the entire backup set. Same for the `$FLIT_USER` console password from §4.2 — generated in > phase 1, escrowed the same way, and equally useless if it exists only on the machine it > recovers. > > Under `--non-interactive` there is no authenticated session and escrow is skipped — which is > safe only because drills also pass `--no-external-state` and generate nothing needing escrow. If the write or the readback > fails, abort phase 9 — do not proceed with a backup system whose keys exist in one place. > > `verify.sh` re-asserts both are present in the vault on every run. This is the one assertion > worth making about somewhere other than the local machine. > > **`verify.sh` deliberately does not assert that `operator-key` is escrowed**, unlike the three > items above. It runs on the server, and the server's access token cannot read > `$OPERATOR_VAULT_NAME` — that inability is the point of §7.1's separate-vault design. Adding the > assertion would require handing the server a credential that can read the vault holding its own > break-glass key, which defeats the reason the vault is separate. Leave this a non-assertion. **3. `restic init`**, once, before the first timer run: ```bash restic -r sftp://uXXXXXX@uXXXXXX.your-storagebox.de:23/backups init ``` Guard all three on already-done state so phase 9 stays idempotent (§5.3): skip the upload if the key already authenticates, skip the escrow if readback already matches, skip `init` if `restic cat config` succeeds. Without the init step the first backup fails with *repository does not exist* — at 03:00, unattended, reported only as a healthchecks failure. #### Rotating credentials `--replace NAME` on `bootstrap.sh` and on `laptop/backup-credentials.sh` rotates one of the two credentials above without touching the other; NAME is `restic-password` or `storagebox-key`, repeatable to rotate both in one run. `bootstrap.sh` validates NAME and accumulates it into `FLIT_REPLACE`; `phase_9_backups()` forwards it to `laptop/backup-credentials.sh` as an explicit `--replace` argument. `wants_replace()` in `lib/common.sh` tests exact membership of that space-separated list, so `restic-password` does not match a `FLIT_REPLACE` entry of `restic-password-next`. > **Rotation adds a key rather than replacing one.** restic keeps its master key encrypted under > any number of passwords, so `restic key add` introduces the replacement while every existing > snapshot stays readable under the old password — re-initialising the repository would have cost > every snapshot in it. `ssh-copy-id` gives the Storage Box key the same overlap for free: it only > appends to `authorized_keys`, so the old key keeps authenticating throughout. > > **That overlap window is the property the design depends on.** The vault must never hold a > credential the repository or the box refuses, because §8.3's hand recovery reads both out of the > vault with no live session to fall back on. So the replacement is escrowed under a `-next` item, > proved against the live target, and only then promoted over the live item — generate, escrow, > verify the readback, prove against the live target, promote, verify the readback again — the same > order as first-time provisioning above, with a working credential in the vault at every step in > between. > > **Removing the old restic key runs last, and behind `confirm()`.** Under `--non-interactive` > there is no prompt to decline, so the removal is deferred (§5.8) rather than skipped silently, > and an interrupted or automated rotation rests at two working keys. Two is a safe resting state; > zero is not. > > **Withdrawing the old Storage Box public key is scripted, not manual.** `ssh-copy-id` only > appends, and neither the `hcloud` CLI's `storage-box` subcommand nor the Hetzner API exposes > key management — Hetzner documents adding a key and says nothing about removing one. That is a > gap in the management plane, not a property of the box: `.ssh/authorized_keys` is reachable > over sftp like any other file on it, so removal needs no console at all. > `storagebox_withdraw_key()` downloads the file, removes the superseded key by matching its > material rather than its comment, backs up the pre-withdrawal file to a timestamped path on the box, re-uploads the > filtered file, and re-proves the current key over sftp before returning — a failure at any step > defers to the operator instead of leaving the box in an unproven state. Matching on material is > not a stylistic choice: `ssh-keygen`'s `-C` was a constant here, so every key generated before > `storagebox_key_comment()` carries the same comment, and a comment match would delete the wrong > one. That constant is the reason keys accumulated indistinguishably in the first place — a > generation marker in the comment is what makes them tellable apart by eye from now on. > > **Nothing here deletes the `-next` item either.** `lib/common.sh` has no vault-deletion > primitive, so `restic-password-next` and `storagebox-key-next` are left behind once promoted, > stale, and their removal is also deferred to the operator by name (§7.1). > > **`FLIT_REPLACE` carries the project prefix because it is read the same way `FLIT_NAME` is** — > defaulted by `lib/common.sh` and inherited by every child process. An unprefixed `REPLACE`, set > in an operator's shell for some unrelated reason, would reach `wants_replace()` regardless and > rotate a credential nobody asked to rotate (§13.2). Both rotation procedures have run against a live server. `--replace restic-password` adds a new key and promotes it, leaving the superseded key in place under the deferred-removal path above, exercised for real under `--non-interactive`. `--replace storagebox-key` generates, escrows, uploads, proves and promotes a new keypair. A backup job run to completion under the replacement credentials, writing a fresh snapshot, is the proof that matters more than either rotation's exit code. The two deferrals that path leaves behind — removing the superseded restic key, and deleting the promoted `-next` vault items — are closed out by hand, not by `laptop/backup-credentials.sh`: `restic key remove` drops the superseded key from the repository, and the `-next` items, once confirmed byte-identical to what they were promoted over, move to the vault's trash. The restic-key removal behind `confirm()` is therefore still unexercised in practice — `confirm()` itself has since run interactively elsewhere, which is how standing rule 16 in §13 was found — and `lib/common.sh` still has no vault-deletion primitive — both remain deferred as a design fact for the next rotation. The superseded Storage Box public keys were untouched by any of this run: `storagebox_withdraw_key()` did not exist yet. Withdrawal is scripted now — see the rotation design above — rather than something a console has to be opened for. A deployment's record of what a specific rotation left behind, and what has since been cleaned up by hand, belongs in `deployment.md`, not here. The design's fail-safe property held in practice as well as on paper: an early `--replace storagebox-key` attempt, run before `storagebox_key_works()` existed (§12.1, §13.2), failed its proof step and promoted nothing, leaving the old key working throughout. The vault never held a credential the Storage Box refused. **Single destination is sufficient.** The Storage Box shares a provider with the server, so an account-level suspension takes out both — but repos live on a `tangled.org`-managed knot, off Hetzner entirely. The only exposure is uncommitted work between pushes. No second off-provider target for now. Let restic own expiry via its retention flags. Do not add server-side deletion rules on the Storage Box; the two would disagree and prune snapshots restic still references. ### 6.3 Alerting Every alert in this spec — `verify.sh` drift (§5.3), the 80% disk threshold (§5.6), backup failures (§6.2), Node pin staleness (§5.4), drill results (§8) — flows through **one path**: ``` job on $FLIT_SERVER → ping healthchecks.io → ntfy push + email ``` **healthchecks.io is the load-bearing piece; ntfy is a delivery channel.** ntfy is a transport: it delivers when something calls it, and nothing calls it when a timer stops firing. Only healthchecks detects *absence*, which is the failure this layer exists to catch. Do not invert this — a design where jobs publish to ntfy directly cannot see a job that never ran. #### Everything is a check, including thresholds Eight checks, comfortably inside the 20 the free tier allows: Every check maps to exactly one script. A slug with no implementation goes DOWN within its period; an assertion implemented twice means one of the two never pings. **`lib/jobs.sh` is the single committed table both machines derive this from** — one row per job, pipe-delimited: name, `OnCalendar` expression, grace, cadence suffix, command. `laptop/healthchecks.sh` upserts a check from a row's schedule and grace; `70-timers.sh` installs a systemd timer from the same row's `OnCalendar` and command; both derive the same slug from `jobs_slug()` — `$FLIT_NAME--`. Before this file existed, the two lived as separate arrays that had to match by convention alone; a row here now has exactly one place to go wrong. | Slug | Runs | Where | Schedule | Grace | | --- | --- | --- | --- | --- | | `$FLIT_NAME-backup-daily` | `restic backup` | server | daily | 6 h | | `$FLIT_NAME-restic-check-weekly` | `restic check` | server | weekly | 2 d | | `$FLIT_NAME-verify-weekly` | `verify.sh` | server | weekly | 2 d | | `$FLIT_NAME-disk-daily` | `bin/check-disk.sh` | server | daily | 6 h | | `$FLIT_NAME-node-pin-weekly` | `bin/check-node-pin.sh` | server | weekly | 2 d | | `$FLIT_NAME-project-config-weekly` | `bin/check-project-config.sh` | server | weekly | 2 d | | `$FLIT_NAME-drill-restore-quarterly` | `drill-restore.sh` | **laptop** | quarterly | 7 d | | `$FLIT_NAME-drill-rebuild-semiannual` | `drill-rebuild.sh` | **laptop** | twice yearly | 14 d | All server-side jobs run through `bin/job-wrapper.sh` (below) as `User=$FLIT_USER`. The last five rows include what would otherwise be event-driven alerts. Disk usage and the Node pin comparison run on timers, so they are checks like anything else: ping `/0` when healthy, `/1` when not. The project-config check starts Claude Code in a throwaway project whose `.claude` is a relative symlink and whose `SessionStart` hook writes a marker. **The marker is the evidence, not the model response:** hooks run before inference, so once the marker exists, a later quota, session-limit, or API failure is irrelevant to whether Claude Code followed the symlink and loaded the config. The probe disables persistence and tools and sets a negligible budget; it fails only when the hook did not fire (distinguishing a pre-hook launch failure from a session that ran while silently ignoring the configuration). Routing them this way rather than curling ntfy directly buys three things: 1. **One code path.** No event-driven branch to maintain alongside the scheduled one. 2. **The threshold checkers become dead-man-covered themselves.** A disk-check timer that dies would otherwise be invisible — the only alert being watched for is one that fires solely on bad news. 3. **No ntfy credential on the server.** ntfy is configured as a healthchecks *integration*, so the token lives at healthchecks.io and never touches the box — which is why it is absent from the §7.3 residue list. #### Where the alerting work runs Split along the credential boundary, not by convenience: | Work | Runs on | Why | | --- | --- | --- | | Create and upsert checks, verify bindings, delivery test | **Laptop** — `laptop/healthchecks.sh`, called by `bootstrap.sh` phase 8 | Needs the healthchecks **read-write** API key, which stays in the vault and never reaches the server (§5.3, §7.1) | | Install timers, the job wrapper, and the job scripts | **Server** — `70-timers.sh`, run by `bootstrap.sh`'s `run_server_lib()` (phase 10) | Needs only the ping URL from `/etc/$FLIT_NAME/env` | Putting check creation in `server/lib/` would place a script needing a laptop-only credential on the machine that cannot obtain it — the recurring mistake this spec guards against. The server half must work with no healthchecks API access at all. #### Creating the checks **Use the Management API v3 in `bootstrap.sh`, not auto-provisioning.** Ping-URL auto-provisioning (`?create=1`) cannot set period or grace — it defaults every check to a 1-day period and 1-hour grace. That is right for the daily jobs and actively wrong for the other five, which would alert continuously. Phase 8 runs four steps, in order. Steps 1 and 4 are the ones that matter. **1. Assert the integrations exist — before creating anything.** ```bash curl -sf -H "X-Api-Key: $HC_API_KEY" https://healthchecks.io/api/v3/channels/ # → {"channels":[{"id":"…","name":"…","kind":"ntfy"}, …]} ``` Require at least one channel with `kind: "ntfy"` and one with `kind: "email"`. If either is missing, **abort before creating checks** with a message naming which one and where to add it. Also check headroom: the free tier allows 20 checks, this design uses 8 plus 1 for the delivery test — the read-data verification folded into the existing weekly check rather than adding a tenth (§6.2) — and the account may already hold others. Creation past the limit returns **403**, which is easy to misread as an authentication problem. Count first and fail with an accurate message. This is the automation that matters. Integrations cannot be created via the API — `channels/` is GET-only — so this step cannot install them. What it *can* do is convert a silent misconfiguration into a loud precondition failure. Checks created with no channel attached fire into nothing, which looks exactly like checks that never fire. **2. Create each check with explicit channel UUIDs**, taken from step 1 — not `channels: "*"`. The wildcard silently binds whatever happens to exist, including nothing. Use `slug` plus `unique: ["slug"]` for upsert semantics, so re-running `bootstrap.sh` updates the existing checks instead of creating duplicates — required by §5.3. ```json { "name": "$FLIT_NAME backup daily", "slug": "$FLIT_NAME-backup-daily", "schedule": "0 3 * * *", "tz": "Europe/Berlin", "grace": 21600, "channels": ",", "unique": ["slug"] } ``` `schedule` accepts cron or systemd `OnCalendar` expressions — use the same expression as the systemd timer that drives the job, so the two cannot drift apart. **3. Read each check back and assert `channels` is non-empty.** Cheap, and it catches a typo in a UUID that step 2 would otherwise accept. **4. Prove delivery end to end.** Ping a scratch check `/fail`, wait for the alert, then `/0`. Pause and ask the operator to confirm **both** a push notification and an email arrived, then delete the scratch check. Use a **fixed slug** (`$FLIT_NAME-delivery-test`) with `unique: ["slug"]` so repeated builds reuse one check rather than accumulating them, and delete it from a trap so a failure midway does not leave it behind. Skipped under `--non-interactive`, which has no operator to confirm. **But it cannot be skipped forever.** Record the date it last passed in the vault, and have `bootstrap.sh` refuse to complete an interactive run if that record is missing or older than a year. Writing that record needs the operator's authenticated `pass-cli` session, not the read-only access token — the same constraint as escrow (§6.2), and the reason this step lives in `laptop/healthchecks.sh` rather than anywhere on the server. Otherwise the entire notification chain — the thing every other alert depends on — could go permanently unproven simply by always passing `--non-interactive`. Nothing before this step proves the chain works — it only proves the pieces are configured. This is the same principle as the §8 drills: an untested notification path is not a notification path. > **Note on API keys.** Keys are project-specific; there are no account-wide keys. The read-only > key omits the `channels` field from responses, so all of the above needs the read-write key — > which is fine, since it runs in `bootstrap.sh` while the operator is present and stays in the > vault. **Do not add a channel-binding assertion to the weekly `verify.sh` run**, which is > unattended and would need that key on disk. Step 4 is what validates this, once. #### Job wrapper Every scheduled job follows this shape: ```bash log=$(mktemp); trap 'rm -f "$log"' EXIT curl -fsS -m 10 --retry 5 "$HC_URL/$SLUG/start" || true restic backup ... > "$log" 2>&1; rc=$? curl -fsS -m 10 --retry 5 --data-binary @"$log" "$HC_URL/$SLUG/$rc" || true ``` Two details that are easy to get wrong and both fail quietly: - **`|| true` on every ping.** Under `set -e`, a failed `/start` ping aborts the script *before the backup runs* — a network blip would silently skip that night's backup. Monitoring must never be able to prevent the work it monitors. - **`mktemp`, not a fixed path.** All eight jobs share this wrapper, and the daily backup and daily disk check can overlap. A shared `/tmp/last.log` means one clobbers the other and the wrong output is posted to healthchecks. Three things this buys beyond a bare ping: - **Ping the exit code, not a fixed success URL.** A job that fails but pings `/0` looks healthy forever. Passing `$?` makes failures register as failures. - **`/start` measures duration** and lets healthchecks flag a job that overruns its grace — a backup quietly growing from 4 minutes to 40 is a signal otherwise missed. - **The POST body is retained** (first 100 kB), so failure output is waiting in the dashboard rather than needing to be reproduced. #### Two delivery channels, both attached to every check - **ntfy** — push, with its own sound. The channel that actually gets noticed, and the reason a DOWN on the daily backup does not sit unread for three days while travelling. - **Email** — the backstop. Slow and easy to ignore, which are the right properties for a second channel and the wrong ones for a primary. Without it, an ntfy outage produces silence indistinguishable from health. Use a high-entropy ntfy topic name: public topics are readable by anyone who guesses them, and these messages carry hostnames, disk state, and backup outcomes. **Do not self-host either service.** A watchdog on the machine it watches reports nothing in the failure it exists to catch, and a second box for monitoring defeats the single-box design. **Where the layering stops:** healthchecks.io itself going down, unnoticed. The §8.1 quarterly restore drill is the backstop for the backstop. --- ## 7. Secrets ### 7.1 Runtime injection — Proton Pass CLI No **application** secret values on disk — the five §7.3 items are the documented exception, and this section is about everything else. `.env` files hold `pass://vault/item/field` references and are safe to commit: ``` DB_HOST=localhost DB_PASSWORD=pass://Dev/Database/password API_KEY=pass://Dev/External/api_key ``` Invoke via: ```bash pass-cli run --env-file .env -- pnpm dev ``` `pass-cli run` resolves the references, injects real values into the child process environment, masks secret values in stdout/stderr by default, and forwards stdin/stdout/stderr and SIGTERM/SIGINT to the child. Do not pass `--no-masking`; masking keeps leaked values out of terminal output that Claude Code may read. The same mechanism wraps `bootstrap.sh` on the laptop (§5.1), so infrastructure credentials follow the same path as application secrets. **Every reference names a field** — `pass://$FLIT_NAME//password`, never a bare item path. The vault has no default field to fall back to, so an item-only reference is invalid. The vault itself is named `$FLIT_NAME` (`VAULT_NAME` in `lib/common.sh`, no separate override). **`.env` may only reference items that exist before `bootstrap.sh` starts.** `pass-cli run --env-file` resolves every reference up front and fails the whole run on the first missing one, so nothing `bootstrap.sh` itself creates can be named there. The Storage Box credentials, the restic repository URL, and the restic password are all created during the run (phase 9) and are read with `vault_get` at the point of use instead — see the standing rule in §13. **A second deployment needs its own `.env`.** `pass-cli run --env-file` reads each `pass://` prefix literally — no shell expands it — so overriding `FLIT_NAME` to stand up a second deployment does not redirect those prefixes: a `.env` still reading `pass://flit/...` would resolve every credential against the first deployment's live vault instead of the new one, silently, since a mismatched vault name is still a valid one to query. `vault_check_env()` in `lib/common.sh` turns that mistake into an immediate failure — it dies naming the offending reference, the vault it names, and `VAULT_NAME` — rather than a silent cross-deployment read. `bootstrap.sh`'s `preflight()`, both drill scripts (via `drill_precondition()` in `lib/drill-common.sh`), and `laptop/healthchecks.sh` all call it before touching the vault, since all four are documented as running under `pass-cli run --env-file .env`. **Scripted vault access goes through six functions in `lib/common.sh`**: `vault_session_ok`, `vault_has`, `vault_get`, `vault_put`, `vault_ensure`, `vault_check_env`. Nothing outside that file calls `pass-cli` directly. For config files that need literal values on disk, use `pass-cli inject -i tpl -o out`, and gitignore the output. Caching is via the kernel keyring — roughly 5–6 s on first fetch, ~0.01 s subsequently, expiring after an hour or at logout. No per-command latency concern. #### Vault inventory Every item `bootstrap.sh` touches, and in which direction. Read items must exist before the build starts; written items are created during it. | Item | Direction | Notes | | --- | --- | --- | | `pass-cli` access token | — | The bootstrap credential itself; minted read-only by `laptop/prepare-vault.sh` via `pass-cli personal-access-token create` and scoped with `access grant --role viewer`, 90-day expiry | | Hetzner API token | read | Laptop only, never on the server (§7.3) | | Tailscale auth key — **persistent** | read | For the build; the node must not expire (§4.2). Mark it **Reusable** — the console default is single-use, and a single-use key works once, then fails silently on the next rebuild | | Tailscale auth key — **ephemeral** | read | For drill instances, so they self-remove (§8). **A second, distinct key**, with the Ephemeral toggle set — do not reuse the persistent one, or drill nodes linger in the tailnet. Also mark it **Reusable**, or the second drill fails | | Storage Box username, hostname, password | **write** | Generated by `laptop/provision-storagebox.sh`, escrowed immediately, then the password is read back for `ssh-copy-id` in the same phase (§6.2) | | restic repository URL | **write** | Derived from the Storage Box username and hostname, phase 9 (§6.2) | | healthchecks read-write API key | read | Laptop only; this is why remediation runs from the laptop (§5.3) | | healthchecks ping key | read | Written on to `/etc/$FLIT_NAME/env` by `05-flit-env.sh`, phase 4 | | restic repository password | **write** | Generated phase 9, escrowed (§6.2) | | Storage Box private key | **write** | Generated phase 9, escrowed (§6.2) | | `restic-password-next`, `storagebox-key-next` | **write** | Exist only mid-rotation, under `--replace` (§6.2); each is promoted over the item above it and then left behind stale — no script here deletes them | | `$FLIT_USER` console password | **write** | Generated phase 1, escrowed (§4.2) | | Delivery-test last-passed date | **write** | Recorded by `laptop/healthchecks.sh` (§6.3); read on every interactive run to enforce the one-year limit | | Operator ssh private key (`operator-key`) | **write** | Generated by `laptop/prepare-vault.sh`'s `store_operator_key()`, escrowed in `$OPERATOR_VAULT_NAME` (default `$VAULT_NAME-operator`) — **the one item in this table living in a vault of its own**, not `VAULT_NAME`. The access token above is scoped to `VAULT_NAME` alone, so the server that reads every other row here is denied this one; the credential that opens the box must not be readable by the box. `resolve_operator_key()` restores it from there onto a laptop that has none | Reads use the access token. **Writes require the operator's authenticated session** — see the escrow note in §6.2. ### 7.2 Git authentication — Tangled Tangled authenticates git over SSH using public keys propagated through atproto: the operator's pubkey is written to their AT repo, and the signed record syncs via the firehose to knots they belong to. Standard SSH pubkey auth from the client's perspective. **The box carries its own key**, generated on the laptop rather than reused from the operator's. `laptop/tangled-key.sh` creates `~/.ssh/tangled_ed25519`, escrows it to the deployment vault as `tangled-key` and verifies the escrow by reading it back, and only then pushes the key to the server — the same order `laptop/backup-credentials.sh` uses for the Storage Box key (§6.2), so a rebuild reproduces the key from the vault instead of needing its public half registered by hand a second time. `bootstrap.sh` `phase_7_sshconfig()` installs the key before writing the `~/.ssh/config` block that names it, and skips the key entirely under `--no-external-state` — registering a public key against the Tangled account is state shared with production, which a drill must not touch (§8.2). A forwarded agent has no part in this. Pushes authenticate with the key on disk and work identically whether or not the operator is connected. Add to the server's `~/.ssh/config`: ``` Host tangled.org tangled.sh Hostname tangled.org User git AddressFamily inet IdentityFile ~/.ssh/tangled_ed25519 IdentitiesOnly yes ``` `AddressFamily inet` is per Tangled's documented configuration. The block matches both `tangled.org` and `tangled.sh` because remotes across the projects on this box are written both ways. `IdentitiesOnly yes` restricts SSH to the identity named here, so nothing else offered to the box — an agent forwarded in for an unrelated reason, another key — can answer in its place. This block applies to `tangled.org`-hosted knots. Self-hosted knots need their own configuration — out of scope here. **Account-level, and scoped accordingly.** A Tangled key authenticates for the whole account, not one repository, so this key can push to every repository the account owns — the same reach a compromised operator key would have. It is a distinct keypair, independently revocable, for exactly that reason: this box runs Claude Code against untrusted repositories (§5.7), and a compromised box should cost one key, not the operator's own identity. **What cannot be automated.** There is no API to register a public key against a Tangled account. `laptop/tangled-key.sh` ends by deferring that one step — the key text to paste — the same deferral convention §5.8 uses everywhere else a change cannot be applied surgically. A key that exists on disk but was never registered fails exactly like a missing one, which is why `verify.sh` (§5.3) tests the push itself rather than the file's presence. ### 7.3 Irreducible residue **Five items** must exist on disk. This is a real limit, not an oversight: | Item | Why it cannot be fetched | Value if leaked | | --- | --- | --- | | `pass-cli` access token | Bootstrap credential for everything else; the server needs it to resolve `pass://` references daily, and for `verify.sh` to confirm the escrow | Read access to one vault, expires in 90 days. **Installed by `laptop/backup-credentials.sh`** — nothing else in the build path puts it there, and without it the escrow assertion SKIPs permanently while runtime secret injection silently has no credential. | | restic repository password | The backup timer runs while the operator is disconnected | Useless without the Storage Box key. **Escrowed to Proton Pass at generation (§6.2)** — the on-disk copy is a cache, not the only copy. | | Storage Box SSH key | Same — restic needs it unattended at 03:00 | Write access to the backup target only; dedicated keypair, used nowhere else. **Also escrowed (§6.2).** | | `~/.claude` OAuth token | Written by the tool itself | Revocable from the console. **The one credential with no expiry monitoring** — `claude --version` exits zero without valid auth, so neither `verify.sh` nor the drills detect a lapse. It surfaces the next time Claude Code is actually used, which is acceptable: the failure is immediate, obvious, and costs one re-authentication. | | healthchecks.io ping key | Scheduled jobs ping it unattended (§6.3) | Integrity only — an attacker could forge healthy pings, but reads nothing. Not a secret in the confidentiality sense. | > **Do not route the backup credentials through `pass-cli` to shorten this list.** The > repository password and the Storage Box key are both needed by a timer that fires while the > operator is disconnected. Making them depend on the vault token means backups stop silently on > day 91 and the failure surfaces during a restore — the worst possible moment. **Backups must > not depend on the secrets manager.** Keeping them decoupled confines token expiry to > interactive work, where it surfaces the same morning. #### Where they live One file, one convention. Do not scatter these: ``` /etc/$FLIT_NAME/env root:$FLIT_USER 0640 # config + HC_URL (embeds the ping key) /etc/$FLIT_NAME/secrets.env root:root 0600 # restic password, pass-cli token /home/$FLIT_USER/.ssh/storagebox_ed25519 $FLIT_USER:$FLIT_USER 0600 ``` `env` is written by `05-flit-env.sh` in phase 4 — earlier than the credentials in `secrets.env`, which phase 9 generates — so that anything installed in between has the file it needs already in place; a systemd unit that referenced `env` before phase 9 wrote it was the defect that moved this write earlier (§12.1). The two modes differ deliberately. `env` is `0640 root:$FLIT_USER` because the jobs run as `$FLIT_USER` and read it directly; `secrets.env` stays `0600 root:root` and reaches them only through systemd's `EnvironmentFile=`, which is read as root before privileges drop. > **A job script must not `source` `secrets.env` unconditionally.** Run by hand as `$FLIT_USER` > that is a permission error. Source it only when readable and otherwise rely on the unit. systemd units load them with `EnvironmentFile=`, so values never appear in unit files, in `ps` output, or in shell history. **Units must also run the job under a login shell**: a systemd environment sources no shell profile, so `mise`, `pnpm`, `cargo`, `rustup` and `claude` are all absent from `PATH`, and every scheduled job needs at least one of them. Without it the jobs fail with *command not found* — reported as a genuine failure, from a perfectly healthy box. ```ini [Service] EnvironmentFile=/etc/$FLIT_NAME/env EnvironmentFile=/etc/$FLIT_NAME/secrets.env ``` `verify.sh` asserts the modes and ownership above. A world-readable `secrets.env` is the kind of thing that happens once during a hurried fix and is never noticed again. Three things deliberately **not** on this list: - **The Hetzner API token.** It stays on the laptop (§5.1). This is why the drills are laptop-initiated rather than server timers — see §8. A token that can create and destroy infrastructure has far more blast radius than anything above, and keeping it off the box is worth the loss of a server-side drill schedule. - **The ntfy token.** It lives at healthchecks.io as an integration, never on the server. This is the direct payoff of routing everything through healthchecks rather than publishing to ntfy from the box (§6.3). - **The healthchecks Management API key.** Used only by `bootstrap.sh` while the operator is present, so it stays in the vault. Mitigations: - Scope the Proton Pass access token **read-only to a single vault**, expiring at **90 days**, so expiry enforces rotation. - Each is independently revocable and individually low-value. **No LUKS volume, and no separate volume for this residue.** Both rejected — see §10. --- ## 8. Drills Two drills that answer different questions. Running only the first is the common mistake — it passes happily while the provisioning scripts are quietly incomplete. | Drill | Question | Cadence | | --- | --- | --- | | **8.1 Restore** | Can the work be recovered? | Quarterly | | **8.2 Rebuild** | Do the scripts reproduce the box? | Twice yearly | **Drill instances use ephemeral Tailscale auth keys and unique hostnames** — `drill-restore-` and `drill-rebuild-`. Both drills join the tailnet to reach the live server (§8.2 diffs its package list), and a drill node registering as `$FLIT_SERVER` would give MagicDNS two answers for the same name. The operator's `ssh $FLIT_SERVER` could then land on a throwaway instance minutes from destruction. Ephemeral keys make Tailscale remove the node automatically when it goes offline, so a failed drill cannot leave a stale `$FLIT_SERVER` in the tailnet. #### Check the token before spending anything **Step zero of both drills: assert the `pass-cli` token is valid and has more than a few days left.** If not, stop immediately and tell the operator to regenerate it. This is not defensive padding — the numbers collide by construction. The token expires at **90 days** (§7.3); the restore drill runs **quarterly**, about 91. So the drill will typically fire just *after* expiry. Without a precondition check it fails at step 3, several minutes and one provisioned CX53 into the run, and reports as a drill failure — which reads like a broken restore path rather than routine credential expiry. Both numbers are individually sensible. They were chosen in different sections without reference to each other, which is what makes this worth an explicit guard rather than a tweak to one of them. **Both run on CX53 — identical to production.** Hourly billing puts a one-hour drill at roughly four cents, so there is no reason to economise here. A drill on a smaller instance introduces size as a variable: §8.2 asserts `cargo build` succeeds, and §4.1 argues at length that Rust link steps need the memory a CX23 does not have. A drill that passes on hardware production does not run proves less than it appears to. Both must destroy the test instance on **all** paths including failure, so a failed drill does not leave a billing instance running. #### Drills run from the laptop, not from a server timer **This is a correctness requirement, not a preference.** A drill provisions an instance, so it needs the Hetzner API token. Putting the drills on a server-side timer would put that token on the server — contradicting §5.1's central claim that it never leaves the operator's machine, and adding a sixth item to §7.3 with far more blast radius than the other five: it can create and destroy infrastructure. So both drills run on the laptop, under `pass-cli run`, like `bootstrap.sh`. They also need the Storage Box key and restic password, which the drill pulls from the live server over the tailnet at run time rather than keeping local copies. **Dead-man coverage is preserved by inverting who is being watched.** The healthchecks checks (§6.3) still exist; the ping simply comes from the laptop-run drill rather than a server timer. If a quarter passes without a drill, the check goes DOWN and alerts — a dead-man switch on the **operator** instead of the server. Drilling still cannot be silently stopped, and no infrastructure credential lands on the box. A drill that fails silently is not a drill; neither is one quietly stopped. #### Assert against fixtures, not against `~/workspace` Build assertions on `fixtures/` in the `$FLIT_NAME` repo — a minimal pnpm workspace and a minimal Rust crate, both committed. `~/workspace` is the wrong target for two reasons. At build time it is empty, so phase 13's drills would vacuously pass. And once it has real repos, the assertions become non-deterministic: a drill would start failing because a real project broke, not because the restore did. Restored `~/workspace` content is still checked — but for **presence**, not buildability: assert a known file from the latest snapshot exists with the expected content. That is what proves the restore worked. Buildability is the fixtures' job. ### 8.1 Restore drill — `drill-restore.sh` **Not optional, and not deferrable.** Run before this build is considered complete. 1. Create a CX53 from the most recent automatic backup — or from the golden snapshot when run with `--from-snapshot`, which is required at build time, before any automatic backup exists (§12.1). 2. Make the whole restic backup set writable — `~/workspace`, `~/.claude`, `~/.bash_history`, `~/.vim/undo` and the small state outside them (credentials, cron scripts, directories with no remote), the same paths `server/bin/backup.sh` passes to `restic backup` (`lib/common.test.sh` guards the two lists against drifting apart) — then restore the restic snapshot onto it. git writes its object files mode 0444, and a restore onto a host whose image already contains them must overwrite those files — which fails on a read-only file even for its owner — so write permission is restored first. A path absent from the image (`~/.vim/undo`, in particular) is skipped rather than failed. 3. Supply the `pass-cli` access token — it is §7.3 residue and is **not** in the restic backup set, so a restored host has none. The drill injects it from the vault session it is already running under (`pass-cli run`), which works because drills are laptop-initiated. It also supplies the current Storage Box key from the same vault session, overwriting whatever copy the image carries — an image freezes the key that was current when it was taken, and a rotation since then retires that copy on the box (standing rule 14). 4. Assert: expected file present in `~/workspace` with expected content (restore worked); `pnpm install` and `cargo build` succeed **in `fixtures/`**; `pass-cli run` resolves a secret; `claude --version` exits zero (its token restores with `~/.claude`). 5. Report wall-clock time to working state, and ping `$FLIT_NAME-drill-restore-quarterly` with the exit code (§6.3). 6. Destroy. An untested restore path is not a backup. ### 8.2 Rebuild drill — `drill-rebuild.sh` This is the drill that catches the drift `verify.sh` cannot (§5.3). The restore drill starts from a snapshot that *contains* the drift, so it cannot detect an incomplete `server/lib/` script. 1. Create a CX53 from a **stock Ubuntu 24.04 image** — not a snapshot, not a backup. 2. Run `bootstrap.sh --target --skip-laptop --skip-drills --non-interactive --no-external-state` against it. All four are required: omitting `--skip-drills` recurses infinitely, `--non-interactive` hangs at the first prompt, and `--no-external-state` contaminates production monitoring and backups (§12.1). 3. Restore **data only** from restic — `~/workspace` and `~/.claude`, no system state. 4. Supply the `pass-cli` token as in §8.1 step 3. 5. Assert the same checks as §8.1, plus `verify.sh` passing — with one expected exception. `--no-external-state` skips installing the box's own Tangled key (§7.2), so `verify.sh`'s two Tangled assertions (key mode, push) FAIL against a key that was never written. That failure is specific to the flag, not evidence the rebuild diverged from the live server; assert every other `verify.sh` check passes and tolerate exactly those two. 6. Report wall-clock time, diff the installed package list against the live server, and ping `$FLIT_NAME-drill-rebuild-semiannual` with the exit code — **non-empty diff means nonzero** (§6.3). 7. Destroy. Compare **`apt-mark showmanual`**, not `dpkg -l`. A full package list differs between two hosts for reasons nobody controls — kernel metapackages, recommends pulled at different times — so the diff is noisy, and a noisy diff gets an ignore-list bolted on until it reports nothing. Manually installed packages are exactly the set the `server/lib/` scripts are responsible for. **Any package present on the live server but absent here is uncommitted drift.** Add it to whichever `server/lib/` script installs packages for that area — which one depends on the drift — and commit. That diff is the entire point of this drill; do not skip step 6. A non-empty diff is a **failure**, alerted via ntfy — not a report to note and move on from. Treating it as informational is how this drill becomes ceremony. ### 8.3 Real recovery procedure The drills inject the `pass-cli` token from the operator's live session. **A real disaster has no live session**, so this is the part the drills do not rehearse. Written out because it has to be executable from a phone, with nothing but access to Proton Pass. Restoring the machine restores `~/workspace` and `~/.claude`. It does **not** restore §7.3 residue — that is by design (§3), and it means four things must be re-supplied before the box works: 0. **Provision a replacement machine first.** These steps operate on a host that exists. Either create one from the most recent Hetzner backup, or run `bootstrap.sh` against a stock image. Both need the Hetzner API token, which needs step 1 — so do step 1 before this one. Whichever path is taken, make the whole restic backup set writable (the paths in `server/bin/backup.sh`, starting with `~/workspace` and `~/.claude`) before any restic restore run against it — git's object files are mode 0444, restoring onto a host that already holds them requires overwriting them, and opening a read-only file for writing fails even for its owner (§8.1). 1. **Create a new `pass-cli` access token** in the Proton Pass UI, scoped read-only to the vault. Everything else depends on this. 2. **Write `/etc/$FLIT_NAME/secrets.env`** — restic repository password and the new token, mode `0600`, root-owned (§7.3). 3. **Restore the Storage Box SSH key** from Proton Pass to `~/.ssh/storagebox_ed25519`, mode `0600`. It is there because phase 9 escrowed it at generation (§6.2) — along with the restic password used in step 2. Without both, the backups cannot be decrypted at all. Read `restic-password` and `storagebox-key`, never a `-next` sibling — those exist only mid-rotation (§6.2, §7.1) and are not guaranteed to be the value the repository or the box currently accepts. 4. **Re-authenticate Claude Code** if `~/.claude` did not restore cleanly, and **write the healthchecks ping URL** into `/etc/$FLIT_NAME/env` — done by `05-flit-env.sh` if step 0 ran `bootstrap.sh`; write it by hand only when recovering without it. Then run `verify.sh`. It fails on anything missed. **Order matters at step 1**: without the vault token, `bootstrap.sh` cannot resolve the Hetzner API token, and the replacement instance cannot be provisioned at all. If the Proton Pass account itself is unreachable, nothing else in this design recovers — that is the single point of failure and it is accepted knowingly. > **This procedure becomes the primary path on a move to dedicated hardware (§11).** There are > no snapshots there, so restic is the only route back. Re-read this section before that > migration, not during it. --- ## 9. Claude Code - Interactive use only. Default permission prompts stay on. - Do **not** pass `--dangerously-skip-permissions`, and do not configure auto mode. - Do not build automation that runs `claude -p` on a timer. The operator being present at the keyboard is the security control that lets the rest of this design stay simple. A malicious `CLAUDE.md`, README, or source comment in a cloned repository can attempt to instruct the model; with default permissions this surfaces as a prompt to decline rather than a silent execution. If unattended operation is wanted later, it is a **re-architecture, not a config change**: container isolation, network allowlist, memory limits, scoped push credential, review logging. Reopen this spec. --- ## 10. Rejected alternatives Recorded so they are not helpfully reintroduced. | Rejected | Reason | | --- | --- | | **Docker for the dev environment** | Three reasons, none of them drift. (1) Its main value here was blast-radius containment for unattended agents, which left scope — see §9. (2) One operator, one box: the isolation and portability Docker buys are for problems this setup does not have. (3) It costs the §5.5 pnpm hardlink care and UID mapping *every day*, in exchange for reproducibility a snapshot plus the `server/lib/` scripts already provide. **Note for anyone revisiting:** do not re-reject this on drift grounds. Docker drifts too, but visibly — changes live in a writable layer and vanish on rebuild, so breakage is loud and early. A provisioned host hides drift until rebuild day. On that axis Docker is arguably better, and the honest case against it is the three reasons above. | | **Widening the `pass-cli` token to read-write** | The obvious "fix" when escrow fails with a read-only token. It would turn the on-disk §7.3 residue into a credential that can rewrite the whole vault — a far worse exposure than one interactive step at build time. Escrow uses the operator's authenticated session instead (§6.2). | | **Sharing Terraform state with the drills** | A drill's destroy step could remove the production entry, and concurrent runs corrupt state. Drills use `hcloud` directly and stay outside managed infrastructure (§5.1). | | **Pull-based provisioning** | Requires either a public configuration repository or a bootstrap credential on the server to clone it. Push from the laptop has no bootstrap cycle. | | **LUKS-encrypted volume** | Protects a block device at rest. On a cloud instance whose disk the operator never physically handles, it mostly adds an unlock step and a failure mode where a reboot leaves services quietly broken. | | **Separate un-backed-up volume for §7.3 residue** | The recovery flow is *restore, then re-authenticate*, so credentials in an old snapshot are inert by design — and a 90-day token expiry means snapshots older than a quarter contain a dead token. Not worth the mount ordering and the failure mode where a restored box looks broken because the volume wasn't attached. | | **Self-hosting ntfy or healthchecks** | A watchdog on the machine it watches reports nothing in the failure it exists to catch. A second box for monitoring defeats the single-box design. | | **Publishing to ntfy directly from the server** | ntfy is a transport and cannot detect a timer that never fired — the failure the alerting layer exists to catch. It also puts an ntfy token on the box for no benefit. Everything pings healthchecks; healthchecks fans out to ntfy and email. | | **Dropping ntfy and using email alone** | Workable — healthchecks emails natively — but a DOWN on the daily backup would sit unread for days while travelling. ntfy costs one UI step at setup and no server-side credential, so it is kept as a delivery channel. | | **healthchecks auto-provisioning (`?create=1`)** | Cannot set period or grace; defaults every check to 1-day period and 1-hour grace. Correct only for the daily jobs, and would make the weekly, monthly, quarterly, and semiannual checks alert constantly. Use the Management API. | | **`channels: "*"` when creating checks** | Silently binds whatever integrations happen to exist — including none. Look up UUIDs via `GET /channels/` and pass them explicitly, so a missing integration is an error rather than a check that alerts into nothing. | | **mosh, in any role** | Removed entirely after tracing what it cost. It needs a UTF-8 locale on the server and UDP 60000–61000 open, breaks SSH agent forwarding (so `git push` failed in any mosh session), has no scrollback, and forced a `verify.sh` assertion — all to improve typing latency on occasional long-haul trips. Its removal also simplified two other decisions: agent forwarding now works everywhere (§7.2), and the renderer choice rests on the clipboard rather than on bandwidth (§5.7). The accepted cost is that `vim` from Asia feels the ~180 ms round trip per keystroke. **Do not reintroduce it to fix that**; reopen §11's dedicated-hardware question or accept the latency. | | **1Password** | Equivalent capability; operator already pays for Proton. `pass-cli run` matches `op run`. | | **`fnm` for Node** | Node-only. The box is polyglot; `mise` covers future runtimes at the same install cost. | | **Deploy key or token on the server** | Only needed for unattended pushes. Agent forwarding is strictly better: no credential on the box. | | **Hetzner Object Storage for backups** | Both are API-provisionable now that Storage Box has moved into the Cloud console (§6.2), so the choice is pricing, not automation. Object Storage bills per GB plus egress, so a restore costs money exactly when things are already going wrong; Storage Box is fixed-price with free restores. | | **Second off-provider backup target** | Repos live on a `tangled.org` knot, already off Hetzner. Residual exposure is uncommitted work only. Revisit if that stops being true. | | **ARM (CAX) instances** | CAX31 now prices above CX43 for identical specs after the 2026 increases. No remaining advantage. | | **Hetzner dedicated / server auction** | **Deferred, not rejected — see §11.** The original objection (no snapshots, operator owns recovery) is substantially weakened by `drill-rebuild.sh` existing. Do not treat this row as settled. | | **Mobile targets (iOS, Android)** | Out of scope. iOS is impossible on Linux regardless of sizing — Xcode is required and macOS-only. Android was buildable but the emulator needs `/dev/kvm`, which Hetzner Cloud does not expose. Removing both is what keeps this a single cloud instance. | | **Two servers (EU + Asia)** | State divergence — which box has the uncommitted branch — costs more than the latency. One box, and the Asia round trip is felt in `vim` (§4.3). | --- ## 11. Deferred Recorded so the design does not foreclose it. **Not in scope for this build.** **macOS Electron builds.** electron-builder produces unsigned macOS artifacts on Linux; signing and notarization normally require a Mac. `rcodesign` (the `apple-codesign` Rust crate) can do both from Linux, but it is a real project rather than a configuration step and is sensitive to Apple's periodic changes. **Intended shape:** built artifacts land in a directory on the server that syncs to the operator's local machine. The sync mechanism is not chosen yet. Nothing in this spec blocks it. When picked up, expect work in three places: cross-compilation targets via `rustup` (§5.4), Apple Developer credentials added to the Proton Pass vault (§7.1, not §7.3 — they are fetchable on demand), and disk headroom for additional Electron platform binaries (§4.1). **Migration to dedicated hardware (Hetzner AX line or server auction).** The reasoning that rejected this has changed and should not be read from §10 alone. The original objection was that dedicated has no snapshots, so the operator owns recovery. **`drill-rebuild.sh` largely answers that**: once a from-stock-image rebuild is tested twice yearly, losing snapshots costs time-to-recovery, not data. What it would buy, specifically for this workload: - **Much better single-core clock** — which governs `tsc` and Rust link steps, the things actually waited on. Shared vCPU is the current ceiling. - **2×1 TB NVMe instead of 320 GB** — retiring §4.1's "disk is the resource to watch" entirely, rather than managing it with prune timers and an 80% alert. What it still costs: - **No snapshots**, so §6.1's one-click whole-machine rollback disappears and restic becomes the only recovery path. §8.1 stops being a drill and becomes the actual procedure. - **Provisioning moves off the `hcloud` Cloud API entirely**, to whatever ordering flow Hetzner runs for dedicated boxes, so `infra/` is rewritten and phase 1 changes shape. - **Hardware failure is the operator's problem**, including the wait for a replacement. **Trigger: revisit after two successful `drill-rebuild.sh` runs**, i.e. roughly a year in. At that point the rebuild path is evidenced rather than assumed, and the move is an afternoon rather than a leap. --- ## 12. Execution ### 12.1 Automated — `bootstrap.sh` Run from the laptop, wrapped in `pass-cli run`. Each phase is idempotent; the whole script is safe to re-run. **Required flags:** | Flag | Effect | | --- | --- | | `--target HOST` | Skip phase 1; provision against an existing host. **Required by §8.2**, which creates its own instance before calling `bootstrap.sh`. Without this the rebuild drill is circular. | | `--skip-laptop` | Skip the laptop half of phase 6. Used by drills, which must not touch the operator's machine. | | `--non-interactive` | **Never prompt, anywhere.** Not an enumerated list of prompts — any code path that would wait for input must instead assert, or skip and log. This covers at least the §12.2 manual touches and the §6.3 step 4 delivery confirmation. **Required by both drills**; without it they hang at the first prompt, and restoring `~/.claude` does not help because that happens after bootstrap returns. | | `--no-external-state` | **Required by both drills.** Touch no account-level state shared with production: skip healthchecks check creation and the delivery test, point `HC_URL` at a sink, skip enabling Hetzner Backups, **copy the existing Storage Box key rather than generating and uploading a new one**, and install the restic timer **without enabling it**. Wrapper scripts, unit files, and timers are still written — that is what the drill verifies. See the note below; omitting this corrupts live monitoring, pollutes backups, and leaves write credentials on the backup target. | | `--skip-drills` | Skip phases 12–13. **Required by both drill scripts.** Without it `bootstrap.sh` → `drill-rebuild.sh` → `bootstrap.sh` recurses without terminating. Both drills must pass this, and `bootstrap.sh` must refuse to run drills when `FLIT_IN_DRILL` is set, as a second guard against a caller forgetting the flag. | | `--print-only` | Per §5.8 — mutate nothing, anywhere; laptop config inspected exactly, server phases planned only. Exits 0. | | `--start-at N` | Resume from phase N. Every phase is idempotent, so completed ones are no-ops; this exists because §12.1 leaves a failed build's instance running for inspection. | | `--stop-at N` | Run no phase after N. `--start-at 2 --stop-at 2` delivers server-side code and nothing else — what an edit to a `server/lib/` or `server/bin/` script needs to reach the box, since `laptop/sync-config.sh` then runs the delivered copy of `server/lib/22-claude-config.sh`. Without `--stop-at`, `--start-at 2` continues through every later phase, turning a one-script delivery into a full provisioning run. Refused if it names a phase before `--start-at`, and refused (rather than read as "no limit") if the value is not numeric. Stopping early logs that the box is not fully provisioned, in place of the usual completion line. | | Phase | Actions | Spec | | --- | --- | --- | | **1. Provision** | Create CX53, Ubuntu 24.04, in `${HCLOUD_LOCATION:-fsn1}`. Register `${HCLOUD_SSH_KEY_NAME:-$FLIT_NAME-operator}` from the key `resolve_operator_key()` resolved in preflight if it is not already registered, and select it by name at creation. An existing registration under that name is compared by key material, not trusted by name. Create user `$FLIT_USER` via cloud-init `user_data`. Both must exist before phase 2, or the instance is reachable only by the §4.2 console password — with phase 3 about to close port 22. `ensure_operator_key()` (`lib/common.sh`) does the registration; both `phase_1_provision()` and `drill_create()` call it before creating an instance, so neither path can select a name nothing has registered — registration used to live only on phase 1's server-does-not-exist-yet branch, which a resumed run (`--start-at 2`) or a drill against an existing box never took. | §4.1, §4.4 | | **2. Deliver** | rsync the `$FLIT_NAME` repo **and the dotfiles repo** to `$FLIT_USER@host:~/$FLIT_NAME/`. Prefer the tailnet if `$FLIT_SERVER` already resolves and answers; fall back to public SSH only on a first run. Hard-coding public SSH breaks every resume, since phase 3 closes port 22. `--start-at 2 --stop-at 2` runs this phase alone, for a `server/lib/` or `server/bin/` edit that needs to reach the box without a full provisioning run. | §5.1 | | **3. Network** | Install Tailscale, join tailnet, **assert key expiry is disabled**, then **prove inbound reachability from the laptop over the tailnet**, and only then invoke `40-firewall.sh --tailnet-verified` to close public SSH | §4.2 | | **4. Toolchains** | `05-flit-env.sh` writes the identity block — including `HC_URL` — into `/etc/$FLIT_NAME/env`, ahead of everything downstream that reads it (§7.3); then packages, `mise`, `rustup`, `corepack`, dotfiles | §5.4 | | **5. Caches** | `sccache`, shared `CARGO_TARGET_DIR`, prune timers, disk alert | §5.6 | | **6. Terminal** | Server `~/.tmux.conf`, `CLAUDE_CODE_DISABLE_ALTERNATE_SCREEN`; then `laptop/install.sh` for Ghostty, `~/.ssh/config`, shell helpers | §5.7 | | **7. SSH config** | Server-side Tangled host block | §7.2 | | **8. Alerting** | **Laptop half** (`laptop/healthchecks.sh`): assert both integrations exist and there is headroom under the 20-check limit, abort if not; upsert the eight checks with explicit channel UUIDs from `lib/jobs.sh`; read back and verify bindings; run the delivery test. **Server half** (`70-timers.sh`, phase 10): job wrapper and job scripts, from the same `lib/jobs.sh` table. | §6.3 | | **9. Backups** | **Laptop half** (`laptop/backup-credentials.sh`): generate the restic password and Storage Box keypair, **escrow both and verify the readback**, `ssh-copy-id -s -p 23` the public key, then push the credentials to the server. **Server half** (`60-backup.sh`): `restic init` against the `env` phase 4 already wrote. Then enable Hetzner Backups via the API. | §6.2 | | **10. Timers** | Install server-side timers for the daily backup, weekly `restic check`, daily disk, weekly Node pin, weekly project-config check, and weekly `verify.sh`, from `lib/jobs.sh`. All run `User=$FLIT_USER` (§5.8). **The two drill checks get no server timer** — they are pinged by laptop-run drills (§8); the check still alerts if a quarter passes without one. Use `OnCalendar` expressions and pass the *same* expression to the healthchecks `schedule` field, so the two cannot drift. A check without a timer goes DOWN within a day. | §5.3, §6.3, §8 | | **11. Verify** | Run `verify.sh`; fail loudly on any assertion | §5.3 | | **12. Snapshot** | Snapshot the converged host as the golden image — **temporary, see §6.1**; skipped once an automatic backup of `$FLIT_SERVER` exists | §6.1 | | **13. Drills** | Run `drill-rebuild.sh`, then `drill-restore.sh --from-snapshot` — **from the laptop**, not on the server (§8) | §8 | Phase 3 contains the ordering constraint from §4.2 and must not be reordered. > **Why `--no-external-state` exists.** Phases 8 and 9 reach outside the instance, into accounts > shared with production. Without the flag, a drill would: > > - **Upsert the eight live checks.** `unique: ["slug"]` matches on slug, so a drill edits the > production checks rather than creating its own. > - **Ping production check slugs.** The drill's timers would report `$FLIT_NAME-backup-daily` healthy — > so a drill running while the real backup is broken would actively hide the failure. > - **Write junk into the real restic repository.** A backup timer on a throwaway box would push > snapshots of an empty workspace into the production repo. > - **Enable paid Hetzner Backups** on an instance minutes from destruction. > - **Leave a write-capable key on the Storage Box.** Phase 9 generates a keypair and > `ssh-copy-id`s it. Run in a drill, that adds a new credential with write access to > production backups, and nothing removes it when the instance is destroyed — four drills a > year, four orphaned keys, accumulating indefinitely. Drills copy the live server's existing > key instead. > > Namespacing drill slugs was the alternative and is worse: eight extra checks per drill blows > the 20-check free tier, and they all go DOWN once the instance is destroyed. **Snapshot precedes drills, deliberately.** Enabling Hetzner Backups in phase 9 does not create one — Hetzner does that on its own schedule, so no backup exists minutes later at build time. The phase 12 golden snapshot is what phase 13's restore drill restores from, via `--from-snapshot`. Quarterly runs thereafter use the most recent automatic backup, which is the real thing being tested. Running the drills before the snapshot fails on an empty backup list. **Failure semantics.** On any phase failing, `bootstrap.sh` **leaves the instance running** and exits nonzero, naming the phase. Do not destroy on failure: the instance is the only evidence of what went wrong, and it costs pennies per hour. Re-running resumes — every phase is idempotent, so completed ones are no-ops. Destroying is the operator's explicit call. The drills are the exception (§8): they destroy on all paths, because they create instances nobody is attached to. **A phase's exit status was its last statement's, not its work's.** `set -e` is deliberately not in force, so a phase that never checked one of its own commands reported whatever ran last instead of what mattered. `phase_1_provision()` did not check `hcloud server create`; its status was that of the trailing `rm -f /tmp/$FLIT_NAME-cloud-init.yaml`, which always succeeds. `phase_2_deliver()` did not check `rsync`; its status was `HOST=$dest`, an assignment, also always zero. `phase_3_network()` did not check `30-tailscale.sh`. All three now die on the failing command directly. The consequence chain is why this is worth stating rather than just fixing: an unchecked `server create` let phase 2 run its rsync against a server that did not exist, and report success doing it. **Server-type capacity is per location and changes hour to hour.** cx53 was unavailable in `fsn1`, the default, on the first real run, and available in an alternative location the same minute. Which datacenters currently have it is not in `hcloud server-type describe` — it lives in `server_types.available` per datacenter. `phase_1_provision()` takes `HCLOUD_LOCATION` (default `fsn1`) rather than hard-coding a region, `lib/drill-common.sh`'s `drill_create()` honours the same variable, and `bootstrap.sh`'s `cx53_locations()` looks up current availability to name, in the failure message itself, which locations would work right now. **`hcloud server create` returns before the server can be reached.** The instance exists once the API call returns, well before cloud-init has created the operator account and started sshd — the first real run got "Connection refused" in phase 2, seconds after phase 1 had reported success. `wait_for_ssh()` blocks until sshd answers as `$FLIT_USER`, and both `phase_1_provision()` and `phase_2_deliver()` call it now. `lib/drill-common.sh`'s `drill_create()` had always had this wait; the gap was that the production path had not. **Hetzner reuses IP addresses.** A rebuilt server can answer on an address whose old host key is still trusted, and `StrictHostKeyChecking=accept-new` rejects the mismatch rather than accepting a key that changed for a reason that looks identical to an attack. `phase_1_provision()` drops the stale `known_hosts` entry for the new server's address, but only on the create path — everywhere else a changed host key is exactly the warning it appears to be. `wait_for_ssh()` recognises the mismatch and dies immediately instead of waiting out its full ~400s timeout, since waiting cannot fix it. **A key that exists is not a key that can sign.** Every ssh here runs `BatchMode=yes`, so a key whose private half is encrypted and absent from the ssh agent authenticates nowhere, and it fails late and misleadingly: sshd answers the offer with PK_OK, ssh cannot then produce a signature, and the failure reads as "Permission denied (publickey)" against a key the server does hold. `resolve_operator_key()` resolves the operator's key from the agent instead of guessing a filename, refuses to choose when the agent holds several, and accepts `HCLOUD_SSH_PUBKEY` as either a path or the key line itself — replacing a hardcoded `~/.ssh/id_ed25519.pub` default that the first real run resolved to a key nobody had unlocked. `server/files/cloud-init.sh` now consumes that same resolved key as `OPERATOR_PUBKEY` rather than reading a file on its own; it previously read `~/.ssh/id_ed25519.pub` with a fallback to `id_rsa.pub`, so the key registered with Hetzner and the key written into the server's `authorized_keys` were two independent guesses free to disagree. `phase_1_provision()` also compares the key material behind an existing Hetzner key name rather than trusting the name, since a name that already exists says nothing about which key is registered under it. **Evidence outlives the terminal.** Leaving the instance running only helps if there is also a record of what the run said. `run_log_start()` in `lib/common.sh` tees both streams of `bootstrap.sh` and of both drills into `~/.local/state/$FLIT_NAME/-.log`, with `-latest.log` pointing at the run in progress. The first real bootstrap run failed and left nothing behind — no phase number, no message — which is what this exists to prevent. Three properties are load-bearing: - **Capture at the file descriptor, not in `log()`/`info()`/`warn()`/`die()`.** Half the output of a run is not written by those: server-side output arrives through `ssh`, and `rsync`, `hcloud` and the `laptop/` scripts write on their own. All of it is a child's stdout or stderr, so redirecting both fds is the one point that catches every source. - **`--print-only` writes no log.** It promises to mutate nothing anywhere, and a file under `$HOME` is a mutation. Both drills previously took a `mktemp` log unconditionally and broke that promise. - **Everything caught is on disk, and for drills is uploaded** — `drill_report()` POSTs the log to healthchecks.io. `vault_get()` prints to stdout, so a bare call would durably leak a credential, and `set -x` would leak every expanded command. Neither exists in the repo; both are recorded as hazards where `run_log_start()` is defined. **Ownership of the lockout constraint.** `40-firewall.sh` refuses to run unless `30-tailscale.sh` has completed *and* tailnet connectivity is verified from the server. Enforce it there, in the script that would cause the lockout — not in `bootstrap.sh`, which is one caller among several, and not by phase ordering alone, which a re-run with `--target` could bypass. **A `-c` tmux cannot enter is not a `-c` tmux refuses.** The `$FLIT_ALIAS` helper that phase 6's `laptop/install.sh` generates ran `tmux new-session -A -s '$p' -c '~/workspace/$p'`. The tilde sat inside single quotes, so no shell ever expanded it — tmux received the literal string `~/workspace/`, naming a directory that cannot exist — and nothing created `~/workspace/` in the first place. tmux does not fail on a `-c` it cannot enter; it silently starts the session in `$HOME` instead, so every session opened in the wrong directory while appearing to work, which is what happened on first real use of the box. `laptop/install.sh`'s `POSIX_HELPERS` and `FISH_HELPERS` now create the directory before invoking tmux and leave the tilde unquoted for the remote shell to expand. **A deferral is not a failure.** `phase_9_backups()` called both laptop scripts with a bare `|| die`. Exit code 3 means the script did its job and left one edit to the operator (§5.8); `phase_6_terminal()` had always allowed for that, but phase 9 had not. `--replace` under `--non-interactive` is *designed* to defer removing the superseded restic key, because there is no prompt to decline and two working keys is a safe resting state — so on the first real rotation the deferral fired exactly as intended, phase 9 died on it, and `60-backup.sh` and the Hetzner backup enable never ran. The one path the deferral exists to serve was the one path that could not complete. `phase_9_backups()` now treats `EXIT_DEFERRED` as success and dies only on any other nonzero status. **A Storage Box key was proved with a shell command it cannot run.** A Storage Box runs a restricted shell that answers every command with `Command not found` and exit 8, so `ssh ... true` fails against a key that works perfectly — confirmed against the live box, where the key restic had been backing up with for a day also failed it. `laptop/backup-credentials.sh` used that idiom twice, wrong in opposite directions: the rotation's proof step read the failure as a bad key and refused to promote, so `--replace storagebox-key` could never have succeeded, and the "already authorised" probe read the same failure as an unauthorised key, so its true branch never ran and every invocation re-uploaded — which is what `ssh-copy-id`'s "All keys were skipped" line had been reporting all along. Both now go through `storagebox_key_works()`, which proves the key over sftp with `-b /dev/null` — connect, authenticate, run nothing, exit — the same protocol restic itself speaks to this target. ### 12.2 Manual touches **Four**, all interactive by nature. `bootstrap.sh` must *drive* these — pause at the right moment, print what to do, wait for confirmation, then verify the result — rather than leaving them for the operator to remember out of band. Under `--non-interactive` it must instead **assert and fail fast**, naming the unmet prerequisite. Drills run this way; a prompt there is a hang, not a pause. | Step | Why it cannot be scripted | When | | --- | --- | --- | | Create ntfy and healthchecks.io accounts, and attach both integrations | Signup, token generation, and integration creation are all UI-only — the Management API's `channels/` endpoint is GET-only and cannot create integrations. Attach one ntfy and one email integration at project level. Store the healthchecks ping key and read-write API key in Proton Pass; the ntfy token stays at healthchecks and never reaches the server. **Phase 8 verifies this and aborts if either integration is missing**, so a mistake here fails loudly rather than producing checks that alert into nothing. | Before phase 1 | | Authenticate `pass-cli` on the laptop | An authenticated session — not the read-only access token — is required to *write* the escrow items in phases 1 and 9 (§6.2), and to mint the access token itself in `laptop/prepare-vault.sh`. `bootstrap.sh` must verify this before phase 1, not discover it after generating credentials it cannot store. | Before phase 1 | | Disable Tailscale node key expiry | Not exposed by the auth-key API — the node inherits expiry from the key, and only the admin console can clear it. `server/lib/30-tailscale.sh` asserts it rather than trusting it, printing the console instruction and refusing to continue, so the omission cannot pass silently. This is the step guarding the §4.2 lockout. | During phase 3 | | Authenticate Claude Code | Interactive OAuth browser flow | After phase 4 | **Eliminated, and worth keeping eliminated:** - **Creating the Proton Pass access token in the web UI** — `laptop/prepare-vault.sh` mints its own with `pass-cli personal-access-token create` and scopes it read-only with `access grant --role viewer`. Nothing left for the operator to copy by hand. - **Ordering a Hetzner Storage Box** — Storage Box moved off Hetzner's legacy Robot console into Hetzner Console and the Hetzner API, and `hcloud` (1.67+) exposes it as `hcloud storage-box`. `laptop/provision-storagebox.sh` creates it in phase 9 and escrows its username, hostname, and password itself (§6.2). - **Tailscale device approval** — use pre-generated auth keys stored in Proton Pass and pass `--authkey` to `tailscale up`. **Two keys are needed**: persistent for the production node, ephemeral for drill instances (§7.1) — the Ephemeral toggle on the key, not a separate key type. Mark both **Reusable**; the console default is single-use, and a single-use key works once, then fails silently on the next rebuild or drill. Fully unattended, and it removes a manual step from the phase with the §4.2 lockout risk, which is exactly where an operator waiting on a web UI is most dangerous. - **SSH pubkey on Tangled** — one-time per identity, not per rebuild. Already done; `verify.sh` asserts a test push succeeds rather than re-registering the key. - **Ghostty and shell configuration** — moved into `laptop/install.sh` (§5.7). - **Creating individual healthchecks** — the Management API creates all eight with correct schedules and grace periods in phase 8. Only the accounts and their keys are manual. If a future change adds a manual step, that is a signal the change is wrong, not that this table needs another row. **Separately from this table:** any script may defer a config edit to the operator under §5.8 when it cannot apply it surgically — most likely where dotfiles are symlinked from a repository. That is correct behaviour, not a defect. `bootstrap.sh` collects every deferred change and prints one consolidated list at the end. --- ## 13. Change guards This spec went through roughly ten review rounds using seven distinct techniques, plus implementation — which found more than all but one of them. Each technique found a defect class the others were blind to, with almost no overlap: reading found contradictions but missed interaction bugs entirely; tracing found those but was blind to artifacts named once and never defined; extraction found those but not schedules that collide. **Running one technique repeatedly converges quickly and misleadingly** — severity here went *up* across successive reading passes, not down, because the technique had stopped matching the remaining defects. **Run these against any amendment, with no exemption for amendments made using this list.** Ordered by how much they found. ### 13.1 Trace, do not read Reading verifies that each section is internally consistent. Most serious defects were *between* sections and invisible to reading: - Pick an execution path and walk it end to end: a cold build, a resume after failure at phase 5, a drill, a recovery from total loss. - Pick one credential and follow it to every consumer. Ask what each consumer needs *at the moment it runs*. This is what surfaced `bootstrap.sh` recursing into itself, drills mutating production monitoring, and the restic password existing only on the machine it was meant to protect. ### 13.2 Check the execution context, in every dimension **By far the most repeated defect here — twenty-five occurrences and counting.** This is the only place in the repository that states the count; everywhere else points here. This section originally said only *"can it reach the credential it needs?"* That framing was too narrow, and the narrowness itself cost five further bugs: the credential boundary was clean while the same class of mistake sat on four other dimensions nobody was checking. For anything new, ask all six: | Dimension | The question | What it missed here | | --- | --- | --- | | **Machine** | Laptop or server? Can it reach what it needs *there*? | Check creation and escrow placed in `server/lib/`, needing laptop-only keys | | **User** | Which account runs this, and can that account read what it opens? | Jobs running as `$FLIT_USER` sourcing a `0600 root:root` file. Also `install -d -m755 -o "$U" -g "$U" "$H/.config/mise"` in `20-toolchains.sh` setting ownership only on the leaf directory it names, leaving the `~/.config` parent it had to create along the way `root:root` while `~/.config/mise` beneath it was correctly `$FLIT_USER:$FLIT_USER` — the failure surfaced two phases later and against the wrong script, as `50-caches.sh` reporting `pnpm` unusable because `EACCES: permission denied, mkdir '/home/$FLIT_USER/.config/pnpm'`. Also the tailscale clone guard (§6.1): a root-owned systemd unit with `ExecStart` pointing into `/home/$U/$FLIT_NAME`, the delivered tree `$U` can write — a root unit executing a `$U`-writable path is root for anything that can write it (§5.7). Fixed by installing the guard to `/usr/local/sbin` instead | | **Environment** | Login shell, or systemd's minimal one? | Every scheduled job: no profile, so `mise`/`pnpm`/`cargo`/`claude` absent from `PATH`. Also a key usable from an interactive shell is not a key usable under `BatchMode=yes` — `resolve_operator_key()` exists because a key an operator can unlock by typing a passphrase is not a key that can sign in a script that never prompts. Also `drill_assert_fixtures()` invoking `pnpm` and `cargo` over plain ssh, which is the same non-interactive shell every scheduled job gets: the assertions reported "pnpm: command not found" against a box where pnpm was installed and working, so `drill_ssh_toolchain()` now exports the PATH `verify.sh` exports. Also a credential the box HOLDS is not a credential the box has USED — `pass-cli` takes a token only as `--pat` and reads no environment variable, so the token in `/etc/$FLIT_NAME/secrets.env` authenticates nothing until `vault_session_ok()` redeems it. Also an inherited environment variable is a channel no caller opened on purpose: `lib/common.sh` defaults `FLIT_REPLACE` from the environment the same way it defaults `FLIT_NAME`, so a bare `REPLACE` — generic enough that something else in an operator's shell could set it for an unrelated reason — would reach `wants_replace()` in a child process and rotate a credential nobody asked to rotate; the `FLIT_` prefix is what keeps the inheritance narrow enough to trust | | **Binary** | Which machine needs the tool installed? | `sshpass` installed on the server; used on the laptop. Also `drill_precondition()` requiring `restic` on the laptop when every restic call runs through `drill_ssh` on the server. Also `30-tailscale.sh` (phase 3) needing `jq` to read `tailscale status --json` before `10-packages.sh`, which installs it, runs in phase 4 — ascending numbers in `server/lib/` describe the file listing, not phase order. Also a binary that EXISTS is not a binary that is REACHABLE: `corepack` installs `pnpm` into the node installation's own bin directory, and mise generates its shims at install time, so the freshly installed `pnpm` had no shim and `pnpm add -g` failed two lines after corepack reported success. An established box hides this, because some later mise operation reshims | | **Userland** | Is the tool the same *implementation* on both sides? | `awk -v` with a multi-line value: gawk tolerates it, BSD awk exits 2. Also `ncurses`/`terminfo`: the same tool on both machines, but Ubuntu 24.04's database predates several terminals in current use, so a `TERM` that resolves on the laptop does not resolve on the server — `bootstrap.sh`'s `push_terminfo()` exists because ssh forwards `TERM` verbatim and tmux exits with `missing or unsuitable terminal` on the first session, which is exactly what happened on the first real use of the box. Also a Storage Box: the same `ssh` client proving a key against the server and against the Storage Box, but the Storage Box answers every command with `Command not found` and exit 8 rather than running it, so `ssh ... true` reports a working key as broken — confirmed against the live box, where the key already backing up production failed the identical probe. `storagebox_key_works()` (`laptop/backup-credentials.sh`) proves the key over sftp instead, the protocol restic itself speaks to that target | | **Direction** | Does the guard test the direction the risk runs in? | Lockout guard verifying *outbound* reachability when the risk is *inbound*. Also `bootstrap.sh` `phase_3_network()` ending with `HOST=$FLIT_SERVER`: the phase switches its ssh target from an address to a NAME, and the shipped default names the production box, so a drill — which creates its own instance and passes `--target` — had phases 4 through 11 reprovision production while the instance under test sat untouched. Reachability was proven; *which machine* answered was not. The guard now compares `/etc/machine-id` across the two targets before switching | The sixth dimension is the newest and the least intuitive: the laptop is macOS and the server is Ubuntu, so `awk`, `date`, `stat`, `sed`, `wc`, `install` and `mktemp` are *different programs* with the same names on the two sides. The first laptop run of `lib/common.test.sh` — a suite that had passed on Linux throughout — failed five assertions, and three of them were `cfg_apply()` truncating the file it was editing to zero bytes while returning 0 and logging success. Anything under `lib/` is the exposed surface, because it is the only code that runs on both machines. Server-only GNU-isms in `verify.sh` and `server/lib/` are correct and must not be "fixed". **Do this mechanically, not by reading.** Build a table of every script against the side it runs on and the capabilities it references — a short script over the source finds these in seconds and does not get bored. Cross-reference §7.1's vault inventory. Then re-run it after any change that moves work between the two machines. ### 13.3 Extract artifacts and check each is defined List every file, script, slug, and flag the spec names. Anything mentioned **exactly once** is suspect — it has a reference but no definition, or a definition nobody uses. This found `70-alerting.sh` sitting in the wrong directory, two checks with no implementing script, and the Node pin assertion specified at two cadences with one implementation. ### 13.4 Cross-check every quantity Counts drift as sections are edited. Verify each stated number against the table it summarises, and check numbers chosen in different sections against each other. The 90-day token expiry and the ~91-day drill cadence were each sensible alone and collided by construction. ### 13.5 Walk the timeline, not the document Read the spec as a sequence of states over a year: day 1, day 90, day 180, month 6, month 12. Recurring schedules interact in ways no single section shows. This found the Tailscale key expiring into a lockout at ~180 days, a check that would sit DOWN most of the year and mask the dead-man property, and a golden snapshot billing forever for one use. ### 13.6 Ask what the laziest passing implementation looks like For every assertion, ask what the cheapest thing that satisfies it literally while being wrong would be. If that exists, the assertion is underspecified. This found that a `verify.sh` skipping every check would report healthy, that `restic backup` exits 0 on an empty directory, and that an escrow readback could echo a variable rather than query the vault. ### 13.7 Write the code **The single most productive technique, and it is not a review at all.** Implementation has found more defects than any reading pass, because a script cannot hand-wave about where it runs: the filesystem either grants the permission or does not, the binary is either on `PATH` or is not. Two examples it caught that six review passes had not: phase 9 needed the same laptop/server split as phase 8, because a server-side script can generate the restic password but can never escrow it; and the console password must be escrowed *before* the server is created, or a failure mid-create leaves a machine whose break-glass password exists nowhere. Corollaries worth keeping: - **Feed findings back into the spec immediately.** Half of these changed the design, not the code. A spec that lags its implementation is worse than no spec. - **Write the tests for the contract, not the function.** `lib/common.test.sh` asserts that the SSH block lands *above* an existing `Host *` and the Ghostty block *below* — the two failures that produce no error at all. - **Lint and cross-check mechanically.** Timer `OnCalendar` expressions were verified equal to their healthchecks `schedule` fields by script; a mismatch there yields phantom DOWN alerts that look like real failures. - **A `--print-only` plan line is not evidence the code performs it.** Phase 1's plan line already claimed the operator's SSH key was registered while `hcloud ssh-key list | head -1` merely hoped one already existed — a dry run printed the intent while the real run did something else. ### 13.8 Cross-check the spec against the code Once an implementation exists, the two can disagree — and the spec is the one that lies, because the code has to actually run. Check mechanically: - Check slugs in §6.3 against `laptop/healthchecks.sh`. - The `bootstrap.sh` flag table against the flags the script actually parses. - File modes and ownership in §7.3 against what the scripts write. - Every §7.3 residue item against the step that installs it. On its first run this found the spec claiming `0640` files were `0600`, an undocumented `--start-at` flag, and a §7.3 entry with no installer — **all caused by a batch of spec edits that threw partway through and was silently discarded in full.** Scripted edits fail atomically or not at all; verify the write landed rather than trusting the exit code of the editor. ### 13.9 Then read it Reading still catches contradictions, stale wording, and claims that were true two edits ago. It is a poor first pass and a necessary last one. --- ### Standing rules these produced Violating any of these has broken something before: 1. **Each check in §6.3 maps to exactly one script.** No assertion at two cadences. 2. **Guards live on the side that would cause the failure**, and can reach what they need there. 3. **Idempotency guards inspect state, never a sentinel file.** 4. **A zero exit is not proof of work** — assert content. 5. **Anything that can be skipped forever will be.** Record when it last actually ran. 6. **Credentials generated on the server are escrowed off it, at generation time.** 7. **A permanently red check is not a check.** Separate advisory from failure. 8. **`--print-only` and dry-run paths mutate nothing, anywhere.** 9. **Scheduled jobs run under a login shell**, or the toolchain is not on `PATH`. 10. **A credential the design assumes exists must be *installed* by some step.** Naming it in §7.1 is not the same as putting it on the machine. 11. **`.env` may only reference vault items that exist before `bootstrap.sh` starts.** `pass-cli run --env-file` resolves every reference up front and fails the whole run on the first missing one. Items the run itself creates — the Storage Box credentials, the restic repository URL, the restic password — are read with `vault_get` at the point of use instead. 12. **A guard is only real if every code path that can bypass it is also guarded.** `--print-only` was asserted at the top level while `phase_13_drills()` invoked both drills unconditionally, and `laptop/healthchecks.sh`'s delivery test created a check and sent live pings regardless. 13. **An early-exit consumer at the end of a pipeline fails the pipeline under `pipefail`.** `grep -q`, `head` and their kin close the pipe at the first match; a producer still writing takes SIGPIPE, and `pipefail` reports its 141 as the pipeline's status. Small output hides this completely — the producer finishes writing into the pipe buffer before the consumer quits — so the bug ships green and surfaces the day the data grows. Materialise the output and search it instead. This is worse than an ordinary false alarm when the pipeline *is* an assertion: the failure is indistinguishable from the condition being asserted against. 14. **A restored or rebuilt host obtains every credential from the vault, never from the image.** An image freezes whatever a credential's value was at capture time; a rotation that runs afterward retires that value everywhere the rotation reaches, but the image's copy still matches what it captured. The gap between "image taken" and "rotation run" produces no failure until something restores from that image, so it ships green and fires only during a real recovery. 15. **Name resolution is not reachability, and a probe by name must fall back to an address.** A bare MagicDNS name stops resolving whenever the laptop loses the tailnet search domain — macOS on a phone hotspot is the ordinary case — while the node stays up and answers on its tailnet address. Three sites treated the two as one thing: `bootstrap.sh` `phase_2_deliver()` concluded the tailnet was down and fell back to a public address that a converged box firewalls, so the run hung rather than failed; `phase_3_network()` refused to close port 22 over a tailnet that was not broken; and `lib/drill-common.sh` `drill_wait()` spent forty attempts before blaming the ephemeral auth key. Ask `tailscale ip -4` before giving up. Try the name first — when it works it proves what the operator types — and warn when the address is what answered. This weakens no identity check: what those defend against is a NAME reaching some other machine, and an address cannot be repointed by DNS. 16. **A prompt must not open its own reader on the terminal.** Every credential-touching script here runs under `pass-cli run`, which hands its child a pipe for stdin and reads the terminal itself to fill it. A prompt that reads `/dev/tty` directly becomes a second reader racing that wrapper: the first answer typed is consumed by the wrapper and delivered into a pipe nobody drains, so the prompt looks ignored and has to be answered twice. Read the stdin the process was given. Writing the prompt to `/dev/tty` is a separate matter and still required, because `run_log_start` routes fd 2 through an awk filter that emits only complete lines and a prompt has no trailing newline. `lib/common.sh` `confirm()` is the one implementation. 17. **An exclusion list matches names; git knows facts. Let git overrule it.** A pattern like `build` or `dist` is a guess that a directory holds regenerated output, and a repository is free to track a source file inside one — `apps/publisher/build/docs.ts` is a real case. The damage is not the missing file but that it cannot be repaired: the server's checkout reads as dirty from then on, every later push stops to ask about overwriting changes nobody made, and pushing again cannot fix it, because the exclusion that caused the deletion also blocks the file. Name git's tracked files, and every ancestor directory of one, ahead of any general pattern — `git_protect_list()` in `lib/common.sh`. Ancestors matter because rsync prunes an excluded directory during the walk and never descends into it. Where a filter is still allowed to strand a tracked file — a project's own `.flit-push-exclude`, which outranks the protection on purpose — the run reports what went missing rather than leaving the next push to discover it, and it asks the far side's git for that answer instead of re-deriving it from the patterns. 18. **A whole-content comparison cannot stand in for a managed block.** `cfg_apply()`'s effective-state check (§5.8) works because a managed block gives it a scope to ignore everything else in the file. `deliver_rendered()` (`server/lib/22-claude-config.sh`) compares whole content instead, for formats with no comment syntax to carry the markers, and that comparison has no scope to ignore anything: a byte written by any third party — not the operator, not a managed block — reads as a real conflict and defers, and keeps deferring on every later run, because nothing changed the byte back. Claude Code writing a `"theme"` key into `~/.claude/settings.json` at first run is a live case of this, not a hypothetical one. Where a format cannot carry a managed block, keep box-local state out of the delivered file entirely — `settings.local.json` for Claude Code — rather than trying to make the comparison smarter. ### When an amendment is large Re-run §13.1 and §13.2 at minimum, rebuild §13.2's script-versus-context table, and run §13.8. The removal of mosh looked like a deletion and silently invalidated two arguments elsewhere — the renderer choice and the agent-forwarding caveat — both of which had been resting on it.