This repository has no description
Shell 99%
<1%
Rust <1%

README.md

flit #

Provisioning for one remote development machine.

A laptop pushes configuration to a single Hetzner server; the server never fetches its own. All work then happens on the server — the laptop is a terminal, a network stack, and a key. One machine is the whole design, not a starting point: a single source of truth is what removes the "which machine has the uncommitted branch" problem, and every tradeoff below follows from it.

Three documents, with no overlap:

  • This file — how to build the box, use it daily, and operate it.
  • flit-spec.md — the design and the reasoning behind it. Section numbers (§4.2, §6.2, …) referenced throughout this file point there.
  • CLAUDE.md — the brief for an agent making changes to this repository.

How it works #

bootstrap.sh              RUNS ON THE LAPTOP — provision, push, execute
verify.sh                 read-only drift detection (PASS/FAIL/SKIP)
drill-restore.sh          quarterly    — can the work be recovered?
drill-rebuild.sh          twice-yearly — do the scripts reproduce the box?
lib/                      shared primitives + tests
laptop/                   anything needing a laptop-only credential
server/lib/               numbered provisioning steps, run in order over ssh
server/bin/               the jobs the timers run
fixtures/                 minimal projects the drills build against

Credentials flow one way. Every secret is held by the laptop and pushed; the server never reaches back for its own configuration. A script's directory is therefore a claim about which machine can run it. laptop/healthchecks.sh and laptop/backup-credentials.sh stay out of server/lib/ because each needs something only the laptop holds — the healthchecks read-write API key, and an authenticated pass-cli session, respectively. A script under server/lib/ would be structurally unable to do that work. Checking all six dimensions of execution context — machine, user, environment, binary, userland, direction — before moving anything across that laptop/server line is spec §13.2's standing rule, and the most repeated defect in the project.

Commands below are written to be pasted as-is, so they use default values rather than the variables that produce them. Those variables are defined in lib/common.sh under "THE VARIABLE SET"; a deployment that overrode one substitutes its own value, and records what it holds in its own deployment.md.

Appears below as Variable What it names
flit FLIT_ALIAS the ssh Host alias, and the flit/flits shell helpers
flit FLIT_NAME the project name, in paths like /etc/flit/ and in job names
$FLIT_SERVER FLIT_SERVER the Hetzner server and ssh HostName — deployment-specific, so left as a variable
$FLIT_USER FLIT_USER the operator account on the server — likewise

What it needs #

Tools on the laptop. bootstrap.sh preflight() requires ssh, rsync, jq, curl, hcloud (Hetzner's CLI), pass-cli, and sshpass (which uploads the Storage Box key non-interactively in phase 9). preflight() dies naming the first one missing, so installing them up front is a convenience, not a trap.

Accounts. laptop/prepare-vault.sh collects credentials but cannot sign up for anything, so these exist first (spec §12.2 has the reasoning for why each stays a manual, web-UI step):

  • Hetzner Cloud project.
  • Tailscale account and tailnet. Generate two auth keys in the admin console before running prepare-vault.sh: one for the persistent production node, one with the Ephemeral toggle set for drills. Mark both Reusable — the console default is single-use, and authenticates once before failing silently on the next rebuild or drill.
  • healthchecks.io and ntfy accounts, with one ntfy and one email integration attached at healthchecks project level. Phase 8 queries the API and aborts if either is missing.
  • Proton Pass account and vault.

Repositories delivered alongside this one. Phase 2 ships ~/.claude to the server with git archive HEAD, so it must exist as a git repository with nothing uncommitted, or the phase refuses to run. ../dotfiles is delivered too when it exists, and skipped when it does not. Chook is not delivered here: it reaches the box as a normal project, pushed with laptop/push-project.sh like any other, carrying its own .git so work done on it there has somewhere to be committed (see "Moving a project onto the box").

From a clone to a working box #

Activate the pre-commit hook. bootstrap.sh preflight() runs git config core.hooksPath .githooks on every invocation, so the hook that rejects a literal secret in a staged .env is active from the first run. Git does not carry hooks across a clone, so on a clone where bootstrap has not run yet, set it by hand first:

git config core.hooksPath .githooks

Fill the vault.

pass-cli login
./laptop/prepare-vault.sh

pass-cli login opens an authenticated session; writes need it (spec §6.2), and so does prepare-vault.sh. That script is the one thing here not run under pass-cli run --env-file .env: that form resolves every reference up front, and on a first run the items it needs do not exist yet. It prompts for five values, mints a sixth itself (pass-cli-token, scoped read-only), and generates a seventh (fixture-secret), storing each at the path .env already references. ./laptop/prepare-vault.sh --print-only reports which of the seven are present and which are missing, and mutates nothing.

.env needs no editing. Every value in it is already a pass:// reference, which is what lets this repository be public — the pre-commit hook rejects a staged .env holding a literal. It can only reference items that exist before bootstrap.sh starts, so it names neither the Storage Box credentials nor restic-repository/restic-password; phase 9 creates those and reads them with vault_get instead.

A second deployment needs its own .env: pass-cli run --env-file reads each pass:// prefix literally, so overriding FLIT_NAME alone does not redirect them — a .env still reading pass://flit/... would resolve every credential against the first deployment's vault instead of the new one. vault_check_env() (lib/common.sh) checks this before bootstrap.sh, both drills, and laptop/healthchecks.sh touch the vault, and fails loudly on a mismatch rather than reading the wrong vault silently.

local.conf is optional. Copy local.conf.example to local.conf — gitignored, laptop-only. It matters in two cases: the ssh agent holds more than one key (resolve_operator_key() refuses to guess among them, so HCLOUD_SSH_PUBKEY, set in local.conf or exported per invocation, makes the choice explicit), or the default HCLOUD_LOCATION (fsn1) has no capacity for the instance size at the time.

Dry run, then the real run.

pass-cli run --env-file .env -- ./bootstrap.sh --print-only
pass-cli run --env-file .env -- ./bootstrap.sh

--print-only mutates nothing anywhere — laptop config is inspected exactly, server phases print a plan only. The real run needs three things done by hand along the way:

  • Disable node key expiry in the Tailscale admin console when phase 3 demands it. Skipping it schedules a lockout roughly six months out, with public SSH already closed (spec §4.2).
  • Authenticate Claude Code through its browser flow after phase 4 installs it.
  • Answer phase 8's interactive prompt, which asks whether both the ntfy push and the email arrived. The run continues once it is answered.

When a phase fails, resume with --start-at N. The failed instance is left running deliberately — it is the only evidence of what went wrong — and so is the log of the run, at ~/.local/state/flit/bootstrap-latest.log (see "The log of a run"). For N greater than 2, check_delivered_commit() refuses the resume outright when the server's .delivered-commit does not match the laptop's HEAD, rather than silently running server-side code that the fix being resumed toward was never shipped to.

To ship an edit to a server/lib/ or server/bin/ script without a full provisioning run, pair --start-at with --stop-at:

pass-cli run --env-file .env -- ./bootstrap.sh --target $FLIT_SERVER --start-at 2 --stop-at 2
./laptop/sync-config.sh

--stop-at N runs no phase after N. --start-at 2 --stop-at 2 delivers the repository and stops, which is only phase 2 — no later phase runs, so nothing else on the box changes. laptop/sync-config.sh then runs the box's delivered copy of server/lib/22-claude-config.sh to pick up the change. --stop-at before --start-at is refused, as is a non-numeric value — read as "no limit" would defeat the point of the flag. Stopping early logs that the box is not fully provisioned, rather than the usual completion line.

A successful first run leaves the box without chook. Nothing in bootstrap.sh requires or delivers it, so the checkout that every hook in the installed settings.json points at does not exist yet. Pushing it is the next step, and is covered as the first example under "Moving a project onto the box" in Daily use, below.

Daily use #

ssh flit reaches the box because laptop/install.sh writes a Host flit block mapping the alias to HostName $FLIT_SERVER and User $FLIT_USER — that indirection is the point, since the alias stays stable while the server it points at is free to change. The same script adds two shell functions:

$ flit api          # attaches or creates session 'api', cd'd to ~/workspace/api
$ flit web          # completely separate session
$ flits             # list every running session

One tmux session per project is the primary way to use the box, not an edge case. Each session has its own windows, panes, working directories, and running processes, and sessions do not see each other — three terminals in three projects means three independent sets of watchers and language servers, all surviving disconnection independently. new-session -A attaches if the session exists and creates it otherwise, so one command covers both cases. Detach with Ctrl-b d, or just close the terminal; the session keeps running either way.

Attaching two terminals to the same session mirrors them instead of giving independent views, and tmux sizes windows to the smallest attached client. That is a feature — phone-to-desktop handoff works this way — but it surprises anyone expecting independence. Independence comes from different session names, not different connections.

Each project session with a running dev server holds a tsserver at 2–6 GB plus watchers, so watchers, not sessions, are the memory cost: four active TypeScript projects can exhaust 32 GB on their own. Kill dev servers in projects not actively being worked on rather than leaving five running because detaching is free. flits plus htop diagnose things once they feel slow.

Moving a project onto the box #

Chook is the first project pushed onto a newly provisioned box, before using agents there at all. The installed settings.json wires every Claude Code hook to chook.mjs under ~/workspace/chook, and each one fails until that checkout exists:

$ ./laptop/push-project.sh ~/misc/chook

$ ./laptop/push-project.sh ~/misc/myproject          # -> ~/workspace/myproject
$ ./laptop/push-project.sh --print-only ~/misc/myproject
$ ./laptop/push-project.sh ~/misc/myproject other-name

Cloning the repository on the server would be shorter and would lose the half of a working directory that git does not track: a gitignored .env, a handoff.md, .claude/settings.local.json, an edit made and not committed. This copies the directory instead, so what arrives is the working directory rather than a fresh checkout of it. Claude Code's session transcripts for the project come too, re-keyed from the laptop's path to the server's, unless --no-sessions.

What does not come is anything the server can rebuild from something that did: node_modules, target/, dist/, .venv, .wrangler, build caches. Those are most of the bytes, and a node_modules built on an arm64 laptop is the wrong one for an amd64 server anyway. A project needing more exclusions lists them, one pattern per line, in a .flit-push-exclude file in its own root.

That list matches directory names, and a name can be wrong: a repository is free to track a source file under build/ or dist/. Anything git tracks is carried regardless, so a blunt pattern cannot strand it. Your own .flit-push-exclude still outranks that — it is an explicit decision about this project, and a repository may well track a data set the server has its own copy of. When one of your patterns does leave a tracked file behind, the push says which files and what to do about it, because the server's checkout reads as dirty from then on and pushing again will not fix it.

The push is one-way, like everything else here: the laptop is authoritative and the server never reaches back. It is not --delete by default, so a file that exists only on the server survives; pass --delete to make the server an exact mirror. If the server's copy has uncommitted changes, the push lists them and asks before overwriting — --force answers yes in advance. Work done on the box is real work, and this is the one thing that can quietly destroy it.

A project can hold the push off entirely with a .flit-no-push file in its root. This covers the case the dirty-tree guard cannot see: an agent working on the server that commits its work leaves a clean tree, and a clean tree that is ahead of the laptop looks exactly like one that is behind. push-project.sh refuses before transferring anything and prints the file's first five lines as the reason. --force does not lift the hold — it exists to override the dirty-tree guard, not this one — and --print-only is refused too, since even a dry run against a project on hold has nothing useful to report. Deleting the file is what lifts it.

Then open it:

$ flit myproject

Holding syncing off entirely #

A .sync-off file at the repository root refuses every laptop-to-server write, not just one project's. bootstrap.sh, laptop/sync-config.sh and laptop/push-project.sh all overwrite rather than merge, so one switch covers all three; check_sync_hold() in lib/common.sh checks it once, at startup, rather than in each script, so it cannot be lifted halfway. Create it with the reason as its content:

$ echo "rebasing the server's dotfiles by hand, do not overwrite" > .sync-off

The refusing script prints the file's first five lines back, which is what makes the reason worth writing down — it is what shows up weeks later when something else hits the hold. Delete the file to resume:

$ rm .sync-off

This is coarser than .flit-no-push: that file holds off pushes to one project, this one stops provisioning and the configuration sync too. There is no flag to override it — push-project.sh --force already overrides the dirty-tree check, and .sync-off exists for exactly the case that check cannot see, so overriding it the same way would defeat the point. Deleting the file is the only way to lift the hold, and doing so is a decision worth leaving a trace of, unlike a flag on a command line.

Secrets in a project #

Commit .env files holding pass:// references, never values:

DB_PASSWORD=pass://Dev/Database/password

Run anything needing them through pass-cli:

pass-cli run --env-file .env -- pnpm dev

First resolution takes ~5 s; subsequent calls are instant from the kernel keyring, expiring after an hour or at logout. The access token expires every 90 days and fails during interactive work, in the morning, with a clear error — deliberate timing (spec §7.3). Rotate it with ./laptop/prepare-vault.sh --replace pass-cli-token, which mints a new token through pass-cli personal-access-token create and re-escrows it in one step rather than leaving the vault holding the expired one. Backups do not depend on this token, so they keep running through an expiry. The quarterly drill lands at roughly 91 days and so usually wants a fresh token too — in practice, refreshing the token and running the drill are the same quarterly ritual.

Clipboard #

Ghostty owns the macOS pasteboard; vim runs on the server. OSC 52 carries a yank back over the terminal byte stream — yank → escape sequence → tmux → SSH → Ghostty → macOS clipboard — which is why both Ghostty (copy-on-select, clipboard-read/write, set by laptop/install.sh) and tmux (allow-passthrough, set-clipboard, set by server/lib/25-terminal.sh in phase 6) need configuring; tmux is not a dumb pipe and has to be told to forward the sequence through.

tmux is set to set-clipboard external, not on, deliberately: under on, any process running inside tmux — including code from an unfamiliar repository running under Claude Code — can set the system clipboard regardless of which user runs it, and a clipboard write followed by a shell paste is a real attack. external restricts clipboard-setting to tmux itself, so copying goes through copy-mode rather than vim's yank register:

Ctrl-b [        enter copy-mode
                Space to start selection, motion keys, Enter to copy

Paste needs none of this — Cmd-V works natively; OSC 52 paste is deliberately awkward for security reasons and is not worth using. Large copies vanishing while small ones land is an OSC string length limit that terminals impose and drop silently, not a broken config.

Claude Code renderer — stay on classic #

server/lib/25-terminal.sh (phase 6) sets CLAUDE_CODE_DISABLE_ALTERNATE_SCREEN=1 on the server, where Claude Code actually runs; setting it on the laptop does nothing. The fullscreen renderer draws into the alternate screen buffer, moving the conversation out of native scrollback — replacing Cmd+f and tmux search with Ctrl+o transcript mode — and out of reach of tmux copy-mode, which is the clipboard path above. The classic renderer keeps everything in native scrollback, where copy-mode works normally.

Fullscreen virtualizes the viewport and transmits only changed regions, cutting ANSI escape data — real savings on a slow link. Classic still wins: losing clipboard access costs more every day than the bandwidth costs on occasional slow links. To use fullscreen anyway, set CLAUDE_CODE_DISABLE_MOUSE=1 alongside it, since fullscreen's mouse capture fights copy-mode directly. It cannot be escaped entirely either way: background sessions opened from agent view or claude attach always use fullscreen, and the env var does not apply to them.

Terminal type #

bootstrap.sh's push_terminfo() (phase 6, right after server/lib/25-terminal.sh) pushes the laptop's terminal description to the server, because Ubuntu 24.04's terminfo database predates several terminals in current use — xterm-ghostty arrived after it — and ssh forwards $TERM verbatim, so without the push tmux exits with missing or unsuitable terminal the first time flit runs. The push writes ~/.terminfo rather than the system database, so it needs no sudo, and only the laptop is touched for reading: it knows $TERM, the server does not. If it ever needs redoing by hand:

infocmp -x "$TERM" | ssh flit -- tic -x -

Browser previews #

Dev servers bind on the box; Tailscale reaches them from the laptop browser directly:

http://$FLIT_SERVER:5173

No tunnel, no port forwarding, no -L flags. Bind to 0.0.0.0, not 127.0.0.1, or the tailnet cannot reach it — Vite needs --host 0.0.0.0 or server.host: true. This is per-project, so it is the one piece of the setup above that the scripts cannot do.

Travel #

ssh flit, same as always — there is no second transport to remember, since mosh was removed deliberately (spec §10), so git push works from every session and nothing about the workflow changes across a trip. What changes is typing: at a distance of, say, Asia from the server's datacenter, the round trip shows up in vim insert mode, roughly 180 ms per keystroke. Normal-mode motions and commands are unaffected; sustained prose entry is the unpleasant part. In order of usefulness: compose long prose locally and paste it in (Cmd-V works natively into insert mode); stay in normal mode, since motions, ., macros, and : commands send one round trip for a whole operation rather than one per character; run builds and tests as usual, since only interactive echo is latency-bound.

A dropped connection costs a reconnect, never state — tmux holds everything, and ControlPersist makes reattaching near-instant. One side effect of that multiplexing: changing networks while a master connection is open leaves the socket stale, and the next ssh flit hangs rather than erroring; ssh -O exit flit clears it. There is no offline mode (see Known limits), so push everything before departure.

Checking and troubleshooting #

verify.sh is read-only drift detection, reporting PASS/FAIL/SKIP. It runs on the server: bootstrap.sh phase 11 invokes it there over ssh, a weekly timer runs it there unattended, and it can be run there by hand at ~/flit/verify.sh. Remediation for anything it reports is bootstrap.sh --target $FLIT_SERVER from the laptop — configuration is pushed, never edited on the box.

Alerts #

Everything on this box runs on a timer, and every timer checks in with healthchecks.io. It fans out to two channels:

job on $FLIT_SERVER  →  healthchecks.io  →  ntfy push  +  email

The ntfy topic is meant to be subscribed to on phone and desktop. Nothing publishes to ntfy from the server — the token lives at healthchecks, which is why it is not among the credentials on the box. The topic name is high-entropy on purpose: public ntfy topics are readable by anyone who guesses the name, and these messages carry the hostname, disk state, and backup outcomes.

Alert Meaning
verify.sh drift Something diverged from the scripts — run bootstrap.sh --target $FLIT_SERVER from the laptop; configuration is pushed, never edited on the box
Backup failed restic errored to the Storage Box; failure output is in the healthchecks dashboard
Rebuild drill diff non-empty Packages on the live box are not in server/lib/10-packages.sh's package list — add them there, then re-run bootstrap.sh --target $FLIT_SERVER
Node pin stale ~2 Electron majors behind. Bump server/files/mise.toml, then bootstrap.sh --target $FLIT_SERVER. Routine divergence never alerts — it just shows in the dashboard
Project config check failed The throwaway SessionStart hook did not fire through a symlinked .claude; quota/session-limit failures after it fires do not fail the check — see server/bin/check-project-config.sh
Backup credentials missing from vault Urgent. Re-escrow immediately — backups are otherwise undecryptable if the box dies (see "If the box is gone")
verify.sh reports SKIP, not FAIL A check could not run — expired token, network down. Not drift; no action unless it repeats
Disk 80% Run cargo sweep and pnpm store prune
… is DOWN A scheduled job did not run at all — see below
Drill DOWN No drill has run this quarter — run it from the laptop (spec §8)

Read DOWN differently from failed. Because everything is a check-in, healthchecks can alert on a job that never ran, not just one that ran badly (spec §6.3): "failed" means the job ran and something went wrong, usually narrow and obvious; "DOWN" means the job never ran — the timer died, the box is unreachable, or something upstream broke, a bigger question than whichever job's name is attached to it. A DOWN on the daily backup after a reboot is the case worth catching, since nothing else in this design would report it.

Every alert also goes to email, deliberately — email is slow and easy to ignore, fine for a backstop, wrong for a primary; if ntfy itself is broken, notification failure would otherwise be indistinguishable from everything being fine. The two drill checks watch whether a drill ran, not the server: drills run from the laptop, since they need the Hetzner API token, which deliberately never touches the box (spec §8), so a drill DOWN means no drill has run this quarter, not that something broke — running it clears the check. Still not covered: healthchecks.io going down, unnoticed. That is where the layering stops; the quarterly restore drill is the backstop for the backstop.

When something looks wrong #

Symptom Check
$FLIT_SERVER will not connect Tailscale up on both ends — public SSH is closed by design. If Tailscale itself is the problem, use break-glass (below)
Tailscale key expiry warning Act on it. When the key lapses the box leaves the tailnet, locking out SSH access (spec §4.2)
pass-cli fails Token expired at 90 days — rotate it with ./laptop/prepare-vault.sh --replace pass-cli-token
Copy does not reach macOS vim yank? Use tmux copy-mode — set-clipboard external blocks inner apps by design (see Clipboard)
Large copies vanish, small ones work OSC string length limit, not config (see Clipboard)
Claude Code scrollback gone Env var set on laptop instead of server (see Claude Code renderer)
Browser cannot reach dev server Bound to 127.0.0.1 instead of 0.0.0.0 (see Browser previews)
flit / tmux exits missing or unsuitable terminal Server terminfo lacks the laptop's $TERM (see Terminal type); by hand: infocmp -x "$TERM" | ssh flit -- tic -x -
flit <name> opens in $HOME, not ~/workspace/<name> Old helper left the tilde quoted, so no shell expanded it, and nothing created the directory — tmux falls back to $HOME silently rather than failing (see Daily use)
Builds slow, machine sluggish Too many watchers (see Daily use), or disk near full (see Known limits)
verify.sh alert Drift from the scripts — run bootstrap.sh --target $FLIT_SERVER from the laptop; configuration is pushed, never edited on the box
ssh flit hangs after changing networks Stale multiplexed socket. ssh -O exit flit, then reconnect
ntfy alert: Node pin stale ~2 Electron majors behind. Bump server/files/mise.toml, then bootstrap.sh --target $FLIT_SERVER (spec §5.4)
Backup DOWN after a reboot Timer did not come back — check systemctl list-timers (see Alerts)
A setting from Clipboard or Claude Code renderer seems absent Did install.sh defer it? Re-run with --print-only and check the deferred list
bootstrap.sh or a drill failed Read the log, not the scrollback — see below

The log of a run #

bootstrap.sh, drill-restore.sh and drill-rebuild.sh each write every line either stream carries to:

~/.local/state/flit/<name>-<timestamp>.log

with <name>-latest.log pointing at the run in progress, so a long bootstrap.sh can be watched from a second terminal:

tail -f ~/.local/state/flit/bootstrap-latest.log

A failed run names its own log path before it exits. --print-only writes no log at all, since it mutates nothing anywhere. Ordering within one stream is exact; between stdout and stderr it is not, so a line that looks out of place next to its neighbour probably came from the other stream. The files are 0600 and are never pruned — the drills each keep two runs a year, and a failed bootstrap is worth more than the bytes.

Recovery and rebuild #

If the box is gone #

Restoring recovers the work; it does not restore the credentials the machine needs to function. Four things must be re-supplied by hand — the full procedure is spec §8.3. Short version: create a new pass-cli token in Proton Pass first (nothing else works without it), provision a replacement, write /etc/flit/secrets.env, restore the Storage Box SSH key, then run verify.sh and let it report what is still missing.

The restic password and Storage Box key live in Proton Pass, escrowed there when the box was built (spec §6.2) — the copies on the server are a cache, which is what makes recovery possible at all. verify.sh re-checks that both are still in the vault on every run; if that alert ever fires, treat it as urgent rather than cosmetic, since it means the backups are one disk failure from being permanently undecryptable. Proton Pass is the single point of failure for this whole design — the account's recovery codes belong somewhere that is not this machine.

Break-glass #

If $FLIT_SERVER is unreachable — Tailscale down, key expired, an over-restrictive firewall rule — Hetzner's Cloud Console has a browser terminal that bypasses SSH entirely. The $FLIT_USER password is in Proton Pass as $FLIT_USER-console-password, put there at build time for exactly this; it does not work over SSH by design, only in that console. Confirm it is actually in the vault before an outage is the first occasion to need it.

Drills #

pass-cli run --env-file .env -- ./drill-restore.sh
pass-cli run --env-file .env -- ./drill-rebuild.sh

Assertion is not evidence, so each drill proves something a green bootstrap.sh run does not. drill-restore.sh (spec §8.1, quarterly) proves the work is recoverable, restoring the most recent Hetzner automatic backup — or, with --from-snapshot, the golden snapshot, which is the only source available before any automatic backup exists. drill-rebuild.sh (spec §8.2, twice-yearly) proves the scripts reproduce the box from a stock image.

Step zero of both is a precondition check: drill_precondition() (lib/drill-common.sh) asserts the pass-cli token still works before anything is provisioned, because the token's 90-day expiry and the drill's quarterly cadence collide by construction, and an expiry discovered several minutes and one provisioned CX53 into a run reads like a broken restore path rather than routine credential expiry (spec §8).

Each drill destroys its instance on every path, including failure, and pings its own healthchecks slug with the exit code.

Four flags always travel together: --skip-drills --non-interactive --no-external-state --skip-laptop. drill-rebuild.sh passes all four to bootstrap.sh itself, and any other invocation of bootstrap.sh against a throwaway instance needs the same set — omitting any one fails differently, and none fails loudly (spec §12.1).

Tearing down and recreating production #

hcloud server delete $FLIT_SERVER removes only the compute instance. Re-running

pass-cli run --env-file .env -- ./bootstrap.sh

recreates it: phase 1 finds no server named $FLIT_SERVER, creates a fresh CX53 from the same cloud-init, registers the operator SSH key, and mints a new console password, overwriting the escrowed one. Phase 2 onward reprovision it exactly as a first build would, reusing whatever the vault already holds rather than starting over — phase 9 finds the Storage Box already exists and pushes the existing restic-password and storagebox-key onto the new box, so backups resume against the same repository, and phase 8 upserts the same eight healthchecks slugs rather than creating duplicates.

Three things are not cleaned up by any of this, and need separate attention:

  • The Storage Box is untouched either way — a separate Hetzner resource, not part of the compute instance, and laptop/provision-storagebox.sh only creates one when none exists.
  • The healthchecks checks are untouched; they go DOWN while the old server is gone and recover once the new one's timers start pinging again.
  • The old Tailscale node is not deregistered by hcloud server delete. Nothing here does that either — drill_cleanup() (lib/drill-common.sh) only ever destroys the Hetzner instance a drill created, and a drill's own node uses the ephemeral key and self-removes; a torn-down production node used the persistent key and does not. Remove the old device from the Tailscale admin console before or right after rebuilding, or the new instance — booting under the same hostname — registers under a suffixed name instead of the one ssh flit and MagicDNS expect.

Rotating credentials #

The pass-cli access token is the only one of these with a scheduled expiry — 90 days, described under "Secrets in a project". The other four have no schedule, and each has a different answer for how to rotate it:

  • Hetzner API token — create the replacement in the Hetzner Cloud console, then ./laptop/prepare-vault.sh --replace hcloud-token. It never reaches the server, so nothing downstream needs re-running.
  • healthchecks ping key — regenerate it on the same project page as the API key, ./laptop/prepare-vault.sh --replace healthchecks-ping-key, then push it to the server by re-running the phase that writes it: pass-cli run --env-file .env -- ./bootstrap.sh --target $FLIT_SERVER --start-at 10.
  • restic password — ./laptop/backup-credentials.sh --replace restic-password $FLIT_SERVER generates a replacement, adds it as a second key on the live repository while authenticating with the one already there, proves the replacement works, then promotes it in the vault and pushes it to the server, all in one run — the same laptop half phase 9 already runs. Reaching it through phase 9 instead uses bootstrap.sh's own --replace NAME flag, which only affects phase 9: pass-cli run --env-file .env -- ./bootstrap.sh --target $FLIT_SERVER --start-at 9 --replace restic-password. Removing the superseded key needs a confirmation prompt; declining it, or running --non-interactive, leaves both keys working rather than removing one that still works, and prints the removal command to run by hand later.
  • Storage Box ssh key — same script, same shape: ./laptop/backup-credentials.sh --replace storagebox-key $FLIT_SERVER, or pass-cli run --env-file .env -- ./bootstrap.sh --target $FLIT_SERVER --start-at 9 --replace storagebox-key. It generates a replacement, uploads it alongside the key already authorized, proves it authenticates, then promotes it in the vault. ssh-copy-id only appends, so it cannot retire the superseded key. Neither hcloud storage-box nor the Hetzner API exposes key management, but that is not the same as removal being manual: .ssh/authorized_keys is reachable over sftp like any other file on the box, so storagebox_withdraw_key() downloads it, removes the superseded key by matching its material rather than its comment (keys generated before storagebox_key_comment() existed all carry the identical comment, and a comment match would delete the wrong one), backs up the original to a timestamped path on the box, and re-uploads the filtered file, proving the current key still authenticates before returning.

Adding or changing a scheduled job #

Every timer and the healthchecks check that watches it come from one row of the JOBS table in lib/jobs.sh — its header comment has the name|OnCalendar|grace|cadence-suffix|command row format and what jobs_slug() derives from it. Add or edit a row there, commit it, then:

pass-cli run --env-file .env -- ./laptop/healthchecks.sh
pass-cli run --env-file .env -- ./bootstrap.sh --target $FLIT_SERVER --start-at 2 --skip-drills

The first upserts the check — keyed on slug, so existing checks are untouched — before anything can ping it. The second redelivers the repository (only phase 2 pushes lib/jobs.sh to the server) and re-applies every phase from 3 on, idempotently, including phase 10 (server/lib/70-timers.sh), which installs the new systemd unit and timer and enables it. --skip-drills skips phase 13, so a routine job change does not trigger a rebuild and a restore drill. Phase 12 skips itself once an automatic backup exists (spec §6.1).

Known limits #

No offline work: no connection means no development, traded away deliberately for a single source of truth — the same reasoning that rejected a second server in another region. A local clone would reintroduce exactly the "which machine has the uncommitted branch" problem this design exists to avoid; flights are reading time.

Watchers compete for RAM (see Daily use). Hetzner's 7-day rollback only covers "broke it just now" — Automatic Backups keep seven daily slots and verify.sh runs weekly, so drift can be older than anything the rollback reaches; restic and the rebuild drill are what actually cover the rest (spec §6.1). Disk fills quietly: Rust target/ directories and the pnpm store grow without bound. A monthly prune timer and a daily 80% check are configured (spec §5.6); if the alert fires, run cargo sweep and pnpm store prune before adding disk.

Electron tests run headless under xvfb-run, no display or forwarding needed, and Linux artifacts build natively. macOS builds are deferred (spec §11) — the intended shape is artifacts landing in a directory that syncs to the laptop, not built yet.


Development #

Everything below is about the repository rather than about running a box. flit-spec.md §13 has the change guards that any modification here is expected to follow; CLAUDE.md is the working brief.

Tests #

bash lib/common.test.sh
bash lib/reset-tailscale-clone.test.sh
find . -name '*.sh' ! -name '*.test.sh' -exec shellcheck -e SC1091 -e SC2154 -e SC2064 {} +

shellcheck is not a dependency of anything here, so its absence makes the find command above print nothing and exit 0 — a clean result and a green one look identical. Check command -v shellcheck before trusting one.

lib/ is the only code that runs on both laptop and server, so a change there also needs running under the server's userland, where /usr/bin/awk resolves to mawk rather than gawk:

docker run --rm -v "$PWD:/repo:ro" ubuntu:24.04 bash -c 'cd /repo; bash lib/common.test.sh'

No Terraform #

Provisioning is driven by the hcloud CLI, in bootstrap.sh phase 1 and in both drills — there is no state file to lose or to share. infra/README.md has the reasoning and the trigger for revisiting it.

Facts about a specific deployment — its region, instance size, resource names, key fingerprints, verification counts, and what has been drilled against it — are deliberately not part of this repository. A deployment keeps them in deployment.md, a local file gitignored for exactly this reason: the repository is public, and those facts are not meant to be.

Publishing this repository #