flit #
Provisioning for one remote development machine.
A laptop pushes configuration to a single Hetzner server; the server never fetches its own. All work then happens on the server — the laptop is a terminal, a network stack, and a key. One machine is the whole design, not a starting point: a single source of truth is what removes the "which machine has the uncommitted branch" problem, and every tradeoff below follows from it.
Three documents, with no overlap:
- This file — how to build the box, use it daily, and operate it.
flit-spec.md— the design and the reasoning behind it. Section numbers (§4.2, §6.2, …) referenced throughout this file point there.CLAUDE.md— the brief for an agent making changes to this repository.
How it works #
bootstrap.sh RUNS ON THE LAPTOP — provision, push, execute
verify.sh read-only drift detection (PASS/FAIL/SKIP)
drill-restore.sh quarterly — can the work be recovered?
drill-rebuild.sh twice-yearly — do the scripts reproduce the box?
lib/ shared primitives + tests
laptop/ anything needing a laptop-only credential
server/lib/ numbered provisioning steps, run in order over ssh
server/bin/ the jobs the timers run
fixtures/ minimal projects the drills build against
Credentials flow one way. Every secret is held by the laptop and pushed;
the server never reaches back for its own configuration. A script's directory
is therefore a claim about which machine can run it.
laptop/healthchecks.sh and laptop/backup-credentials.sh stay out of
server/lib/ because each needs something only the laptop holds — the
healthchecks read-write API key, and an authenticated pass-cli session,
respectively. A script under server/lib/ would be structurally unable to do
that work. Checking all six dimensions of execution context — machine, user,
environment, binary, userland, direction — before moving anything across that
laptop/server line is spec §13.2's standing rule, and the most repeated defect
in the project.
Commands below are written to be pasted as-is, so they use default values
rather than the variables that produce them. Those variables are defined in
lib/common.sh under "THE VARIABLE SET"; a deployment that overrode one
substitutes its own value, and records what it holds in its own
deployment.md.
| Appears below as | Variable | What it names |
|---|---|---|
flit |
FLIT_ALIAS |
the ssh Host alias, and the flit/flits shell helpers |
flit |
FLIT_NAME |
the project name, in paths like /etc/flit/ and in job names |
$FLIT_SERVER |
FLIT_SERVER |
the Hetzner server and ssh HostName — deployment-specific, so left as a variable |
$FLIT_USER |
FLIT_USER |
the operator account on the server — likewise |
What it needs #
Tools on the laptop. bootstrap.sh preflight() requires ssh, rsync,
jq, curl, hcloud (Hetzner's CLI), pass-cli, and sshpass (which uploads
the Storage Box key non-interactively in phase 9). preflight() dies naming the
first one missing, so installing them up front is a convenience, not a trap.
Accounts. laptop/prepare-vault.sh collects credentials but cannot sign up
for anything, so these exist first (spec §12.2 has the reasoning for why each
stays a manual, web-UI step):
- Hetzner Cloud project.
- Tailscale account and tailnet. Generate two auth keys in the admin
console before running
prepare-vault.sh: one for the persistent production node, one with the Ephemeral toggle set for drills. Mark both Reusable — the console default is single-use, and authenticates once before failing silently on the next rebuild or drill. - healthchecks.io and ntfy accounts, with one ntfy and one email integration attached at healthchecks project level. Phase 8 queries the API and aborts if either is missing.
- Proton Pass account and vault.
Repositories delivered alongside this one. Phase 2 ships ~/.claude to the
server with git archive HEAD, so it must exist as a git repository with
nothing uncommitted, or the phase refuses to run. ../dotfiles is delivered
too when it exists, and skipped when it does not. Chook is not delivered here:
it reaches the box as a normal project, pushed with laptop/push-project.sh
like any other, carrying its own .git so work done on it there has somewhere
to be committed (see "Moving a project onto the box").
From a clone to a working box #
Activate the pre-commit hook. bootstrap.sh preflight() runs
git config core.hooksPath .githooks on every invocation, so the hook that
rejects a literal secret in a staged .env is active from the first run. Git
does not carry hooks across a clone, so on a clone where bootstrap has not run
yet, set it by hand first:
git config core.hooksPath .githooks
Fill the vault.
pass-cli login
./laptop/prepare-vault.sh
pass-cli login opens an authenticated session; writes need it (spec §6.2), and
so does prepare-vault.sh. That script is the one thing here not run under
pass-cli run --env-file .env: that form resolves every reference up front, and
on a first run the items it needs do not exist yet. It prompts for five values,
mints a sixth itself (pass-cli-token, scoped read-only), and generates a
seventh (fixture-secret), storing each at the path .env already references.
./laptop/prepare-vault.sh --print-only reports which of the seven are present
and which are missing, and mutates nothing.
.env needs no editing. Every value in it is already a pass://
reference, which is what lets this repository be public — the pre-commit hook
rejects a staged .env holding a literal. It can only reference items that
exist before bootstrap.sh starts, so it names neither the Storage Box
credentials nor restic-repository/restic-password; phase 9 creates those and
reads them with vault_get instead.
A second deployment needs its own .env: pass-cli run --env-file reads each
pass:// prefix literally, so overriding FLIT_NAME alone does not redirect
them — a .env still reading pass://flit/... would resolve every credential
against the first deployment's vault instead of the new one. vault_check_env()
(lib/common.sh) checks this before bootstrap.sh, both drills, and
laptop/healthchecks.sh touch the vault, and fails loudly on a mismatch rather
than reading the wrong vault silently.
local.conf is optional. Copy local.conf.example to local.conf —
gitignored, laptop-only. It matters in two cases: the ssh agent holds more than
one key (resolve_operator_key() refuses to guess among them, so
HCLOUD_SSH_PUBKEY, set in local.conf or exported per invocation, makes the
choice explicit), or the default HCLOUD_LOCATION (fsn1) has no capacity for
the instance size at the time.
Dry run, then the real run.
pass-cli run --env-file .env -- ./bootstrap.sh --print-only
pass-cli run --env-file .env -- ./bootstrap.sh
--print-only mutates nothing anywhere — laptop config is inspected exactly,
server phases print a plan only. The real run needs three things done by hand
along the way:
- Disable node key expiry in the Tailscale admin console when phase 3 demands it. Skipping it schedules a lockout roughly six months out, with public SSH already closed (spec §4.2).
- Authenticate Claude Code through its browser flow after phase 4 installs it.
- Answer phase 8's interactive prompt, which asks whether both the ntfy push and the email arrived. The run continues once it is answered.
When a phase fails, resume with --start-at N. The failed instance is left
running deliberately — it is the only evidence of what went wrong — and so is
the log of the run, at ~/.local/state/flit/bootstrap-latest.log (see "The log
of a run"). For N greater than 2, check_delivered_commit() refuses the
resume outright when the server's .delivered-commit does not match the
laptop's HEAD, rather than silently running server-side code that the fix being
resumed toward was never shipped to.
To ship an edit to a server/lib/ or server/bin/ script without a full
provisioning run, pair --start-at with --stop-at:
pass-cli run --env-file .env -- ./bootstrap.sh --target $FLIT_SERVER --start-at 2 --stop-at 2
./laptop/sync-config.sh
--stop-at N runs no phase after N. --start-at 2 --stop-at 2 delivers the
repository and stops, which is only phase 2 — no later phase runs, so nothing
else on the box changes. laptop/sync-config.sh then runs the box's delivered
copy of server/lib/22-claude-config.sh to pick up the change. --stop-at
before --start-at is refused, as is a non-numeric value — read as "no limit"
would defeat the point of the flag. Stopping early logs that the box is not
fully provisioned, rather than the usual completion line.
A successful first run leaves the box without chook. Nothing in
bootstrap.sh requires or delivers it, so the checkout that every hook in the
installed settings.json points at does not exist yet. Pushing it is the next
step, and is covered as the first example under "Moving a project onto the
box" in Daily use, below.
Daily use #
ssh flit reaches the box because laptop/install.sh writes a Host flit
block mapping the alias to HostName $FLIT_SERVER and User $FLIT_USER — that
indirection is the point, since the alias stays stable while the server it
points at is free to change. The same script adds two shell functions:
$ flit api # attaches or creates session 'api', cd'd to ~/workspace/api
$ flit web # completely separate session
$ flits # list every running session
One tmux session per project is the primary way to use the box, not an edge
case. Each session has its own windows, panes, working directories, and
running processes, and sessions do not see each other — three terminals in three
projects means three independent sets of watchers and language servers, all
surviving disconnection independently. new-session -A attaches if the session
exists and creates it otherwise, so one command covers both cases. Detach with
Ctrl-b d, or just close the terminal; the session keeps running either way.
Attaching two terminals to the same session mirrors them instead of giving independent views, and tmux sizes windows to the smallest attached client. That is a feature — phone-to-desktop handoff works this way — but it surprises anyone expecting independence. Independence comes from different session names, not different connections.
Each project session with a running dev server holds a tsserver at 2–6 GB plus
watchers, so watchers, not sessions, are the memory cost: four active TypeScript
projects can exhaust 32 GB on their own. Kill dev servers in projects not
actively being worked on rather than leaving five running because detaching is
free. flits plus htop diagnose things once they feel slow.
Moving a project onto the box #
Chook is the first project pushed onto a newly provisioned box, before
using agents there at all. The installed settings.json wires every Claude
Code hook to chook.mjs under ~/workspace/chook, and each one fails until
that checkout exists:
$ ./laptop/push-project.sh ~/misc/chook
$ ./laptop/push-project.sh ~/misc/myproject # -> ~/workspace/myproject
$ ./laptop/push-project.sh --print-only ~/misc/myproject
$ ./laptop/push-project.sh ~/misc/myproject other-name
Cloning the repository on the server would be shorter and would lose the half
of a working directory that git does not track: a gitignored .env, a
handoff.md, .claude/settings.local.json, an edit made and not committed.
This copies the directory instead, so what arrives is the working directory
rather than a fresh checkout of it. Claude Code's session transcripts for the
project come too, re-keyed from the laptop's path to the server's, unless
--no-sessions.
What does not come is anything the server can rebuild from something that did:
node_modules, target/, dist/, .venv, .wrangler, build caches. Those
are most of the bytes, and a node_modules built on an arm64 laptop is the
wrong one for an amd64 server anyway. A project needing more exclusions lists
them, one pattern per line, in a .flit-push-exclude file in its own root.
That list matches directory names, and a name can be wrong: a repository is free
to track a source file under build/ or dist/. Anything git tracks is carried
regardless, so a blunt pattern cannot strand it. Your own .flit-push-exclude
still outranks that — it is an explicit decision about this project, and a
repository may well track a data set the server has its own copy of. When one of
your patterns does leave a tracked file behind, the push says which files and
what to do about it, because the server's checkout reads as dirty from then on
and pushing again will not fix it.
The push is one-way, like everything else here: the laptop is authoritative and
the server never reaches back. It is not --delete by default, so a file that
exists only on the server survives; pass --delete to make the server an exact
mirror. If the server's copy has uncommitted changes, the push lists them and
asks before overwriting — --force answers yes in advance. Work done on the
box is real work, and this is the one thing that can quietly destroy it.
A project can hold the push off entirely with a .flit-no-push file in its
root. This covers the case the dirty-tree guard cannot see: an agent working
on the server that commits its work leaves a clean tree, and a clean tree that
is ahead of the laptop looks exactly like one that is behind. push-project.sh
refuses before transferring anything and prints the file's first five lines as
the reason. --force does not lift the hold — it exists to override the
dirty-tree guard, not this one — and --print-only is refused too, since even
a dry run against a project on hold has nothing useful to report. Deleting the
file is what lifts it.
Then open it:
$ flit myproject
Holding syncing off entirely #
A .sync-off file at the repository root refuses every laptop-to-server
write, not just one project's. bootstrap.sh, laptop/sync-config.sh and
laptop/push-project.sh all overwrite rather than merge, so one switch covers
all three; check_sync_hold() in lib/common.sh checks it once, at startup,
rather than in each script, so it cannot be lifted halfway. Create it with the
reason as its content:
$ echo "rebasing the server's dotfiles by hand, do not overwrite" > .sync-off
The refusing script prints the file's first five lines back, which is what makes the reason worth writing down — it is what shows up weeks later when something else hits the hold. Delete the file to resume:
$ rm .sync-off
This is coarser than .flit-no-push: that file holds off pushes to one
project, this one stops provisioning and the configuration sync too. There is
no flag to override it — push-project.sh --force already overrides the
dirty-tree check, and .sync-off exists for exactly the case that check
cannot see, so overriding it the same way would defeat the point. Deleting the
file is the only way to lift the hold, and doing so is a decision worth
leaving a trace of, unlike a flag on a command line.
Secrets in a project #
Commit .env files holding pass:// references, never values:
DB_PASSWORD=pass://Dev/Database/password
Run anything needing them through pass-cli:
pass-cli run --env-file .env -- pnpm dev
First resolution takes ~5 s; subsequent calls are instant from the kernel
keyring, expiring after an hour or at logout. The access token expires every 90
days and fails during interactive work, in the morning, with a clear error —
deliberate timing (spec §7.3). Rotate it with
./laptop/prepare-vault.sh --replace pass-cli-token, which mints a new token
through pass-cli personal-access-token create and re-escrows it in one step
rather than leaving the vault holding the expired one. Backups do not depend on
this token, so they keep running through an expiry. The quarterly drill lands at
roughly 91 days and so usually wants a fresh token too — in practice, refreshing
the token and running the drill are the same quarterly ritual.
Clipboard #
Ghostty owns the macOS pasteboard; vim runs on the server. OSC 52 carries a yank
back over the terminal byte stream — yank → escape sequence → tmux → SSH →
Ghostty → macOS clipboard — which is why both Ghostty (copy-on-select,
clipboard-read/write, set by laptop/install.sh) and tmux
(allow-passthrough, set-clipboard, set by server/lib/25-terminal.sh in
phase 6) need configuring; tmux is not a dumb pipe and has to be told to forward
the sequence through.
tmux is set to set-clipboard external, not on, deliberately: under on, any
process running inside tmux — including code from an unfamiliar repository
running under Claude Code — can set the system clipboard regardless of which
user runs it, and a clipboard write followed by a shell paste is a real attack.
external restricts clipboard-setting to tmux itself, so copying goes through
copy-mode rather than vim's yank register:
Ctrl-b [ enter copy-mode
Space to start selection, motion keys, Enter to copy
Paste needs none of this — Cmd-V works natively; OSC 52 paste is deliberately
awkward for security reasons and is not worth using. Large copies vanishing
while small ones land is an OSC string length limit that terminals impose and
drop silently, not a broken config.
Claude Code renderer — stay on classic #
server/lib/25-terminal.sh (phase 6) sets
CLAUDE_CODE_DISABLE_ALTERNATE_SCREEN=1 on the server, where Claude Code
actually runs; setting it on the laptop does nothing. The fullscreen renderer
draws into the alternate screen buffer, moving the conversation out of native
scrollback — replacing Cmd+f and tmux search with Ctrl+o transcript mode —
and out of reach of tmux copy-mode, which is the clipboard path above. The
classic renderer keeps everything in native scrollback, where copy-mode works
normally.
Fullscreen virtualizes the viewport and transmits only changed regions, cutting
ANSI escape data — real savings on a slow link. Classic still wins: losing
clipboard access costs more every day than the bandwidth costs on occasional
slow links. To use fullscreen anyway, set CLAUDE_CODE_DISABLE_MOUSE=1
alongside it, since fullscreen's mouse capture fights copy-mode directly. It
cannot be escaped entirely either way: background sessions opened from agent
view or claude attach always use fullscreen, and the env var does not apply to
them.
Terminal type #
bootstrap.sh's push_terminfo() (phase 6, right after
server/lib/25-terminal.sh) pushes the laptop's terminal description to the
server, because Ubuntu 24.04's terminfo database predates several terminals in
current use — xterm-ghostty arrived after it — and ssh forwards $TERM
verbatim, so without the push tmux exits with missing or unsuitable terminal
the first time flit runs. The push writes ~/.terminfo rather than the
system database, so it needs no sudo, and only the laptop is touched for
reading: it knows $TERM, the server does not. If it ever needs redoing by
hand:
infocmp -x "$TERM" | ssh flit -- tic -x -
Browser previews #
Dev servers bind on the box; Tailscale reaches them from the laptop browser directly:
http://$FLIT_SERVER:5173
No tunnel, no port forwarding, no -L flags. Bind to 0.0.0.0, not
127.0.0.1, or the tailnet cannot reach it — Vite needs --host 0.0.0.0 or
server.host: true. This is per-project, so it is the one piece of the setup
above that the scripts cannot do.
Travel #
ssh flit, same as always — there is no second transport to remember, since
mosh was removed deliberately (spec §10), so git push works from every session
and nothing about the workflow changes across a trip. What changes is typing: at
a distance of, say, Asia from the server's datacenter, the round trip shows up in
vim insert mode, roughly 180 ms per keystroke. Normal-mode motions and commands
are unaffected; sustained prose entry is the unpleasant part. In order of
usefulness: compose long prose locally and paste it in (Cmd-V works natively
into insert mode); stay in normal mode, since motions, ., macros, and :
commands send one round trip for a whole operation rather than one per character;
run builds and tests as usual, since only interactive echo is latency-bound.
A dropped connection costs a reconnect, never state — tmux holds everything, and
ControlPersist makes reattaching near-instant. One side effect of that
multiplexing: changing networks while a master connection is open leaves the
socket stale, and the next ssh flit hangs rather than erroring;
ssh -O exit flit clears it. There is no offline mode (see Known limits), so
push everything before departure.
Checking and troubleshooting #
verify.sh is read-only drift detection, reporting PASS/FAIL/SKIP. It runs on
the server: bootstrap.sh phase 11 invokes it there over ssh, a weekly timer
runs it there unattended, and it can be run there by hand at
~/flit/verify.sh. Remediation for anything it reports is
bootstrap.sh --target $FLIT_SERVER from the laptop — configuration is pushed,
never edited on the box.
Alerts #
Everything on this box runs on a timer, and every timer checks in with healthchecks.io. It fans out to two channels:
job on $FLIT_SERVER → healthchecks.io → ntfy push + email
The ntfy topic is meant to be subscribed to on phone and desktop. Nothing publishes to ntfy from the server — the token lives at healthchecks, which is why it is not among the credentials on the box. The topic name is high-entropy on purpose: public ntfy topics are readable by anyone who guesses the name, and these messages carry the hostname, disk state, and backup outcomes.
| Alert | Meaning |
|---|---|
verify.sh drift |
Something diverged from the scripts — run bootstrap.sh --target $FLIT_SERVER from the laptop; configuration is pushed, never edited on the box |
| Backup failed | restic errored to the Storage Box; failure output is in the healthchecks dashboard |
| Rebuild drill diff non-empty | Packages on the live box are not in server/lib/10-packages.sh's package list — add them there, then re-run bootstrap.sh --target $FLIT_SERVER |
| Node pin stale | ~2 Electron majors behind. Bump server/files/mise.toml, then bootstrap.sh --target $FLIT_SERVER. Routine divergence never alerts — it just shows in the dashboard |
| Project config check failed | The throwaway SessionStart hook did not fire through a symlinked .claude; quota/session-limit failures after it fires do not fail the check — see server/bin/check-project-config.sh |
| Backup credentials missing from vault | Urgent. Re-escrow immediately — backups are otherwise undecryptable if the box dies (see "If the box is gone") |
verify.sh reports SKIP, not FAIL |
A check could not run — expired token, network down. Not drift; no action unless it repeats |
| Disk 80% | Run cargo sweep and pnpm store prune |
… is DOWN |
A scheduled job did not run at all — see below |
| Drill DOWN | No drill has run this quarter — run it from the laptop (spec §8) |
Read DOWN differently from failed. Because everything is a check-in, healthchecks can alert on a job that never ran, not just one that ran badly (spec §6.3): "failed" means the job ran and something went wrong, usually narrow and obvious; "DOWN" means the job never ran — the timer died, the box is unreachable, or something upstream broke, a bigger question than whichever job's name is attached to it. A DOWN on the daily backup after a reboot is the case worth catching, since nothing else in this design would report it.
Every alert also goes to email, deliberately — email is slow and easy to ignore, fine for a backstop, wrong for a primary; if ntfy itself is broken, notification failure would otherwise be indistinguishable from everything being fine. The two drill checks watch whether a drill ran, not the server: drills run from the laptop, since they need the Hetzner API token, which deliberately never touches the box (spec §8), so a drill DOWN means no drill has run this quarter, not that something broke — running it clears the check. Still not covered: healthchecks.io going down, unnoticed. That is where the layering stops; the quarterly restore drill is the backstop for the backstop.
When something looks wrong #
| Symptom | Check |
|---|---|
$FLIT_SERVER will not connect |
Tailscale up on both ends — public SSH is closed by design. If Tailscale itself is the problem, use break-glass (below) |
| Tailscale key expiry warning | Act on it. When the key lapses the box leaves the tailnet, locking out SSH access (spec §4.2) |
pass-cli fails |
Token expired at 90 days — rotate it with ./laptop/prepare-vault.sh --replace pass-cli-token |
| Copy does not reach macOS | vim yank? Use tmux copy-mode — set-clipboard external blocks inner apps by design (see Clipboard) |
| Large copies vanish, small ones work | OSC string length limit, not config (see Clipboard) |
| Claude Code scrollback gone | Env var set on laptop instead of server (see Claude Code renderer) |
| Browser cannot reach dev server | Bound to 127.0.0.1 instead of 0.0.0.0 (see Browser previews) |
flit / tmux exits missing or unsuitable terminal |
Server terminfo lacks the laptop's $TERM (see Terminal type); by hand: infocmp -x "$TERM" | ssh flit -- tic -x - |
flit <name> opens in $HOME, not ~/workspace/<name> |
Old helper left the tilde quoted, so no shell expanded it, and nothing created the directory — tmux falls back to $HOME silently rather than failing (see Daily use) |
| Builds slow, machine sluggish | Too many watchers (see Daily use), or disk near full (see Known limits) |
verify.sh alert |
Drift from the scripts — run bootstrap.sh --target $FLIT_SERVER from the laptop; configuration is pushed, never edited on the box |
ssh flit hangs after changing networks |
Stale multiplexed socket. ssh -O exit flit, then reconnect |
| ntfy alert: Node pin stale | ~2 Electron majors behind. Bump server/files/mise.toml, then bootstrap.sh --target $FLIT_SERVER (spec §5.4) |
| Backup DOWN after a reboot | Timer did not come back — check systemctl list-timers (see Alerts) |
| A setting from Clipboard or Claude Code renderer seems absent | Did install.sh defer it? Re-run with --print-only and check the deferred list |
bootstrap.sh or a drill failed |
Read the log, not the scrollback — see below |
The log of a run #
bootstrap.sh, drill-restore.sh and drill-rebuild.sh each write every line
either stream carries to:
~/.local/state/flit/<name>-<timestamp>.log
with <name>-latest.log pointing at the run in progress, so a long
bootstrap.sh can be watched from a second terminal:
tail -f ~/.local/state/flit/bootstrap-latest.log
A failed run names its own log path before it exits. --print-only writes no
log at all, since it mutates nothing anywhere. Ordering within one stream is
exact; between stdout and stderr it is not, so a line that looks out of place
next to its neighbour probably came from the other stream. The files are 0600
and are never pruned — the drills each keep two runs a year, and a failed
bootstrap is worth more than the bytes.
Recovery and rebuild #
If the box is gone #
Restoring recovers the work; it does not restore the credentials the machine
needs to function. Four things must be re-supplied by hand — the full procedure
is spec §8.3. Short version: create a new pass-cli token in Proton Pass first
(nothing else works without it), provision a replacement, write
/etc/flit/secrets.env, restore the Storage Box SSH key, then run verify.sh
and let it report what is still missing.
The restic password and Storage Box key live in Proton Pass, escrowed there when
the box was built (spec §6.2) — the copies on the server are a cache, which is
what makes recovery possible at all. verify.sh re-checks that both are still
in the vault on every run; if that alert ever fires, treat it as urgent rather
than cosmetic, since it means the backups are one disk failure from being
permanently undecryptable. Proton Pass is the single point of failure for this
whole design — the account's recovery codes belong somewhere that is not this
machine.
Break-glass #
If $FLIT_SERVER is unreachable — Tailscale down, key expired, an
over-restrictive firewall rule — Hetzner's Cloud Console has a browser terminal
that bypasses SSH entirely. The $FLIT_USER password is in Proton Pass as
$FLIT_USER-console-password, put there at build time for exactly this; it does
not work over SSH by design, only in that console. Confirm it is actually in the
vault before an outage is the first occasion to need it.
Drills #
pass-cli run --env-file .env -- ./drill-restore.sh
pass-cli run --env-file .env -- ./drill-rebuild.sh
Assertion is not evidence, so each drill proves something a green
bootstrap.sh run does not. drill-restore.sh (spec §8.1, quarterly) proves
the work is recoverable, restoring the most recent Hetzner automatic backup —
or, with --from-snapshot, the golden snapshot, which is the only source
available before any automatic backup exists. drill-rebuild.sh (spec §8.2,
twice-yearly) proves the scripts reproduce the box from a stock image.
Step zero of both is a precondition check: drill_precondition()
(lib/drill-common.sh) asserts the pass-cli token still works before
anything is provisioned, because the token's 90-day expiry and the drill's
quarterly cadence collide by construction, and an expiry discovered several
minutes and one provisioned CX53 into a run reads like a broken restore path
rather than routine credential expiry (spec §8).
Each drill destroys its instance on every path, including failure, and pings its own healthchecks slug with the exit code.
Four flags always travel together: --skip-drills --non-interactive --no-external-state --skip-laptop. drill-rebuild.sh passes all four to
bootstrap.sh itself, and any other invocation of bootstrap.sh against a
throwaway instance needs the same set — omitting any one fails differently, and
none fails loudly (spec §12.1).
Tearing down and recreating production #
hcloud server delete $FLIT_SERVER removes only the compute instance.
Re-running
pass-cli run --env-file .env -- ./bootstrap.sh
recreates it: phase 1 finds no server named $FLIT_SERVER, creates a fresh CX53
from the same cloud-init, registers the operator SSH key, and mints a new
console password, overwriting the escrowed one. Phase 2 onward reprovision it
exactly as a first build would, reusing whatever the vault already holds rather
than starting over — phase 9 finds the Storage Box already exists and pushes the
existing restic-password and storagebox-key onto the new box, so backups
resume against the same repository, and phase 8 upserts the same eight
healthchecks slugs rather than creating duplicates.
Three things are not cleaned up by any of this, and need separate attention:
- The Storage Box is untouched either way — a separate Hetzner resource, not
part of the compute instance, and
laptop/provision-storagebox.shonly creates one when none exists. - The healthchecks checks are untouched; they go DOWN while the old server is gone and recover once the new one's timers start pinging again.
- The old Tailscale node is not deregistered by
hcloud server delete. Nothing here does that either —drill_cleanup()(lib/drill-common.sh) only ever destroys the Hetzner instance a drill created, and a drill's own node uses the ephemeral key and self-removes; a torn-down production node used the persistent key and does not. Remove the old device from the Tailscale admin console before or right after rebuilding, or the new instance — booting under the same hostname — registers under a suffixed name instead of the onessh flitand MagicDNS expect.
Rotating credentials #
The pass-cli access token is the only one of these with a scheduled expiry —
90 days, described under "Secrets in a project". The other four have no
schedule, and each has a different answer for how to rotate it:
- Hetzner API token — create the replacement in the Hetzner Cloud console,
then
./laptop/prepare-vault.sh --replace hcloud-token. It never reaches the server, so nothing downstream needs re-running. - healthchecks ping key — regenerate it on the same project page as the API
key,
./laptop/prepare-vault.sh --replace healthchecks-ping-key, then push it to the server by re-running the phase that writes it:pass-cli run --env-file .env -- ./bootstrap.sh --target $FLIT_SERVER --start-at 10. - restic password —
./laptop/backup-credentials.sh --replace restic-password $FLIT_SERVERgenerates a replacement, adds it as a second key on the live repository while authenticating with the one already there, proves the replacement works, then promotes it in the vault and pushes it to the server, all in one run — the same laptop half phase 9 already runs. Reaching it through phase 9 instead usesbootstrap.sh's own--replace NAMEflag, which only affects phase 9:pass-cli run --env-file .env -- ./bootstrap.sh --target $FLIT_SERVER --start-at 9 --replace restic-password. Removing the superseded key needs a confirmation prompt; declining it, or running--non-interactive, leaves both keys working rather than removing one that still works, and prints the removal command to run by hand later. - Storage Box ssh key — same script, same shape:
./laptop/backup-credentials.sh --replace storagebox-key $FLIT_SERVER, orpass-cli run --env-file .env -- ./bootstrap.sh --target $FLIT_SERVER --start-at 9 --replace storagebox-key. It generates a replacement, uploads it alongside the key already authorized, proves it authenticates, then promotes it in the vault.ssh-copy-idonly appends, so it cannot retire the superseded key. Neitherhcloud storage-boxnor the Hetzner API exposes key management, but that is not the same as removal being manual:.ssh/authorized_keysis reachable over sftp like any other file on the box, sostoragebox_withdraw_key()downloads it, removes the superseded key by matching its material rather than its comment (keys generated beforestoragebox_key_comment()existed all carry the identical comment, and a comment match would delete the wrong one), backs up the original to a timestamped path on the box, and re-uploads the filtered file, proving the current key still authenticates before returning.
Adding or changing a scheduled job #
Every timer and the healthchecks check that watches it come from one row of the
JOBS table in lib/jobs.sh — its header comment has the
name|OnCalendar|grace|cadence-suffix|command row format and what jobs_slug()
derives from it. Add or edit a row there, commit it, then:
pass-cli run --env-file .env -- ./laptop/healthchecks.sh
pass-cli run --env-file .env -- ./bootstrap.sh --target $FLIT_SERVER --start-at 2 --skip-drills
The first upserts the check — keyed on slug, so existing checks are untouched —
before anything can ping it. The second redelivers the repository (only phase 2
pushes lib/jobs.sh to the server) and re-applies every phase from 3 on,
idempotently, including phase 10 (server/lib/70-timers.sh), which installs the
new systemd unit and timer and enables it. --skip-drills skips phase 13, so a
routine job change does not trigger a rebuild and a restore drill. Phase 12
skips itself once an automatic backup exists (spec §6.1).
Known limits #
No offline work: no connection means no development, traded away deliberately for a single source of truth — the same reasoning that rejected a second server in another region. A local clone would reintroduce exactly the "which machine has the uncommitted branch" problem this design exists to avoid; flights are reading time.
Watchers compete for RAM (see Daily use). Hetzner's 7-day rollback only covers
"broke it just now" — Automatic Backups keep seven daily slots and verify.sh
runs weekly, so drift can be older than anything the rollback reaches; restic and
the rebuild drill are what actually cover the rest (spec §6.1). Disk fills
quietly: Rust target/ directories and the pnpm store grow without bound. A
monthly prune timer and a daily 80% check are configured (spec §5.6); if the
alert fires, run cargo sweep and pnpm store prune before adding disk.
Electron tests run headless under xvfb-run, no display or forwarding needed,
and Linux artifacts build natively. macOS builds are deferred (spec §11) — the
intended shape is artifacts landing in a directory that syncs to the laptop, not
built yet.
Development #
Everything below is about the repository rather than about running a box.
flit-spec.md §13 has the change guards that any modification here is expected
to follow; CLAUDE.md is the working brief.
Tests #
bash lib/common.test.sh
bash lib/reset-tailscale-clone.test.sh
find . -name '*.sh' ! -name '*.test.sh' -exec shellcheck -e SC1091 -e SC2154 -e SC2064 {} +
shellcheck is not a dependency of anything here, so its absence makes the
find command above print nothing and exit 0 — a clean result and a green one
look identical. Check command -v shellcheck before trusting one.
lib/ is the only code that runs on both laptop and server, so a change there
also needs running under the server's userland, where /usr/bin/awk resolves to
mawk rather than gawk:
docker run --rm -v "$PWD:/repo:ro" ubuntu:24.04 bash -c 'cd /repo; bash lib/common.test.sh'
No Terraform #
Provisioning is driven by the hcloud CLI, in bootstrap.sh phase 1 and in
both drills — there is no state file to lose or to share. infra/README.md has
the reasoning and the trigger for revisiting it.
Facts about a specific deployment — its region, instance size, resource names,
key fingerprints, verification counts, and what has been drilled against it —
are deliberately not part of this repository. A deployment keeps them in
deployment.md, a local file gitignored for exactly this reason: the
repository is public, and those facts are not meant to be.