services in here:
- bobbin
- svfe
- hydrant
- seo-reporter
dev is configured as a pair of bobbin and svfe without anything in front. they get automatically deployed from next branch, see more on that below.
prod is configured as multiple bobbins and svfe's. bobbins have an LB in
front and so do the svfe's. its configured such that tangled.org will point
to nearest svfe, and api.tangled.org will point to nearest bobbin. this
way all the svfe configs are the same because they just take api.tangled.org.
prod is deployed via the deploy app on the nix flake, so just run
nix run .#deploy -- --target prod --ref <branch, tag or commit> (or
--path <checkout> to deploy a local tree as it is), optionally passing
--build-on user@host to use a specific host for building the images on.
to be able to work with terraform here, outside of gcloud auth login run
gcloud auth application-default login as well, because tf wants ADC.
layout #
main.tfis what infra manager applies. it picks an env out ofmodules/envsand builds it withmodules/environment.modules/envshas one file per env with all the 'actual' config. its basically the place to edit to add new hydrants or bobbins etc.modules/controlandcontrol/<env>are the part we apply ourselves, see below for why.- the
deployapp is../deploy.nu. modules/nodesandnodes/<env>are the nixos vms, see the end.modules/workersandworkers/<env>are the cloudflare workers, see the very end.
why there are two layers #
we want ci to deploy dev without handing it the keys to the whole project. so ci never runs terraform itself. it pushes images and asks infra manager to roll a new revision, and infra manager does the actual work with its own service account.
that service account can manage the services, network, registry and lb, but
it can't touch iam. if it could, anything that lands in the tree (and so
anything that can push to the branch) could grant itself more. so all the
identities and grants live in control/<env>, which only a project owner
applies, by hand:
nix run .#deploy -- --target dev --preview # shows what applying would change on live deployment
nix run .#deploy -- --target dev # applies it
don't apply main.tf manually. infra manager only adopts resources that
already exist while its state is clean, so anything you create behind its
back will become an annoyance for you later :p
deploying #
dev deploys on every push to sv-fe that touches bobbin or the web app
(.tangled/workflows/deploy-gcp-dev.yml).
prod runs pinned digests, not the :prod tag. a tag can move under you,
a digest is exactly what we built, and it makes a rollback just "pin the old
digests again". this is also why the control roots go through deploy
and not bare terraform. deploy passes along the pins that are already
there, a bare apply would drop them and the env would go back to its tag.
a build only builds images whose source changed. each image gets a
src-<hash> tag, hashed from its source (for bobbin, only the crates it depends
on, as the core flake narrows them) and its build args, and an image that
already has the tag gets reused. ci's images don't get one, so the first local
deploy after ci still builds. --only bobbin deploys
just that service, and --rebuild builds anyway, e.g. to pick up a newer base
image.
to roll back, run nix run .#deploy -- --target prod --rollback (or
dev) and pick a build.
rollback can only reach images the registry still has. prod keeps the last 16 builds, and anything older gets deleted after 90 days. dev keeps the last 8, and deletes the rest after 3 days.
how ci gets into gcp #
the spindle signs a short-lived token for each workflow run that asks
for one (with audience: in the workflow), and gcp trusts tokens from
https://tokens.spindle.tangled.sh through workload identity federation.
nothing long-lived to leak, rotate, or paste into a secret.
gcp only lets a token act as the ci deployer if it's for one exact repo and branch. note that checking just the repo isn't enough, a manually triggered run has no branch, so it would be able to deploy otherwise.
this is only turned on for dev (ci_oidc in modules/envs).
setting up a new env #
- enable these apis: config, run, compute, artifactregistry, iam, iamcredentials, storage, cloudresourcemanager. infra manager needs the last one to write iam, and the error you get without it doesn't say so.
- get an org admin to allow public members on the project. the org blocks
allUsersby default and project owners can't change that. prod needs this too, the lb can't authenticate to cloud run, so the services are public but only accept traffic from the lb. - make a versioned
gs://<project>-tfstatebucket for the control state. - add the env to
modules/envs, copycontrol/devtocontrol/<env>with its own bucket and env name, and apply it. that creates the deployment and infra manager builds everything else. if the first revision fails on permissions, wait a few minutes, new grants take a while to land.
gotchas #
- infra manager runs terraform 1.5.7, so don't put anything newer in
main.tfand its modules. - the lock file needs linux hashes too, since infra manager runs on linux:
terraform providers lock -platform=linux_amd64 -platform=darwin_arm64. - only one apply can run per deployment at a time. a second one fails with "unable to queue the operation".
nodes #
nodes/<env> runs hydrant on nixos vms. in prod they sit behind one global lb,
which sends each client to the closest healthy node. dev leaves the lb out
(lb = false), so its node is only reachable from inside the vpc and over ssh.
it's a plain terraform root that infra
manager never sees (the blueprint excludes it). deploy applies it after the
services, except for --rollback, and asks first when the plan would destroy anything
beyond redeploying a changed system. --no-nodes skips it,
and ci never touches it:
nix run .#deploy -- --target dev # services and nodes
terraform -chdir=terraform/nodes/dev apply # just the nodes, which have no pins to lose
a new vm boots debian, and nixos-anywhere
replaces it with nixosConfigurations.hyd-gcp, getting in over a throwaway key
that dies with debian. after that terraform deploys again whenever the system
changes, over your own ssh agent as tangler. every node is that one config with
its name passed in, and none of them are in the colmena hive.
- every plan builds the system, so you need nix and an x86_64-linux builder.
- the flake is read from git, so new files need at least
git add -Nfirst. - a replaced vm is reinstalled from scratch and starts its index over, so the debian image is ignored after creation.
hubsinmodules/envsgroups the cells around the node their bobbins read, so a new region goes into the hub whose hydrant it should use. the node itself lives in one of its hub's cells, and the vm is named after itsname.- bobbin can't get ready without its hydrant, so the blueprint leaves it out of a cell
until the hub's node exists (its
<node>-internaladdress, really). a new hub's first deploy makes the subnet and svfe, the nodes make the vm, anddeploythen rolls the blueprint once more to put bobbin in. with--no-nodes, the next deploy or ci push does. --previewcan't plan the nodes of a hub whose subnet doesn't exist yet, pass--no-nodesuntil its first deploy.- the nodes themselves are listed in
modules/envs. to grow a node's disk, raise itsdisk_gbthere and apply. the disk grows in place, and the node hands the space to hydrant on its next boot, or right away withsystemctl start grow-hydrant-volume.
workers #
workers/<env> deploys the cloudflare workers listed under workers in modules/envs,
each keyed by its directory in core. like the nodes it's a plain root that infra manager
never sees. terraform can't bundle typescript, so deploy builds every worker in the tree it
was given (pnpm install and pnpm run build, which leaves dist/index.js) and applies the
root with those bundles. that only happens with --ref or --path, --only workers deploys
just them, and ci never touches them. a worker's wrangler config is for wrangler dev, what
runs is whatever modules/envs says.
control makes the google side: the secrets the workers read, and for a worker with a
google_identity, a service account <worker>-<suffix> and a key for it, which lands in
<worker>-google-key and from there in the binding. its apis get enabled there too.
the first deploy with a worker in it only runs control, because the list of workers comes
from control's last apply. fill in the secrets it made before the next one, without a
trailing newline. cloudflare-workers-token wants workers scripts edit on the account.
printf %s "$VALUE" | gcloud secrets versions add <secret> --data-file=- --project exemplary-proxy-507408-d6
- the key sits in control's state, so the state bucket is as secret as the key. to rotate it,
terraform -chdir=terraform/control/dev taint 'module.control.google_service_account_key.worker["<worker>"]'and deploy with--ref/--path(a bare apply of control would drop the image pins). - if the org enforces
iam.disableServiceAccountKeyCreation, the key needs an exception for the project. --previewdoesn't plan the workers, since that would mean building them.