Tangled infrastructure definitions in Nix
README.md

services in here:

  • bobbin
  • svfe

dev is configured as a pair of bobbin and svfe without anything in front. they get automatically deployed from next branch, see more on that below.

prod is configured as multiple bobbins and svfe's. bobbins have an LB in front and so do the svfe's. its configured such that tangled.org will point to nearest svfe, and api.tangled.org will point to nearest bobbin. this way all the svfe configs are the same because they just take api.tangled.org. prod is deployed via the deploy-gcp app on the nix flake, so just run nix run .#deploy-gcp -- --target prod --ref <branch, tag or commit> (or --path <checkout> to deploy a local tree as it is), optionally passing --build-on user@host to use a specific host for building the images on.

to be able to work with terraform here, outside of gcloud auth login run gcloud auth application-default login as well, because tf wants ADC.

layout #

  • main.tf is what infra manager applies. it picks an env out of modules/envs and builds it with modules/environment.
  • modules/envs has one file per env with all the values. everything else reads from here, so adding an env or changing a region is one file.
  • modules/control and control/<env> are the part we apply ourselves, see below for why.
  • the deploy-gcp app is ../deploy-gcp.nu.

why there are two layers #

we want ci to deploy dev without handing it the keys to the whole project. so ci never runs terraform itself. it pushes images and asks infra manager to roll a new revision, and infra manager does the actual work with its own service account.

that service account can manage the services, network, registry and lb, but it can't touch iam. if it could, anything that lands in the tree (and so anything that can push to the branch) could grant itself more. so all the identities and grants live in control/<env>, which only a project owner applies, by hand:

nix run .#deploy-gcp -- --target dev --preview   # shows what applying would change on live deployment
nix run .#deploy-gcp -- --target dev --no-build  # applies it

don't apply main.tf manually. infra manager only adopts resources that already exist while its state is clean, so anything you create behind its back will become an annoyance for you later :p

deploying #

dev deploys on every push to sv-fe that touches bobbin or the web app (.tangled/workflows/deploy-gcp-dev.yml).

prod runs pinned digests, not the :prod tag. a tag can move under you, a digest is exactly what we built, and it makes a rollback just "pin the old digests again". this is also why the control roots go through deploy-gcp and not bare terraform. --no-build passes along the pins that are already there, a bare apply would drop them and the env would go back to its tag.

to roll back, run nix run .#deploy-gcp -- --target prod --rollback (or dev) and pick a build.

rollback can only reach images the registry still has. prod keeps the last 16 builds, and anything older gets deleted after 90 days. dev keeps the last 8, and deletes the rest after 3 days.

how ci gets into gcp #

the spindle signs a short-lived token for each workflow run that asks for one (with audience: in the workflow), and gcp trusts tokens from https://tokens.spindle.tangled.sh through workload identity federation. nothing long-lived to leak, rotate, or paste into a secret.

gcp only lets a token act as the ci deployer if it's for one exact repo and branch. note that checking just the repo isn't enough, a manually triggered run has no branch, so it would be able to deploy otherwise.

this is only turned on for dev (ci_oidc in modules/envs).

setting up a new env #

  1. enable these apis: config, run, compute, artifactregistry, iam, iamcredentials, storage, cloudresourcemanager. infra manager needs the last one to write iam, and the error you get without it doesn't say so.
  2. get an org admin to allow public members on the project. the org blocks allUsers by default and project owners can't change that. prod needs this too, the lb can't authenticate to cloud run, so the services are public but only accept traffic from the lb.
  3. make a versioned gs://<project>-tfstate bucket for the control state.
  4. add the env to modules/envs, copy control/dev to control/<env> with its own bucket and env name, and apply it. that creates the deployment and infra manager builds everything else. if the first revision fails on permissions, wait a few minutes, new grants take a while to land.

gotchas #

  • infra manager runs terraform 1.5.7, so don't put anything newer in main.tf and its modules.
  • the lock file needs linux hashes too, since infra manager runs on linux: terraform providers lock -platform=linux_amd64 -platform=darwin_arm64.
  • only one apply can run per deployment at a time. a second one fails with "unable to queue the operation".
  • svfe builds aren't reproducible, so every prod deploy rolls svfe even if nothing changed (should probably attempt to fix this, nbd though)