This repository has no description
README.md

services in here:

  • bobbin
  • svfe
  • hydrant
  • seo-reporter

dev is configured as a pair of bobbin and svfe without anything in front. they get automatically deployed from next branch, see more on that below.

prod is configured as multiple bobbins and svfe's. bobbins have an LB in front and so do the svfe's. its configured such that tangled.org will point to nearest svfe, and api.tangled.org will point to nearest bobbin. this way all the svfe configs are the same because they just take api.tangled.org. prod is deployed via the deploy app on the nix flake, so just run nix run .#deploy -- --target prod --ref <branch, tag or commit> (or --path <checkout> to deploy a local tree as it is), optionally passing --build-on user@host to use a specific host for building the images on.

to be able to work with terraform here, outside of gcloud auth login run gcloud auth application-default login as well, because tf wants ADC.

layout #

  • main.tf is what infra manager applies. it picks an env out of modules/envs and builds it with modules/environment.
  • modules/envs has one file per env with all the 'actual' config. its basically the place to edit to add new hydrants or bobbins etc.
  • modules/control and control/<env> are the part we apply ourselves, see below for why.
  • the deploy app is ../deploy.nu.
  • modules/nodes and nodes/<env> are the nixos vms, see the end.
  • modules/workers and workers/<env> are the cloudflare workers, see the very end.

why there are two layers #

we want ci to deploy dev without handing it the keys to the whole project. so ci never runs terraform itself. it pushes images and asks infra manager to roll a new revision, and infra manager does the actual work with its own service account.

that service account can manage the services, network, registry and lb, but it can't touch iam. if it could, anything that lands in the tree (and so anything that can push to the branch) could grant itself more. so all the identities and grants live in control/<env>, which only a project owner applies, by hand:

nix run .#deploy -- --target dev --preview   # shows what applying would change on live deployment
nix run .#deploy -- --target dev             # applies it

don't apply main.tf manually. infra manager only adopts resources that already exist while its state is clean, so anything you create behind its back will become an annoyance for you later :p

deploying #

dev deploys on every push to sv-fe that touches bobbin or the web app (.tangled/workflows/deploy-gcp-dev.yml).

prod runs pinned digests, not the :prod tag. a tag can move under you, a digest is exactly what we built, and it makes a rollback just "pin the old digests again". this is also why the control roots go through deploy and not bare terraform. deploy passes along the pins that are already there, a bare apply would drop them and the env would go back to its tag.

a build only builds images whose source changed. each image gets a src-<hash> tag, hashed from its source (for bobbin, only the crates it depends on, as the core flake narrows them) and its build args, and an image that already has the tag gets reused. ci's images don't get one, so the first local deploy after ci still builds. --only bobbin deploys just that service, and --rebuild builds anyway, e.g. to pick up a newer base image.

to roll back, run nix run .#deploy -- --target prod --rollback (or dev) and pick a build.

rollback can only reach images the registry still has. prod keeps the last 16 builds, and anything older gets deleted after 90 days. dev keeps the last 8, and deletes the rest after 3 days.

how ci gets into gcp #

the spindle signs a short-lived token for each workflow run that asks for one (with audience: in the workflow), and gcp trusts tokens from https://tokens.spindle.tangled.sh through workload identity federation. nothing long-lived to leak, rotate, or paste into a secret.

gcp only lets a token act as the ci deployer if it's for one exact repo and branch. note that checking just the repo isn't enough, a manually triggered run has no branch, so it would be able to deploy otherwise.

this is only turned on for dev (ci_oidc in modules/envs).

setting up a new env #

  1. enable these apis: config, run, compute, artifactregistry, iam, iamcredentials, storage, cloudresourcemanager. infra manager needs the last one to write iam, and the error you get without it doesn't say so.
  2. get an org admin to allow public members on the project. the org blocks allUsers by default and project owners can't change that. prod needs this too, the lb can't authenticate to cloud run, so the services are public but only accept traffic from the lb.
  3. make a versioned gs://<project>-tfstate bucket for the control state.
  4. add the env to modules/envs, copy control/dev to control/<env> with its own bucket and env name, and apply it. that creates the deployment and infra manager builds everything else. if the first revision fails on permissions, wait a few minutes, new grants take a while to land.

gotchas #

  • infra manager runs terraform 1.5.7, so don't put anything newer in main.tf and its modules.
  • the lock file needs linux hashes too, since infra manager runs on linux: terraform providers lock -platform=linux_amd64 -platform=darwin_arm64.
  • only one apply can run per deployment at a time. a second one fails with "unable to queue the operation".

nodes #

nodes/<env> runs hydrant on nixos vms. in prod they sit behind one global lb, which sends each client to the closest healthy node. dev leaves the lb out (lb = false), so its node is only reachable from inside the vpc and over ssh. it's a plain terraform root that infra manager never sees (the blueprint excludes it). deploy applies it after the services, except for --rollback, and asks first when the plan would destroy anything beyond redeploying a changed system. --no-nodes skips it, and ci never touches it:

nix run .#deploy -- --target dev              # services and nodes
terraform -chdir=terraform/nodes/dev apply    # just the nodes, which have no pins to lose

a new vm boots debian, and nixos-anywhere replaces it with nixosConfigurations.hyd-gcp, getting in over a throwaway key that dies with debian. after that terraform deploys again whenever the system changes, over your own ssh agent as tangler. every node is that one config with its name passed in, and none of them are in the colmena hive.

  • every plan builds the system, so you need nix and an x86_64-linux builder.
  • the flake is read from git, so new files need at least git add -N first.
  • a replaced vm is reinstalled from scratch and starts its index over, so the debian image is ignored after creation.
  • hubs in modules/envs groups the cells around the node their bobbins read, so a new region goes into the hub whose hydrant it should use. the node itself lives in one of its hub's cells, and the vm is named after its name.
  • bobbin can't get ready without its hydrant, so the blueprint leaves it out of a cell until the hub's node exists (its <node>-internal address, really). a new hub's first deploy makes the subnet and svfe, the nodes make the vm, and deploy then rolls the blueprint once more to put bobbin in. with --no-nodes, the next deploy or ci push does.
  • --preview can't plan the nodes of a hub whose subnet doesn't exist yet, pass --no-nodes until its first deploy.
  • the nodes themselves are listed in modules/envs. to grow a node's disk, raise its disk_gb there and apply. the disk grows in place, and the node hands the space to hydrant on its next boot, or right away with systemctl start grow-hydrant-volume.

workers #

workers/<env> deploys the cloudflare workers listed under workers in modules/envs, each keyed by its directory in core. like the nodes it's a plain root that infra manager never sees. terraform can't bundle typescript, so deploy builds every worker in the tree it was given (pnpm install and pnpm run build, which leaves dist/index.js) and applies the root with those bundles. that only happens with --ref or --path, --only workers deploys just them, and ci never touches them. a worker's wrangler config is for wrangler dev, what runs is whatever modules/envs says.

control makes the google side: the secrets the workers read, and for a worker with a google_identity, a service account <worker>-<suffix> and a key for it, which lands in <worker>-google-key and from there in the binding. its apis get enabled there too.

the first deploy with a worker in it only runs control, because the list of workers comes from control's last apply. fill in the secrets it made before the next one, without a trailing newline. cloudflare-workers-token wants workers scripts edit on the account.

printf %s "$VALUE" | gcloud secrets versions add <secret> --data-file=- --project exemplary-proxy-507408-d6
  • the key sits in control's state, so the state bucket is as secret as the key. to rotate it, terraform -chdir=terraform/control/dev taint 'module.control.google_service_account_key.worker["<worker>"]' and deploy with --ref/--path (a bare apply of control would drop the image pins).
  • if the org enforces iam.disableServiceAccountKeyCreation, the key needs an exception for the project.
  • --preview doesn't plan the workers, since that would mean building them.