services in here:
- bobbin
- svfe
dev is configured as a pair of bobbin and svfe without anything in front. they get automatically deployed from next branch, see more on that below.
prod is configured as multiple bobbins and svfe's. bobbins have an LB in
front and so do the svfe's. its configured such that tangled.org will point
to nearest svfe, and api.tangled.org will point to nearest bobbin. this
way all the svfe configs are the same because they just take api.tangled.org.
prod is deployed via the deploy-gcp app on the nix flake, so just run
nix run .#deploy-gcp -- --target prod --ref <branch, tag or commit> (or
--path <checkout> to deploy a local tree as it is), optionally passing
--build-on user@host to use a specific host for building the images on.
to be able to work with terraform here, outside of gcloud auth login run
gcloud auth application-default login as well, because tf wants ADC.
layout #
main.tfis what infra manager applies. it picks an env out ofmodules/envsand builds it withmodules/environment.modules/envshas one file per env with all the values. everything else reads from here, so adding an env or changing a region is one file.modules/controlandcontrol/<env>are the part we apply ourselves, see below for why.- the
deploy-gcpapp is../deploy-gcp.nu.
why there are two layers #
we want ci to deploy dev without handing it the keys to the whole project. so ci never runs terraform itself. it pushes images and asks infra manager to roll a new revision, and infra manager does the actual work with its own service account.
that service account can manage the services, network, registry and lb, but
it can't touch iam. if it could, anything that lands in the tree (and so
anything that can push to the branch) could grant itself more. so all the
identities and grants live in control/<env>, which only a project owner
applies, by hand:
nix run .#deploy-gcp -- --target dev --preview # shows what applying would change on live deployment
nix run .#deploy-gcp -- --target dev --no-build # applies it
don't apply main.tf manually. infra manager only adopts resources that
already exist while its state is clean, so anything you create behind its
back will become an annoyance for you later :p
deploying #
dev deploys on every push to sv-fe that touches bobbin or the web app
(.tangled/workflows/deploy-gcp-dev.yml).
prod runs pinned digests, not the :prod tag. a tag can move under you,
a digest is exactly what we built, and it makes a rollback just "pin the old
digests again". this is also why the control roots go through deploy-gcp
and not bare terraform. --no-build passes along the pins that are already
there, a bare apply would drop them and the env would go back to its tag.
to roll back, run nix run .#deploy-gcp -- --target prod --rollback (or
dev) and pick a build.
rollback can only reach images the registry still has. prod keeps the last 16 builds, and anything older gets deleted after 90 days. dev keeps the last 8, and deletes the rest after 3 days.
how ci gets into gcp #
the spindle signs a short-lived token for each workflow run that asks
for one (with audience: in the workflow), and gcp trusts tokens from
https://tokens.spindle.tangled.sh through workload identity federation.
nothing long-lived to leak, rotate, or paste into a secret.
gcp only lets a token act as the ci deployer if it's for one exact repo and branch. note that checking just the repo isn't enough, a manually triggered run has no branch, so it would be able to deploy otherwise.
this is only turned on for dev (ci_oidc in modules/envs).
setting up a new env #
- enable these apis: config, run, compute, artifactregistry, iam, iamcredentials, storage, cloudresourcemanager. infra manager needs the last one to write iam, and the error you get without it doesn't say so.
- get an org admin to allow public members on the project. the org blocks
allUsersby default and project owners can't change that. prod needs this too, the lb can't authenticate to cloud run, so the services are public but only accept traffic from the lb. - make a versioned
gs://<project>-tfstatebucket for the control state. - add the env to
modules/envs, copycontrol/devtocontrol/<env>with its own bucket and env name, and apply it. that creates the deployment and infra manager builds everything else. if the first revision fails on permissions, wait a few minutes, new grants take a while to land.
gotchas #
- infra manager runs terraform 1.5.7, so don't put anything newer in
main.tfand its modules. - the lock file needs linux hashes too, since infra manager runs on linux:
terraform providers lock -platform=linux_amd64 -platform=darwin_arm64. - only one apply can run per deployment at a time. a second one fails with "unable to queue the operation".
- svfe builds aren't reproducible, so every prod deploy rolls svfe even if nothing changed (should probably attempt to fix this, nbd though)