diff --git a/.gitignore b/.gitignore index ae8c727..507064e 100644 --- a/.gitignore +++ b/.gitignore @@ -9,10 +9,10 @@ modules/*/.terraform.lock.hcl # Environment state lives in S3. Bootstrap state is committed on purpose: it # holds no secrets and a fresh clone needs it to plan bootstrap - see -# bootstrap/README.md. +# envs/bootstrap/README.md. *.tfstate *.tfstate.* -!bootstrap/terraform.tfstate +!envs/bootstrap/terraform.tfstate crash.log crash.*.log diff --git a/DEVELOPING.md b/DEVELOPING.md index 39deeb9..ae9c84e 100644 --- a/DEVELOPING.md +++ b/DEVELOPING.md @@ -1,7 +1,8 @@ # Developing -Working on the running infrastructure. Assumes `AWS_PROFILE=lance-blue-apply` -and a checkout of this repo; run `tofu` commands from `envs/lance.blue`. +Working on the running infrastructure. Assumes `AWS_PROFILE=lance-blue` and a +checkout of this repo; run `tofu` commands from `envs/lance.blue`. The scripts +under `scripts/` set their own `AWS_PROFILE` and ignore yours. ## A shell on the api host @@ -19,7 +20,10 @@ The session user is not root; `sudo` works. ## Reading logs without a shell -`send-command` needs no plugin and runs as root: +`scripts/logs.sh` is the working way — it uses `start-session` with +`AWS-StartInteractiveCommand`. The `send-command` route below is broken on +current interpreters (Python 3.14 argparse; see `scripts/health.sh`), kept +here in case the CLI fixes it: CMD_ID=$(aws ssm send-command \ --instance-ids "$(tofu output -raw api_instance_id)" \ diff --git a/README.md b/README.md index e8b5236..d8c69a0 100644 --- a/README.md +++ b/README.md @@ -18,7 +18,7 @@ written here is applied. ## How it is run **By hand, from a laptop.** There is no CI, no OIDC role and no credential in -the repo. Applies run as `lance-blue-apply`, a role created by `bootstrap/` +the repo. Applies run as `lance-blue-apply`, a role created by `envs/bootstrap/` and assumed through the `lance-blue` profile. Anything sensitive is passed as `-var` on the command line. @@ -67,11 +67,13 @@ validation misses — a module edit that breaks the env roots composing it. | `modules/match-cluster/` | ECS cluster and the arena task definition | | `modules/static-site/` | S3 + CloudFront for the headquarters frontend | | `modules/dns/` | the `lance.blue` hosted zone | -| `bootstrap/` | one-shot: the state bucket and the apply role | -| `docs/` | architecture, decisions, cost model, runbook, licensing | +| `modules/api-host/` | the EC2 host running headquarters-api behind Caddy | +| `envs/bootstrap/` | one-shot: the state bucket and the apply role | +| `scripts/` | deploy, logs, a shell on the api host, match operations | +| `docs/` | architecture, runbook | -Modules take inputs and return outputs; they never read remote state or call -`aws` themselves. Each directory under `envs/` is one AWS account. Dev and +Modules take inputs and return outputs; they never read remote state, and the +one place a module calls `aws` itself is the static-site cache invalidation. Each directory under `envs/` is one AWS account. Dev and prod resources share the account and the root, told apart by name suffix; some resources (the hosted zone, ECR) are shared between environments. A second AWS account would be a directory copy. @@ -85,8 +87,8 @@ second AWS account would be a directory copy. - **Tags** come from `default_tags` on the provider: `Project=lance.blue`, `ManagedBy=terraform`, `Repo=infra`, and `Commit` when the sha is passed at plan time (`-var git_commit=$(git rev-parse --short=12 HEAD)`; omitted, the - tag is simply absent). Modules add `Component`, and per-environment - resources add `Environment`. + tag is simply absent). The env root passes `Component` — and, for + per-environment resources, `Environment` — into each module's `tags`. - **Regions**: one, from `var.region`. Anything that must live in `us-east-1` regardless (CloudFront-facing certificates) uses an aliased provider and says so. diff --git a/TODO.md b/TODO.md index 4799bd2..9f3f15b 100644 --- a/TODO.md +++ b/TODO.md @@ -11,9 +11,9 @@ compute of its own yet. - [ ] **Compute for the WebSocket proxy, when it splits out.** The API runs on the api host and proxies matches from inside its own binary today. A separate proxy binary needs somewhere to live, and its - drain-on-deploy story (headquarters M5) decides where. Options and - prices are in - [docs/decisions.md](docs/decisions.md#where-the-control-plane-runs). + drain-on-deploy story (headquarters M5) decides where. The priced + comparison of options was in docs/decisions.md, which was deleted; + it needs redoing when the decision is live. - [ ] **No external backstop for a hung task.** A hung match bills until something kills it. Arena enforces `limits.idleTimeoutSeconds` from inside the container now (`container/watch/idle.sh`), so the common diff --git a/docs/architecture.md b/docs/architecture.md index c53f08b..dd2e4df 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -3,7 +3,7 @@ The target picture. Ported from the spike's notes and the shape arena was built against; refined where building arena settled a question. -Nothing here is deployed. Where a component has no Terraform yet, it says so. +All of this is applied and running, except where a line says otherwise. ## The shape @@ -12,7 +12,7 @@ Nothing here is deployed. Where a component has no Terraform yet, it says so. │ ┌────────────────────┴────────────────────┐ │ │ - www / static hq.lance.blue + www / static api.lance.blue │ │ ▼ ▼ CloudFront headquarters @@ -44,7 +44,7 @@ Nothing here is deployed. Where a component has no Terraform yet, it says so. | network | VPC, public subnets, IGW | `modules/network` | written | | frontend hosting | S3 + CloudFront | `modules/static-site` | written | | DNS | Route 53 | `modules/dns` | written | -| control plane + WebSocket proxy | undecided | — | **not written** | +| control plane + WebSocket proxy | EC2 (Graviton) + Caddy | `modules/api-host` | written | ## One match, one task @@ -64,8 +64,9 @@ Sizing evidence, all from the spike: (384m–2g) *and* match size (2v2 vs 4v4). The cost is per-JVM overhead, not game state, so smaller matches are not cheaper. That measurement is of the headless server plus bots; the Suramadu server and the client JVM are on - top of it and have never been measured together. The task definition asks for - 1 vCPU / 2GB and that number is a guess until a real match is watched. + top of it. The task definition asks for 2 vCPU / 8GB now, from watching + real matches: 1 vCPU starved the client JVM, and a full match peaked at + 3.2GB (see `envs/lance.blue/variables.tf`). - **CPU is not the constraint** — 8 concurrent games used 2.2 of 16 cores. - **`units.cache` is not shipped by MegaMek** — 20.2s to build from 10,989 files, 2.9s to load a prebuilt one. arena's image bakes it, which is worth @@ -87,8 +88,8 @@ were signed elsewhere. The task *execution* role — the one ECS itself uses to pull the image and open a log stream — is the only role with policy attached. That is also why no ATProto token ever reaches a match container, which is a -security property and a licensing one at the same time; see -[licensing.md](licensing.md). +security property and a licensing one at the same time; the licensing side is +covered in arena's `LICENSING.md`. ## Networking @@ -106,9 +107,6 @@ That matters because **MegaMek has no per-player authentication.** Whoever reaches the port is whoever the game thinks they are. The proxy is the auth boundary; the container is never it. -The reasoning and the alternatives considered are in -[decisions.md](decisions.md#no-nat-gateway). - ## We do not host identity Worth stating, because the spike's notes assumed the opposite and the @@ -129,14 +127,15 @@ attributing an uploaded turn report to somebody requires naming them. A DID resolves through whatever service actually holds that identity, which is not one of ours. -## The gap: control plane and proxy +## The gap: where the proxy lives -There is no Terraform for headquarters, because there is no headquarters. The -decision that blocks it is *what it runs on*, and it is a cost decision more -than a technical one — an ALB alone is ~$16/month, which is more than the rest -of the standing cost put together. Options are laid out in -[decisions.md](decisions.md#where-the-control-plane-runs). +The control plane runs on `modules/api-host`: one Graviton instance, Caddy +terminating TLS at api.lance.blue, deploys by image tag. What is still open is +compute for the WebSocket proxy once headquarters splits it out of the api +binary (TODO.md tracks it, waiting on that repo's drain-on-deploy story). The +cost shape that ruled out an ALB — ~$16/month, more than the rest of the +standing cost put together — still applies to whatever the proxy lands on. -Whatever it becomes, it is the only always-on component, it holds every -credential in the system, and it must remain a separate process from anything -running Suramadu code. +The api host is the only always-on component, it holds the OAuth session state +for players' own PDSes — every credential the system keeps — and it must remain +a separate process from anything running Suramadu code. diff --git a/docs/runbook.md b/docs/runbook.md index 18d06ab..11bee00 100644 --- a/docs/runbook.md +++ b/docs/runbook.md @@ -47,7 +47,7 @@ The account is `179302349187`. **2. Bootstrap.** Create the state bucket and the apply role. - cd bootstrap + cd envs/bootstrap tofu init tofu apply -var 'region=us-east-1' @@ -55,7 +55,7 @@ role. `~/.aws/config`, set its `source_profile` to your SSO profile for this account, then check it: - export AWS_PROFILE=lance-blue-apply + export AWS_PROFILE=lance-blue aws sts get-caller-identity # expect assumed-role/lance-blue-apply The administrator credentials are only needed again to change the apply role @@ -110,8 +110,8 @@ lance.blue's Bluesky account when a release id moved — nothing to announce on an infrastructure-only change, and nothing at all when the plan is empty. `--dry-run` stops after the plan; `--no-post` applies silently. -The post is built from three outputs and no others: `releases`, `site_url` and -`api_url`. All three are public knowledge. Bucket names, the instance id and +The post is built from two outputs and no others: `releases` and `site_url`. +Both are public knowledge. Bucket names, the instance id and the security group ids are outputs too and deliberately stay out of it, so adding an output can never leak one by accident — posting more means editing the script. @@ -165,8 +165,8 @@ the basename of its own `IMAGE_NAME`: | docker login --username AWS --password-stdin \ "$(tofu -chdir=$INFRA output -raw ecr_registry_host)" - ./build.sh - REGISTRY=$(tofu -chdir=$INFRA output -raw ecr_push_registry) ./push.sh + ./scripts/build.sh + REGISTRY=$(tofu -chdir=$INFRA output -raw ecr_push_registry) ./scripts/push.sh The same `REGISTRY` works for every repository in the namespace. To see what exists: `tofu -chdir=envs/lance.blue output -json ecr_repositories`. diff --git a/envs/bootstrap/README.md b/envs/bootstrap/README.md index c67fd53..6cdb227 100644 --- a/envs/bootstrap/README.md +++ b/envs/bootstrap/README.md @@ -64,6 +64,6 @@ apart. One IAM role and its policy, per iam.tf. The role has admin access to the services this repo uses, can only write IAM under the `lance-blue-` prefix, -and is explicitly denied from changing its own permissions. The reasoning is -in -[docs/decisions.md](../docs/decisions.md#borrowed-admin-once-then-an-apply-role). +and is explicitly denied from changing its own permissions. Admin credentials +are borrowed exactly once, to create this role; every apply after that runs as +the role, so day-to-day work never holds more than the project needs. diff --git a/envs/bootstrap/iam.tf b/envs/bootstrap/iam.tf index 6f661fd..75303c8 100644 --- a/envs/bootstrap/iam.tf +++ b/envs/bootstrap/iam.tf @@ -59,14 +59,14 @@ data "aws_iam_policy_document" "apply" { # Per-service wildcards instead of a curated action list, because a curated # list breaks every time a module grows a new resource type. Account-wide # s3:* and friends are acceptable because the account contains only this - # project. See docs/decisions.md. + # project. statement { sid = "ServiceAdmin" actions = [ "ec2:*", # network: VPC, subnets, routing, security groups "ecs:*", # match cluster, plus run-task by hand per the runbook "ecr:*", # registry, plus docker login and push from a laptop - "s3:*", # artifacts, the state bucket, the future site bucket + "s3:*", # artifacts, the state bucket, the site and log buckets "route53:*", # the hosted zone "logs:*", # the match log group "cloudfront:*", # static site @@ -85,7 +85,7 @@ data "aws_iam_policy_document" "apply" { } # IAM writes are limited to roles and policies with the project prefix - - # the match task roles today, the control plane's later. This includes iam:PassRole; + # the match task roles and the api host's role. This includes iam:PassRole; # what a passed role can do is still limited by that role's own trust # policy. statement { diff --git a/envs/bootstrap/main.tf b/envs/bootstrap/main.tf index 53bfc5a..a6e2cb4 100644 --- a/envs/bootstrap/main.tf +++ b/envs/bootstrap/main.tf @@ -1,8 +1,9 @@ # The S3 bucket that holds every environment's state. # # This root has no backend of its own: a bucket cannot store the state of its -# own creation. Its state file stays on local disk, gitignored, and losing it -# costs a `tofu import` rather than the bucket - see README.md. +# own creation. Its state file stays in this directory and is committed on +# purpose - it holds no secrets, and a fresh clone needs it to plan bootstrap; +# see README.md. # # Run once per AWS account, ever. diff --git a/envs/lance.blue/main.tf b/envs/lance.blue/main.tf index 6e29cc4..b5c5362 100644 --- a/envs/lance.blue/main.tf +++ b/envs/lance.blue/main.tf @@ -3,8 +3,8 @@ # Dev and prod resources share the account and this root, told apart by an # environment suffix on the name: lance-blue--prod. Shared # resources - the hosted zone, the ECR repositories - carry no environment. -# Only prod's resources exist today. See -# docs/decisions.md#one-account-environments-by-suffix. +# Only prod's resources exist today. One account rather than one per +# environment because the account contains only this project. # # This file composes modules and nothing else: no resources are declared here, # so anything that turns out to be worth reusing already lives somewhere it can diff --git a/envs/lance.blue/outputs.tf b/envs/lance.blue/outputs.tf index 16b553d..c428b65 100644 --- a/envs/lance.blue/outputs.tf +++ b/envs/lance.blue/outputs.tf @@ -16,18 +16,18 @@ output "public_subnet_ids" { value = module.network.public_subnet_ids } -# What arena's docker/push.sh wants: +# What arena's scripts/push.sh wants: # # aws ecr get-login-password --region us-east-1 \ # | docker login --username AWS --password-stdin "$(tofu output -raw ecr_registry_host)" -# REGISTRY=$(tofu output -raw ecr_push_registry) ./docker/push.sh +# REGISTRY=$(tofu output -raw ecr_push_registry) ./scripts/push.sh output "ecr_registry_host" { description = "Registry host, for docker login." value = module.ecr.registry_host } output "ecr_push_registry" { - description = "The REGISTRY value arena's docker/push.sh expects." + description = "The REGISTRY value arena's scripts/push.sh expects." value = module.ecr.push_registry } @@ -109,11 +109,11 @@ output "api_ssm_parameters" { value = module.api_host.ssm_parameter_names } -# Both are unattached. There is no control plane role to attach them to; when -# there is, it needs both. The api host's role does not get them - it starts -# no tasks. +# Both are attached to the api host's role, which is the control plane today - +# it is what launches match tasks. Echoed here so what the role carries is +# readable from state. output "control_plane_policy_arns" { - description = "Unattached IAM policies the control plane's role should carry once it exists: artifacts access, and permission to start a match." + description = "IAM policies the control plane's role carries: artifacts access, and permission to start a match. Attached to the api host's role." value = { artifacts = module.artifacts.control_plane_policy_arn run_match = module.match_cluster.control_plane_policy_arn diff --git a/envs/lance.blue/variables.tf b/envs/lance.blue/variables.tf index f4c7930..7876313 100644 --- a/envs/lance.blue/variables.tf +++ b/envs/lance.blue/variables.tf @@ -18,7 +18,7 @@ variable "account_id" { } variable "region" { - description = "AWS region. us-east-1 is cheapest; it is a poor default for European players. See docs/decisions.md#region-us-east-1." + description = "AWS region. us-east-1 is cheapest; it is a poor default for European players." type = string default = "us-east-1" } @@ -133,18 +133,18 @@ variable "match_task_cpu" { } variable "match_task_memory" { - description = "Fargate memory in MiB for a match. No default: see match_task_cpu. Measured: ~1.1GB steady, 2.0GB peak. Fargate also sets a floor per CPU size - 2 vCPU needs at least 4096." + description = "Fargate memory in MiB for a match. No default: see match_task_cpu. Measured: a full match peaked at 3.2GB at 2 vCPU / 4GB (see TODO.md); 8192 clears arena's ~4.5GB of heap ceilings. Fargate also sets a floor per CPU size - 2 vCPU needs at least 4096." type = number } variable "match_max_concurrent_tasks" { - description = "Most match tasks allowed at once. A cost ceiling the control plane enforces; 20 tasks at 1 vCPU / 2GB is ~$0.99/hour if every slot is full. Must stay below the account's Fargate On-Demand vCPU quota (30 as of 2026-08) divided by match_task_cpu, or AWS refuses RunTask before the cap matters." + description = "Most match tasks allowed at once. A cost ceiling the control plane enforces; 20 tasks at 2 vCPU / 8GB is ~$2.33/hour if every slot is full. Must stay below the account's Fargate On-Demand vCPU quota (30 as of 2026-08) divided by the task's vCPUs - 2 today, so 15, which this default exceeds: AWS refuses RunTask above 15 before the cap matters." type = number default = 20 } variable "match_task_architecture" { - description = "Flip to ARM64 for ~20% off, once arena can build an arm64 image. Tracked in TODO.md." + description = "Flip to ARM64 for ~20% off, once arena can build an arm64 image (arena's TODO tracks the x86_64-only pins)." type = string default = "X86_64" } diff --git a/modules/api-host/main.tf b/modules/api-host/main.tf index 8d80c75..6801a84 100644 --- a/modules/api-host/main.tf +++ b/modules/api-host/main.tf @@ -5,14 +5,12 @@ # swapping the container when it changes - seconds of downtime, not an # instance replacement. Infra changes (user_data, instance type) still # replace the box. What must survive a replacement lives on a separate EBS -# volume: the sqlite stopgap store (a credential store; DynamoDB is the -# tracked follow-up) and Caddy's certificates, so a replacement neither logs -# players out nor re-asks Let's Encrypt. +# volume: the sqlite stopgap store (OAuth session state against players' own +# PDSes; DynamoDB is the tracked follow-up) and Caddy's certificates, so a +# replacement neither logs players out nor re-asks Let's Encrypt. # # TLS terminates in a Caddy container; the Rust binary never sees a # certificate. No SSH: shell access is SSM Session Manager. -# -# See docs/decisions.md#where-the-control-plane-runs. terraform { required_version = ">= 1.11" diff --git a/modules/artifacts/main.tf b/modules/artifacts/main.tf index 1de61e4..fc27508 100644 --- a/modules/artifacts/main.tf +++ b/modules/artifacts/main.tf @@ -152,14 +152,13 @@ data "aws_iam_policy_document" "bucket" { # ceiling on what any manifest can hand a match container. Signing itself is # local and needs no API call - the policy is what makes the resulting URL work. # -# Created unattached: there is no control plane role to attach it to yet. It -# exists now because writing down the contract is most of the value, and an -# unattached IAM policy is free. +# The env root attaches it to the api host's role, which is the control plane +# today. Written as its own policy because the contract is most of the value. resource "aws_iam_policy" "control_plane" { count = var.create_control_plane_policy ? 1 : 0 name = "${var.name_prefix}-artifacts-access-policy-${var.environment}" - description = "Read and write match artifacts, and sign URLs granting the same. Intended for the control plane's role; nothing attaches it yet." + description = "Read and write match artifacts, and sign URLs granting the same. Carried by the control plane's role - the api host today." policy = data.aws_iam_policy_document.control_plane.json tags = var.tags diff --git a/modules/artifacts/outputs.tf b/modules/artifacts/outputs.tf index 55ae9b1..a629264 100644 --- a/modules/artifacts/outputs.tf +++ b/modules/artifacts/outputs.tf @@ -14,6 +14,6 @@ output "bucket_regional_domain_name" { } output "control_plane_policy_arn" { - description = "ARN of the unattached policy the control plane's role should carry. Null when create_control_plane_policy is false." + description = "ARN of the policy the control plane's role carries; the env root attaches it. Null when create_control_plane_policy is false." value = one(aws_iam_policy.control_plane[*].arn) } diff --git a/modules/artifacts/variables.tf b/modules/artifacts/variables.tf index 2a65b34..0b36f6d 100644 --- a/modules/artifacts/variables.tf +++ b/modules/artifacts/variables.tf @@ -26,7 +26,7 @@ variable "camo_cache_retention_days" { } variable "create_control_plane_policy" { - description = "Create the unattached IAM policy describing what the control plane may do here. Free, and writing down the contract is most of the value." + description = "Create the IAM policy describing what the control plane may do here; the env root attaches it to the api host's role." type = bool default = true } diff --git a/modules/dns/main.tf b/modules/dns/main.tf index 3645e81..86caf5d 100644 --- a/modules/dns/main.tf +++ b/modules/dns/main.tf @@ -1,14 +1,15 @@ # The hosted zone for lance.blue. # -# A zone for our own names - the site, and the control plane when it exists. +# A zone for our own names - the site, and the control plane at api.. # Not identity infrastructure: we do not run a PDS, do not create accounts and # do not issue handles under this domain, so nothing writes a record here per # player. See docs/architecture.md#we-do-not-host-identity. # # No certificate either. The site's certificate lives in the static-site # module, and the only other thing that wanted one from us was per-player TLS, -# which is not happening. When the control plane exists it will need a -# certificate for its own name; that belongs with the control plane, not here. +# which is not happening. The control plane's certificate comes from Let's +# Encrypt via Caddy on the api host, not from ACM, so it needs nothing here +# beyond its records (which the api-host module writes into this zone). terraform { required_version = ">= 1.11" diff --git a/modules/match-cluster/control-plane.tf b/modules/match-cluster/control-plane.tf index 4c8de8d..bf5f9d3 100644 --- a/modules/match-cluster/control-plane.tf +++ b/modules/match-cluster/control-plane.tf @@ -1,14 +1,14 @@ # What the control plane needs in order to start a match. # -# Created unattached, like the artifacts policy: there is no headquarters role -# to attach it to yet, and an unattached IAM policy is free. Writing the -# contract down now is what stops it being written as "administrator" later. +# Attached by the env root to the api host's role, like the artifacts policy - +# the api host is the control plane today. Written as its own policy because +# the contract is what stops it being written as "administrator" later. resource "aws_iam_policy" "control_plane" { count = var.create_control_plane_policy ? 1 : 0 name = "${var.name_prefix}-run-match-policy-${var.environment}" - description = "Start, stop and observe match tasks. Intended for the control plane's role; nothing attaches it yet." + description = "Start, stop and observe match tasks. Carried by the control plane's role - the api host today." policy = data.aws_iam_policy_document.control_plane.json tags = var.tags diff --git a/modules/match-cluster/main.tf b/modules/match-cluster/main.tf index b8acfa2..ada856b 100644 --- a/modules/match-cluster/main.tf +++ b/modules/match-cluster/main.tf @@ -41,7 +41,7 @@ resource "aws_ecs_cluster" "this" { # reclaimed match is a match people were playing, and arena has no checkpoint or # resume - the spike found reconnection is exactly where things break. So Spot # is available for benchmark and load-test runs, and the default strategy is -# on-demand. See docs/decisions.md#fargate-on-demand-not-spot. +# on-demand. resource "aws_ecs_cluster_capacity_providers" "this" { cluster_name = aws_ecs_cluster.this.name capacity_providers = ["FARGATE", "FARGATE_SPOT"] diff --git a/modules/match-cluster/outputs.tf b/modules/match-cluster/outputs.tf index f4f1070..5e2bea2 100644 --- a/modules/match-cluster/outputs.tf +++ b/modules/match-cluster/outputs.tf @@ -44,6 +44,6 @@ output "max_concurrent_tasks_parameter" { } output "control_plane_policy_arn" { - description = "ARN of the unattached policy the control plane's role needs to start a match. Null when create_control_plane_policy is false." + description = "ARN of the policy the control plane's role needs to start a match; the env root attaches it. Null when create_control_plane_policy is false." value = one(aws_iam_policy.control_plane[*].arn) } diff --git a/modules/match-cluster/variables.tf b/modules/match-cluster/variables.tf index 885ac7c..649b85e 100644 --- a/modules/match-cluster/variables.tf +++ b/modules/match-cluster/variables.tf @@ -69,7 +69,7 @@ variable "megamek_port" { } variable "proxy_security_group_ids" { - description = "Security groups permitted to reach the container port. Empty until the WebSocket proxy exists, which means no ingress at all." + description = "Security groups permitted to reach the container port. The api host's group today; empty means no ingress at all." type = list(string) default = [] } @@ -81,7 +81,7 @@ variable "log_retention_days" { } variable "max_concurrent_tasks" { - description = "Most match tasks allowed to run at once. A cost ceiling, not a scaling target - ECS has no native cap on standalone RunTask tasks, so this is published for the control plane to enforce before each RunTask. Size it in relation to the account's Fargate On-Demand vCPU quota (30 as of 2026-08): above quota / task_cpu, AWS starts refusing RunTask before this cap is ever reached." + description = "Most match tasks allowed to run at once. A cost ceiling, not a scaling target - ECS has no native cap on standalone RunTask tasks, so this is published for the control plane to enforce before each RunTask. Size it in relation to the account's Fargate On-Demand vCPU quota (30 as of 2026-08): above quota divided by the task's vCPUs (2 today, so 15), AWS starts refusing RunTask before this cap is ever reached." type = number default = 20 @@ -92,7 +92,7 @@ variable "max_concurrent_tasks" { } variable "create_control_plane_policy" { - description = "Create the unattached IAM policy describing what the control plane may do here." + description = "Create the IAM policy describing what the control plane may do here; the env root attaches it to the api host's role." type = bool default = true } diff --git a/modules/network/main.tf b/modules/network/main.tf index 89e2731..a3627ed 100644 --- a/modules/network/main.tf +++ b/modules/network/main.tf @@ -3,7 +3,7 @@ # Public subnets only, no NAT gateway, no interface endpoints. A NAT gateway is # ~$33/month before a byte moves and the three interface endpoints that would # replace it are ~$21/month; a public IP on a task is $0.005/hour and only for -# the length of the match. See docs/decisions.md#no-nat-gateway. +# the length of the match. # # Public IP does not mean reachable. Nothing in this module opens ingress - # security groups are owned by the modules that need them, and the match task's diff --git a/modules/static-site/main.tf b/modules/static-site/main.tf index 81a874c..4724efe 100644 --- a/modules/static-site/main.tf +++ b/modules/static-site/main.tf @@ -5,8 +5,6 @@ # bucket; this module provides the bucket and the distribution, and # var.origin_path selects which prefix is served. Changing it invalidates the # cache as part of the apply. -# -# See docs/decisions.md#s3--cloudfront-for-the-static-site. terraform { required_version = ">= 1.11"