diff --git a/README.md b/README.md index d56b39b..1780cc5 100644 --- a/README.md +++ b/README.md @@ -69,7 +69,7 @@ validation misses — a module edit that breaks the env roots composing it. | `modules/dns/` | the `lance.blue` hosted zone | | `modules/api-host/` | the EC2 host running headquarters-api behind Caddy | | `envs/bootstrap/` | one-shot: the state bucket and the apply role | -| `scripts/` | deploy, logs, a shell on the api host, match operations, health checks | +| `scripts/` | deploy, logs, a shell on the api host, match operations, per-match cost, health checks | | `docs/` | architecture, runbook | Modules take inputs and return outputs; they never read remote state, and the diff --git a/TODO.md b/TODO.md index 5e13fc6..155f500 100644 --- a/TODO.md +++ b/TODO.md @@ -21,11 +21,18 @@ compute of its own yet. dies with the game — nothing outside the container notices. An EventBridge rule or an alarm on task age would close that, and would double as the "a match task exists at all" signal the entry below - wants. + wants. `scripts/matches/cost.sh` now says what it is worth: over the 69 + matches to 2026-08-21, 14 ran longer than three hours and account for + 92% of the $49 spent; the median match that finished cost $0.015. The + longest billed 67 hours and cost $8.28. - [ ] **Container Insights is on at the `enhanced` tier.** Turned on to size the match task, which it did. It bills per metric per task; decide whether the per-task numbers are worth keeping now that - `scripts/matches/list.sh` has done its job. + `scripts/matches/list.sh` has done its job. Measured since: + 0.7% of match spend, about $0.005 a match, and it is the only source + of a task's peak memory and of a start time once ECS has forgotten the + task — `scripts/matches/cost.sh` reads both. Cheap, and now load-bearing + for a day after each match. - [ ] **Match task sizing is measured, not final.** 1 vCPU starved the client JVM (pinned at 100%); 2 vCPU / 4GB ran a full match at 1339 CPU units and a 3.2GB peak. The task now runs 2 vCPU / 8GB, because arena's heap diff --git a/docs/runbook.md b/docs/runbook.md index c73b402..23a4567 100644 --- a/docs/runbook.md +++ b/docs/runbook.md @@ -294,3 +294,41 @@ URL is a bearer credential and does not belong in a task definition revision. it the image pull hangs and then times out — which reads as the wrong problem. Logs land in the group `match_log_group` names. + +## What a match cost + + ./scripts/matches/cost.sh collect # ask AWS, file the facts + ./scripts/matches/cost.sh report # price them + +`collect` reads six sources — a live task from ECS, RunTask from CloudTrail, +the match log group, Container Insights, the artifacts in S3, and the match's +own `result.json` — and writes what each of them said into a ledger under +`.work/`. `report` prices that ledger against `scripts/matches/prices.json` +and never touches AWS, so a corrected assumption re-prices every match already +recorded. + +Run `collect` on a schedule. Every source forgets, and they forget at +different speeds: about an hour for a stopped task in ECS, a day for +Container Insights, seven days for the match logs, thirty for the artifacts, +ninety for CloudTrail. Run daily it keeps exact numbers; run monthly it keeps +inferred ones. Nothing already in the ledger is downgraded by a later run. + +The `src` column says where each match's billed window came from, as +start/end — `ecs` is the task's own record, `trail/s3` is an exact RunTask and +an end taken from the last artifact upload. `--explain ` prints one +match line by line, including which sources were still readable for it. + + ./scripts/matches/cost.sh collect --hours 2160 # backfill CloudTrail's 90 days + ./scripts/matches/cost.sh report --explain 5fe10a12 + ./scripts/matches/cost.sh report --format csv > costs.csv + ./scripts/matches/cost.sh prices --refresh # re-read the AWS price list + +Prices come from the AWS Price List API and are committed, so a report is +reproducible and a price change is a diff. The assumptions beside them — +what a Spot task costs, how long provisioning takes, how often artifacts are +read back — are hand-maintained and a refresh leaves them alone. + +`scripts/matches/cost.py` documents what is counted and what is not. The +short version: Fargate is around 95% of a match, and serving the report pages +afterwards is the one real per-match cost that is measurable and not measured, +because it lives in CloudFront's access logs and needs Athena to read.