about things notes.zzstoatzz.io
notes
notes operations instance-sizing.md
4.0 kB
Markdown
at main

instance sizing #

size a cloud instance to a measured working set, not to an estimate. the estimate is almost always the sum of everyone's worst case, and the difference between it and the measurement is frequently a bug rather than a requirement.

a service provisioned on a 32-core / 128 GB dedicated box measured at roughly one core and a 53.6 GiB resident set, of which about 51 GiB was a single oversized index structure. once the structure was fixed the working set was 2.4 GiB, and the service moved to an 8-vcpu / 16 GB shared instance with six times the headroom. the compute premium existed only to house the bug.

measuring the working set #

  • measure at a fresh process start. allocators retain freed pages, so RSS measured after a fix overstates the requirement until the process restarts.
  • separate resident from required. page cache, arena retention, and preallocated buffers all inflate RSS without representing demand.
  • measure the load you actually serve. a peak during a backfill is a reason to rent capacity for the backfill, not to own it year-round.

the billing asymmetries worth designing around #

rescale is up-only for disks, reversible for cpu and ram. growing a root disk forecloses ever moving to a smaller instance type; a cpu/ram-only rescale round-trips (on hetzner, upgrade_disk=false). therefore start small and rescale up on evidence. being too small costs minutes of downtime once; being too large costs every month until someone re-measures.

hourly billing makes insurance nearly free. keeping the old instance one extra night as a rollback target costs about an hour's rate per hour. upsizing for a one-day heavy job and back costs single-digit currency units. these are scratch allocations, not commitments — price them accordingly.

read the whole bill before celebrating. storage frequently is the product. a 3 TB volume dwarfed the right-sized server that replaced the old one, which turned a 97% saving on the server line into a 71% saving overall. identify the dominant line item before optimizing anything.

shared vcpu is adequate for rate-limited and I/O-bound work. a polite network crawl or a one-core serving load is not measurably affected by contention. reserve dedicated cores for sustained cpu-bound phases, and rent them only for the duration of those phases.

burstable vcpu fails as throttling, not saturation — and the process is not the one that notices. a fly shared-cpu-1x accrues credits against a 1/16-core baseline; ~0.15 cores sustained for ten minutes drains the whole balance (fly_instance_cpu_balance → 0) and the hypervisor pins the machine to baseline — steal jumps to 50%+ while guest user+system stay small. from inside, the app is idle and error-free; from outside, every request crawls (an oauth page went 1.6s, upstream proxy fetches hit 10s timeouts). diagnose from the host metrics (balance, throttle, steal — on fly the prometheus API takes Authorization: FlyV1 <token>), and attribute the burn with a per-process /proc/[pid]/stat diff across the window — request logs can miss it entirely when the cost lives in a streaming loop rather than a handler (here: a client re-pulling a 38MB firehose replay from the same cursor every 40 seconds, forever).

sources #

  • a hetzner rightsizing, 2026-08: €629/mo dedicated → €18.49/mo shared, after fixing the index structure that accounted for 51 of 53.6 GiB
  • zds-pds on fly shared-cpu-1x, 2026-08: half-hourly credit exhaustion from a cursor-stuck firehose consumer; a live demo landed inside a throttle window and the PDS (oauth, createSpace) appeared dead while serving zero errors