# instance sizing size a cloud instance to a measured working set, not to an estimate. the estimate is almost always the sum of everyone's worst case, and the difference between it and the measurement is frequently a bug rather than a requirement. a service provisioned on a 32-core / 128 GB dedicated box measured at roughly one core and a 53.6 GiB resident set, of which about 51 GiB was a single oversized index structure. once the structure was fixed the working set was **2.4 GiB**, and the service moved to an 8-vcpu / 16 GB shared instance with six times the headroom. the compute premium existed only to house the bug. ## measuring the working set - **measure at a fresh process start.** allocators retain freed pages, so RSS measured after a fix overstates the requirement until the process restarts. - **separate resident from required.** page cache, arena retention, and preallocated buffers all inflate RSS without representing demand. - **measure the load you actually serve.** a peak during a backfill is a reason to rent capacity for the backfill, not to own it year-round. ## the billing asymmetries worth designing around **rescale is up-only for disks, reversible for cpu and ram.** growing a root disk forecloses ever moving to a smaller instance type; a cpu/ram-only rescale round-trips (on hetzner, `upgrade_disk=false`). therefore **start small and rescale up on evidence**. being too small costs minutes of downtime once; being too large costs every month until someone re-measures. **hourly billing makes insurance nearly free.** keeping the old instance one extra night as a rollback target costs about an hour's rate per hour. upsizing for a one-day heavy job and back costs single-digit currency units. these are scratch allocations, not commitments — price them accordingly. **read the whole bill before celebrating.** storage frequently is the product. a 3 TB volume dwarfed the right-sized server that replaced the old one, which turned a 97% saving on the server line into a 71% saving overall. identify the dominant line item before optimizing anything. **shared vcpu is adequate for rate-limited and I/O-bound work.** a polite network crawl or a one-core serving load is not measurably affected by contention. reserve dedicated cores for sustained cpu-bound phases, and rent them only for the duration of those phases. **burstable vcpu fails as throttling, not saturation — and the process is not the one that notices.** a fly shared-cpu-1x accrues credits against a 1/16-core baseline; ~0.15 cores sustained for ten minutes drains the whole balance (`fly_instance_cpu_balance` → 0) and the hypervisor pins the machine to baseline — `steal` jumps to 50%+ while guest `user`+`system` stay small. from inside, the app is idle and error-free; from outside, every request crawls (an oauth page went 1.6s, upstream proxy fetches hit 10s timeouts). diagnose from the host metrics (balance, throttle, steal — on fly the prometheus API takes `Authorization: FlyV1 `), and attribute the burn with a per-process `/proc/[pid]/stat` diff across the window — request logs can miss it entirely when the cost lives in a streaming loop rather than a handler (here: a client re-pulling a 38MB firehose replay from the same cursor every 40 seconds, forever). ## related - [home-infra/](./home-infra/) — the same discipline at house scale, where the cost curve is capital rather than rent - [throughput-bottlenecks](./throughput-bottlenecks.md) — an oversized instance is often a serial-core diagnosis that was never made - [bounded-scans](./bounded-scans.md) ## sources - a hetzner rightsizing, 2026-08: €629/mo dedicated → €18.49/mo shared, after fixing the index structure that accounted for 51 of 53.6 GiB - zds-pds on fly shared-cpu-1x, 2026-08: half-hourly credit exhaustion from a cursor-stuck firehose consumer; a live demo landed inside a throttle window and the PDS (oauth, createSpace) appeared dead while serving zero errors