instance sizing #
size a cloud instance to a measured working set, not to an estimate. the estimate is almost always the sum of everyone's worst case, and the difference between it and the measurement is frequently a bug rather than a requirement.
a service provisioned on a 32-core / 128 GB dedicated box measured at roughly one core and a 53.6 GiB resident set, of which about 51 GiB was a single oversized index structure. once the structure was fixed the working set was 2.4 GiB, and the service moved to an 8-vcpu / 16 GB shared instance with six times the headroom. the compute premium existed only to house the bug.
measuring the working set #
- measure at a fresh process start. allocators retain freed pages, so RSS measured after a fix overstates the requirement until the process restarts.
- separate resident from required. page cache, arena retention, and preallocated buffers all inflate RSS without representing demand.
- measure the load you actually serve. a peak during a backfill is a reason to rent capacity for the backfill, not to own it year-round.
the billing asymmetries worth designing around #
rescale is up-only for disks, reversible for cpu and ram. growing a root
disk forecloses ever moving to a smaller instance type; a cpu/ram-only rescale
round-trips (on hetzner, upgrade_disk=false). therefore start small and
rescale up on evidence. being too small costs minutes of downtime once;
being too large costs every month until someone re-measures.
hourly billing makes insurance nearly free. keeping the old instance one extra night as a rollback target costs about an hour's rate per hour. upsizing for a one-day heavy job and back costs single-digit currency units. these are scratch allocations, not commitments — price them accordingly.
read the whole bill before celebrating. storage frequently is the product. a 3 TB volume dwarfed the right-sized server that replaced the old one, which turned a 97% saving on the server line into a 71% saving overall. identify the dominant line item before optimizing anything.
shared vcpu is adequate for rate-limited and I/O-bound work. a polite network crawl or a one-core serving load is not measurably affected by contention. reserve dedicated cores for sustained cpu-bound phases, and rent them only for the duration of those phases.
burstable vcpu fails as throttling, not saturation — and the process is
not the one that notices. a fly shared-cpu-1x accrues credits against a
1/16-core baseline; ~0.15 cores sustained for ten minutes drains the whole
balance (fly_instance_cpu_balance → 0) and the hypervisor pins the machine
to baseline — steal jumps to 50%+ while guest user+system stay small.
from inside, the app is idle and error-free; from outside, every request
crawls (an oauth page went 1.6s, upstream proxy fetches hit 10s timeouts).
diagnose from the host metrics (balance, throttle, steal — on fly the
prometheus API takes Authorization: FlyV1 <token>), and attribute the burn
with a per-process /proc/[pid]/stat diff across the window — request logs
can miss it entirely when the cost lives in a streaming loop rather than a
handler (here: a client re-pulling a 38MB firehose replay from the same
cursor every 40 seconds, forever).
related #
- home-infra/ — the same discipline at house scale, where the cost curve is capital rather than rent
- throughput-bottlenecks — an oversized instance is often a serial-core diagnosis that was never made
- bounded-scans
sources #
- a hetzner rightsizing, 2026-08: €629/mo dedicated → €18.49/mo shared, after fixing the index structure that accounted for 51 of 53.6 GiB
- zds-pds on fly shared-cpu-1x, 2026-08: half-hourly credit exhaustion from a cursor-stuck firehose consumer; a live demo landed inside a throttle window and the PDS (oauth, createSpace) appeared dead while serving zero errors