This repository has no description
systemone docs performance.md
6.9 kB
Markdown
at main

performance #

measured 2026-09-22 against fastmcp's JevSearchTransform over a 187-tool catalog, from chicago. the transform makes two round trips per search. first it sends a wide pass that ranks the catalog in 150-tool chunks, in parallel. then it sends a close read: one choice over the shortlist's full docs, plus a yes/no question per candidate.

where a search's time goes #

per-request timings from the client (httpx event hooks) and from the container (the forward and handler_ms log lines), median of 8 queries:

jev (api.typesafe.ai) kev-4b, H100 kev-4b, L4
transform work before the first request 125 ms 125 ms 130 ms
wide pass round trip ~170 ms ~300 ms ~950 ms
close read round trip ~155 ms ~430 ms ~1,050 ms
search 0.48 s 0.88 to 0.93 s 2.3 s

jev is served from AWS us-west-2 and Modal's ingress is in us-east-1 (checked against AWS's published ranges). jev is farther away and still faster, so distance is not the difference. for a kev close read on the H100:

ms
client round trip ~390
Modal's duration for the request 384
Modal's execution for the request 317
our handler 241
model 215
  • Modal adds ~135 ms per request. about 65 ms goes between ingress and the container, and about 70 ms goes inside the container but outside our handler, in Modal's ASGI adapter. a trivial /health spends ~25 ms in that adapter.
  • kev's model time has a floor of ~60 ms per forward pass on an H100, whatever the size. a 150-tool chunk (2.7k tokens) takes 75 ms, and a 37-tool chunk (0.7k tokens) takes 64 ms. the close read needs three forwards: the state, then one per length group. that makes ~215 ms, more than jev's whole request.
  • the two wide chunks run one after the other. they share a state and the coalescer would merge them, but the client opens a second connection for the second request, so it arrives ~70 ms after the first. by then the first forward has started, and kev has one lock.

so the remaining gap to jev is about half Modal and about half kev's per-forward cost. jev answers whole requests in under ~100 ms server-side. its stack isn't public.

what helped #

median seconds per search. the earlier runs used 50-tool chunks, the last ones 150:

change L4 H100
kev as released, 50-tool chunks 5.7 1.3
causal_conv1d loads, autotune cached, same-state merge 5.7 1.1
close-read rows pad per length group, state computed once 2.0 1.5
container pinned to a US region 2.4 1.1
150-tool chunks 2.3 0.88

runs of the same code vary by a few hundred ms with Modal's routing, so read single-row changes loosely.

  • close-read padding was the largest cause. kev pads every question row to the longest. one choice over 8 full docs (~2k tokens) next to 16 one-line questions meant ~3k real tokens became ~20k padded: 4.7 s on an L4.
  • region. unpinned, Modal put the H100 container in GCP asia-south2 while inputs enter through Virginia. a single 50-tool request took 770 ms against 310 ms once pinned to us. region pinning costs 1.15x for a broad region and 1.75x for a narrow one like us-east.
  • chunk size. 50-tool chunks made four wide requests, and their shortlists often overflowed into a third round trip. 150 matches jev's shape. the L4 only needed 50 while the reference conv ran.
  • the conv kernel. Dao-AILab's prebuilt causal_conv1d wheel needs glibc 2.32. Modal's 2023.12 image builder ships 2.31, so the import failed and transformers fell back without an error. the justfile pins builder 2025.06.
  • autotune cache. TRITON_CACHE_AUTOTUNING=1 with the triton cache on the volume cut cold start from 46 s to 31 s.

what didn't help #

  • kev's short-state path, forward_rows_batch. a 16-question request with a ~250-token state took 591 ms warm there, against ~220 ms through the prefix path. its first call at a new shape spent ~9.5 s compiling. multi-question requests now always go through the prefix path.
  • uvicorn behind modal.web_server instead of asgi_app. about 100 ms slower per search in one run, so it was dropped.
  • merging the wide chunks on an L4. the L4 is compute-bound, so one merged pass costs what the passes cost separately.
  • a longer gather window (30 ms). the second chunk arrives later than that.

what's left #

  • the per-forward floor. CUDA graphs or torch.compile could cut it, but that is kev-level work with dynamic shapes.
  • the Modal hop. a narrow us-east pin might save tens of ms at 1.75x. serving from a GPU box without Modal's input plane means paying for it always-on. modal.experimental.http_server (flash) skips the input plane, but its docs say not to use it unless Modal support says so.
  • serialized chunks. running two forwards at once on one GPU needs more than kev's single lock.

an H100 costs ~5x an L4 per hour and saves ~1.4 s per search here. neither GPU gets below Modal's ~135 ms per request.

cost #

Modal rates on 2026-09-22 (modal billing rates): L4 $0.80/h, H100 $3.95/h, CPU $0.0473 per core-hour, memory $0.008 per GiB-hour. a container is billed for its CPU and memory requests (2 cores, 16 GiB here), and a region pin multiplies usage by 1.15 for us. Modal's docs say it applies to usage pricing, so it's assumed to cover CPU and memory too. all-in, per running container:

per hour per 30-day month, always on
L4 ~$1.18 ~$850
H100 ~$4.80 ~$3,450

jev costs $42 per billion input tokens (jev-1.13, 2026-09-18) with output free. a search here sends ~9k input tokens, so it costs ~$0.0004, or $0.38 per thousand.

kev's cost per search depends on how busy the GPU is. one lock serializes forward passes, so a container's throughput is roughly one search per unit of model time:

model time per search searches per hour, fully busy cost per search, fully busy
jev n/a n/a ~$0.0004
kev-4b, L4 ~1.5 s ~2,400 ~$0.0005
kev-4b, H100 ~0.35 s ~10,000 ~$0.0005

fully busy is the best case, and it still costs slightly more than jev per search. at lower load, scale-to-zero sets a floor instead. a search after idle pays about a minute of cold start plus the 5-minute scaledown_window, about $0.12 on an L4 (~300 jev searches) or $0.48 on an H100. keeping one container warm costs the always-on price.

this spike, measured with modal billing report --for today --show-resources: $2.17 for the L4 app and $4.89 for the H100 app. the jev baselines in the harness runs used about 1.3M input tokens per run, ~$0.05 each.

on cost alone, self-hosting kev on Modal does not beat jev for this workload at any volume. the reasons to self-host are elsewhere: control over the model, fine-tuning on your own labels (kev ships a training path), and keeping requests off a third-party API.