performance #
measured 2026-09-22 against fastmcp's JevSearchTransform over a 187-tool catalog, from chicago. the transform makes two round trips per search. first it sends a wide pass that ranks the catalog in 150-tool chunks, in parallel. then it sends a close read: one choice over the shortlist's full docs, plus a yes/no question per candidate.
where a search's time goes #
per-request timings from the client (httpx event hooks) and from the container (the forward and handler_ms log lines), median of 8 queries:
| jev (api.typesafe.ai) | kev-4b, H100 | kev-4b, L4 | |
|---|---|---|---|
| transform work before the first request | 125 ms | 125 ms | 130 ms |
| wide pass round trip | ~170 ms | ~300 ms | ~950 ms |
| close read round trip | ~155 ms | ~430 ms | ~1,050 ms |
| search | 0.48 s | 0.88 to 0.93 s | 2.3 s |
jev is served from AWS us-west-2 and Modal's ingress is in us-east-1 (checked against AWS's published ranges). jev is farther away and still faster, so distance is not the difference. for a kev close read on the H100:
| ms | |
|---|---|
| client round trip | ~390 |
Modal's duration for the request |
384 |
Modal's execution for the request |
317 |
| our handler | 241 |
| model | 215 |
- Modal adds ~135 ms per request. about 65 ms goes between ingress and the container, and about 70 ms goes inside the container but outside our handler, in Modal's ASGI adapter. a trivial
/healthspends ~25 ms in that adapter. - kev's model time has a floor of ~60 ms per forward pass on an H100, whatever the size. a 150-tool chunk (2.7k tokens) takes 75 ms, and a 37-tool chunk (0.7k tokens) takes 64 ms. the close read needs three forwards: the state, then one per length group. that makes ~215 ms, more than jev's whole request.
- the two wide chunks run one after the other. they share a state and the coalescer would merge them, but the client opens a second connection for the second request, so it arrives ~70 ms after the first. by then the first forward has started, and kev has one lock.
so the remaining gap to jev is about half Modal and about half kev's per-forward cost. jev answers whole requests in under ~100 ms server-side. its stack isn't public.
what helped #
median seconds per search. the earlier runs used 50-tool chunks, the last ones 150:
| change | L4 | H100 |
|---|---|---|
| kev as released, 50-tool chunks | 5.7 | 1.3 |
| causal_conv1d loads, autotune cached, same-state merge | 5.7 | 1.1 |
| close-read rows pad per length group, state computed once | 2.0 | 1.5 |
| container pinned to a US region | 2.4 | 1.1 |
| 150-tool chunks | 2.3 | 0.88 |
runs of the same code vary by a few hundred ms with Modal's routing, so read single-row changes loosely.
- close-read padding was the largest cause. kev pads every question row to the longest. one choice over 8 full docs (~2k tokens) next to 16 one-line questions meant ~3k real tokens became ~20k padded: 4.7 s on an L4.
- region. unpinned, Modal put the H100 container in GCP asia-south2 while inputs enter through Virginia. a single 50-tool request took 770 ms against 310 ms once pinned to
us. region pinning costs 1.15x for a broad region and 1.75x for a narrow one likeus-east. - chunk size. 50-tool chunks made four wide requests, and their shortlists often overflowed into a third round trip. 150 matches jev's shape. the L4 only needed 50 while the reference conv ran.
- the conv kernel. Dao-AILab's prebuilt causal_conv1d wheel needs glibc 2.32. Modal's 2023.12 image builder ships 2.31, so the import failed and transformers fell back without an error. the justfile pins builder 2025.06.
- autotune cache.
TRITON_CACHE_AUTOTUNING=1with the triton cache on the volume cut cold start from 46 s to 31 s.
what didn't help #
- kev's short-state path,
forward_rows_batch. a 16-question request with a ~250-token state took 591 ms warm there, against ~220 ms through the prefix path. its first call at a new shape spent ~9.5 s compiling. multi-question requests now always go through the prefix path. - uvicorn behind
modal.web_serverinstead ofasgi_app. about 100 ms slower per search in one run, so it was dropped. - merging the wide chunks on an L4. the L4 is compute-bound, so one merged pass costs what the passes cost separately.
- a longer gather window (30 ms). the second chunk arrives later than that.
what's left #
- the per-forward floor. CUDA graphs or
torch.compilecould cut it, but that is kev-level work with dynamic shapes. - the Modal hop. a narrow
us-eastpin might save tens of ms at 1.75x. serving from a GPU box without Modal's input plane means paying for it always-on.modal.experimental.http_server(flash) skips the input plane, but its docs say not to use it unless Modal support says so. - serialized chunks. running two forwards at once on one GPU needs more than kev's single lock.
an H100 costs ~5x an L4 per hour and saves ~1.4 s per search here. neither GPU gets below Modal's ~135 ms per request.
cost #
Modal rates on 2026-09-22 (modal billing rates): L4 $0.80/h, H100 $3.95/h, CPU $0.0473 per core-hour, memory $0.008 per GiB-hour. a container is billed for its CPU and memory requests (2 cores, 16 GiB here), and a region pin multiplies usage by 1.15 for us. Modal's docs say it applies to usage pricing, so it's assumed to cover CPU and memory too. all-in, per running container:
| per hour | per 30-day month, always on | |
|---|---|---|
| L4 | ~$1.18 | ~$850 |
| H100 | ~$4.80 | ~$3,450 |
jev costs $42 per billion input tokens (jev-1.13, 2026-09-18) with output free. a search here sends ~9k input tokens, so it costs ~$0.0004, or $0.38 per thousand.
kev's cost per search depends on how busy the GPU is. one lock serializes forward passes, so a container's throughput is roughly one search per unit of model time:
| model time per search | searches per hour, fully busy | cost per search, fully busy | |
|---|---|---|---|
| jev | n/a | n/a | ~$0.0004 |
| kev-4b, L4 | ~1.5 s | ~2,400 | ~$0.0005 |
| kev-4b, H100 | ~0.35 s | ~10,000 | ~$0.0005 |
fully busy is the best case, and it still costs slightly more than jev per search. at lower load, scale-to-zero sets a floor instead. a search after idle pays about a minute of cold start plus the 5-minute scaledown_window, about $0.12 on an L4 (~300 jev searches) or $0.48 on an H100. keeping one container warm costs the always-on price.
this spike, measured with modal billing report --for today --show-resources: $2.17 for the L4 app and $4.89 for the H100 app. the jev baselines in the harness runs used about 1.3M input tokens per run, ~$0.05 each.
on cost alone, self-hosting kev on Modal does not beat jev for this workload at any volume. the reasons to self-host are elsewhere: control over the model, fine-tuning on your own labels (kev ships a training path), and keeping requests off a third-party API.