# performance measured 2026-09-22 against fastmcp's `JevSearchTransform` over a 187-tool catalog, from chicago. the transform makes two round trips per search. first it sends a wide pass that ranks the catalog in 150-tool chunks, in parallel. then it sends a close read: one choice over the shortlist's full docs, plus a yes/no question per candidate. ## where a search's time goes per-request timings from the client (httpx event hooks) and from the container (the `forward` and `handler_ms` log lines), median of 8 queries: | | jev (api.typesafe.ai) | kev-4b, H100 | kev-4b, L4 | |---|---|---|---| | transform work before the first request | 125 ms | 125 ms | 130 ms | | wide pass round trip | ~170 ms | ~300 ms | ~950 ms | | close read round trip | ~155 ms | ~430 ms | ~1,050 ms | | search | 0.48 s | 0.88 to 0.93 s | 2.3 s | jev is served from AWS us-west-2 and Modal's ingress is in us-east-1 (checked against AWS's published ranges). jev is farther away and still faster, so distance is not the difference. for a kev close read on the H100: | | ms | |---|---| | client round trip | ~390 | | Modal's `duration` for the request | 384 | | Modal's `execution` for the request | 317 | | our handler | 241 | | model | 215 | - **Modal adds ~135 ms per request.** about 65 ms goes between ingress and the container, and about 70 ms goes inside the container but outside our handler, in Modal's ASGI adapter. a trivial `/health` spends ~25 ms in that adapter. - **kev's model time has a floor of ~60 ms per forward pass on an H100,** whatever the size. a 150-tool chunk (2.7k tokens) takes 75 ms, and a 37-tool chunk (0.7k tokens) takes 64 ms. the close read needs three forwards: the state, then one per length group. that makes ~215 ms, more than jev's whole request. - **the two wide chunks run one after the other.** they share a state and the coalescer would merge them, but the client opens a second connection for the second request, so it arrives ~70 ms after the first. by then the first forward has started, and kev has one lock. so the remaining gap to jev is about half Modal and about half kev's per-forward cost. jev answers whole requests in under ~100 ms server-side. its stack isn't public. ## what helped median seconds per search. the earlier runs used 50-tool chunks, the last ones 150: | change | L4 | H100 | |---|---|---| | kev as released, 50-tool chunks | 5.7 | 1.3 | | causal_conv1d loads, autotune cached, same-state merge | 5.7 | 1.1 | | close-read rows pad per length group, state computed once | 2.0 | 1.5 | | container pinned to a US region | 2.4 | 1.1 | | 150-tool chunks | 2.3 | 0.88 | runs of the same code vary by a few hundred ms with Modal's routing, so read single-row changes loosely. - **close-read padding** was the largest cause. kev pads every question row to the longest. one choice over 8 full docs (~2k tokens) next to 16 one-line questions meant ~3k real tokens became ~20k padded: 4.7 s on an L4. - **region.** unpinned, Modal put the H100 container in GCP asia-south2 while inputs enter through Virginia. a single 50-tool request took 770 ms against 310 ms once pinned to `us`. region pinning costs 1.15x for a broad region and 1.75x for a narrow one like `us-east`. - **chunk size.** 50-tool chunks made four wide requests, and their shortlists often overflowed into a third round trip. 150 matches jev's shape. the L4 only needed 50 while the reference conv ran. - **the conv kernel.** Dao-AILab's prebuilt causal_conv1d wheel needs glibc 2.32. Modal's 2023.12 image builder ships 2.31, so the import failed and transformers fell back without an error. the justfile pins builder 2025.06. - **autotune cache.** `TRITON_CACHE_AUTOTUNING=1` with the triton cache on the volume cut cold start from 46 s to 31 s. ## what didn't help - **kev's short-state path, `forward_rows_batch`.** a 16-question request with a ~250-token state took 591 ms warm there, against ~220 ms through the prefix path. its first call at a new shape spent ~9.5 s compiling. multi-question requests now always go through the prefix path. - **uvicorn behind `modal.web_server`** instead of `asgi_app`. about 100 ms slower per search in one run, so it was dropped. - **merging the wide chunks on an L4.** the L4 is compute-bound, so one merged pass costs what the passes cost separately. - **a longer gather window (30 ms).** the second chunk arrives later than that. ## what's left - **the per-forward floor.** CUDA graphs or `torch.compile` could cut it, but that is kev-level work with dynamic shapes. - **the Modal hop.** a narrow `us-east` pin might save tens of ms at 1.75x. serving from a GPU box without Modal's input plane means paying for it always-on. `modal.experimental.http_server` (flash) skips the input plane, but its docs say not to use it unless Modal support says so. - **serialized chunks.** running two forwards at once on one GPU needs more than kev's single lock. an H100 costs ~5x an L4 per hour and saves ~1.4 s per search here. neither GPU gets below Modal's ~135 ms per request. ## cost Modal rates on 2026-09-22 (`modal billing rates`): L4 $0.80/h, H100 $3.95/h, CPU $0.0473 per core-hour, memory $0.008 per GiB-hour. a container is billed for its CPU and memory requests (2 cores, 16 GiB here), and a region pin multiplies usage by 1.15 for `us`. Modal's docs say it applies to usage pricing, so it's assumed to cover CPU and memory too. all-in, per running container: | | per hour | per 30-day month, always on | |---|---|---| | L4 | ~$1.18 | ~$850 | | H100 | ~$4.80 | ~$3,450 | jev costs $42 per billion input tokens (jev-1.13, 2026-09-18) with output free. a search here sends ~9k input tokens, so it costs ~$0.0004, or $0.38 per thousand. kev's cost per search depends on how busy the GPU is. one lock serializes forward passes, so a container's throughput is roughly one search per unit of model time: | | model time per search | searches per hour, fully busy | cost per search, fully busy | |---|---|---|---| | jev | n/a | n/a | ~$0.0004 | | kev-4b, L4 | ~1.5 s | ~2,400 | ~$0.0005 | | kev-4b, H100 | ~0.35 s | ~10,000 | ~$0.0005 | fully busy is the best case, and it still costs slightly more than jev per search. at lower load, scale-to-zero sets a floor instead. a search after idle pays about a minute of cold start plus the 5-minute `scaledown_window`, about $0.12 on an L4 (~300 jev searches) or $0.48 on an H100. keeping one container warm costs the always-on price. this spike, measured with `modal billing report --for today --show-resources`: $2.17 for the L4 app and $4.89 for the H100 app. the jev baselines in the harness runs used about 1.3M input tokens per run, ~$0.05 each. on cost alone, self-hosting kev on Modal does not beat jev for this workload at any volume. the reasons to self-host are elsewhere: control over the model, fine-tuning on your own labels (kev ships a training path), and keeping requests off a third-party API.