benchmarking python event loops #
one percentage from one in-process run is mostly a story about that process. event-loop microbenchmarks are sensitive to cyclic GC timing, allocator history, run order, kernel buffer configuration, and whether every contender did the same useful work.
harness rules #
- run warmups separately from measurement;
- run every measured sample in a fresh python process — installing multiple event-loop policies sequentially in one interpreter gives later loops a different allocator and GC history, and a fresh worker also contains crashes and hangs;
- rotate contender order between rounds;
- assert a parity counter or checksum before accepting a timing — "same primitive" is not parity when the work or the kernel configuration differs (e.g. socket buffer sizes set by one runtime and not the other);
- report median, min, max, and coefficient of variation;
- preserve timeouts and errors in the result instead of dropping them — compatibility failures are benchmark results too;
- record python, package, OS, CPU, GC, iteration, and payload provenance.
span different mechanisms rather than repeating one scheduler benchmark: scheduler ops (call_soon, timer cancellation, sleep(0), task fanout, call_soon_threadsafe), network (TCP request/response, bulk streams, getaddrinfo), and process (spawn/exit, piped bulk I/O).
keep the headline number boring: default runtime settings, identical work, fresh processes, interleaved order. GC-disabled runs, isolated probes, and buffer inspection are diagnostics that explain a result — never merge them into the headline.
watch for name collisions: astral uv (rust package tool) is unrelated to libuv or uvloop; a benchmark that shells out to uv run ... is using it as a subprocess workload, not as an event loop.
measured example #
zuv (a zig/libuv asyncio loop) vs uvloop on an apple M5 pro, python 3.14.5, two warmups and five fresh-process samples per case, GC enabled: 0.94x on 1 KiB TCP echo up to 2.05x on timer cancellation; sleep(0) and process spawn/exit were ties; bulk TCP 1.28x. an early in-process run had reported 16.8% faster on the scheduler — bimodal around GC collection timing — while a GC-disabled diagnostic showed the scheduler-only region ~43% faster; neither number belongs in the headline. the piped-I/O case (1,492 vs 671 MiB/s, ~6% CoV) only materialized after the adopted socketpairs matched libuv's 64 KiB buffer requests — a parity bug, not a speedup. rloop and rsloop won several scheduler microbenchmarks but failed or hung TCP and pipe workloads in the same run.
related low-level failure: libuv process lifetimes