subscriber: keepalive pinger detects half-open upstream connections under zio master
the 2026-08-18 and 2026-08-27 wedges were mass half-open TCP: the network path died without FIN/RST, the kernel kept ~2,900 connections ESTABLISHED, and level-triggered epoll stayed silent forever. the existing readLoopWithHeartbeat covers this under Io.Threaded, but its SO_RCVTIMEO trigger is structurally inert under zio — netRead parks the fiber in epoll and the socket timeout never fires — so production had no liveness at all. under zio the read loop now runs as a cancelable task while the subscriber fiber drives a ping loop: silent past the interval → ping; max_failures unanswered pings → cancel the reader and reconnect. teardown is fiber cancellation, not a cross-fiber socket close, so the reader observes cancel_requested before ever touching the fd again (no closed-fd-reuse race) and the fd is closed exactly once by client.deinit. sending the ping also arms the kernel's own tcp_retries2 backstop, which pure reading never does. silence alone never kills a connection (only ~120 of ~3,000 hosts deliver in a given 30s; a silence-judging sweep caused the 2026-08-19 outage): the pong is the discriminator, any received frame resets the count, and a reader blocked inside serverMessage (rate-limit/backpressure, by design) suspends judgement entirely. policy is a pure decide() function with tests. knobs (RELAY_WS_PING_ENABLED/INTERVAL_SEC/MAX_FAILURES, defaults on/30/4 matching indigo) are runtime-settable via GET/POST /admin/ws-ping, not persisted; {"enabled":false} is the kill switch. new counters: relay_ws_pings_sent_total, relay_ws_ping_timeout_closes_total, relay_ws_ping_write_failures_total. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>