diff --git a/docs/handoffs/HANDOFF-2026-08-01-pg-dep-bump.md b/docs/handoffs/HANDOFF-2026-08-01-pg-dep-bump.md new file mode 100644 index 0000000..7777db6 --- /dev/null +++ b/docs/handoffs/HANDOFF-2026-08-01-pg-dep-bump.md @@ -0,0 +1,85 @@ +# operator handoff 2026-08-01 — pg.zig dependency bump, ready to deploy + +`main` is at `03391c8`. One dependency change plus a three-line call-site +edit. No zlay logic changed. Safe to deploy in my judgement, with one +unverified item called out below. + +## what changed + +`build.zig.zon` moves pg.zig from `github.com/zzstoatzz/pg.zig@dev` (5ce2355) +to `tangled.org/zzstoatzz.io/pg.zig` at `c5a5607`. + +The old pin was an independent zig-0.16 port that diverged from upstream +(`karlseguin/pg.zig`) five months ago. The new pin is current upstream master +plus our test/CI work and one pool fix. + +Call sites: `Pool.initUri(io, allocator, ...)` instead of +`(allocator, io, ...)`, three places in `event_log.zig`. Argument order only. +`exec` / `query` / `row` / `stats` / `deinit` are unchanged, which covers the +other 23 usages. + +## why it is worth deploying + +- **`52b9f8a` fix double-unlock in the pool's reconnector.** zlay drives this + pool continuously; the old pin does not have this. +- **`4215baf` Pool owns all opts.** zlay passes `DATABASE_URL` through + `initUri` — precisely the pointer-retention path this fixes. +- **`c9213c2` / `234ded5`** propagate `error.Canceled` and shutdown errors + rather than swallowing them. +- **`95fbbf9` openssl becomes a stub module when disabled.** zlay passes no + openssl options, so pg no longer `@cImport`s `openssl/ssl.h` + unconditionally and translate-c is skipped entirely. One fewer thing the + cross-compiled production build depends on. + +## the risky-sounding part, and why it is not + +Upstream had rewritten `Pool.acquire` to bound its wait with an `Io.Select` +racing `Io.Condition.wait` against `Io.sleep`. `Select.concurrent` guarantees +a unit of concurrency — under `Io.Threaded`, a real OS thread — so every +waiter cost two extra threads per wait iteration. Measured on a harness: + +``` +1024 callers, 10-connection pool upstream as-is after fix + peak threads 3045 1028 + failed acquires 24-55 per run 0 + throughput ~1700 ops/sec ~9500 ops/sec + max acquire wait 10.0s (=timeout) ~2.0s +``` + +That was fixed on the fork *before* this bump (`c5a5607`), by reverting to +the monotonic-counter + `futexWaitTimeout` design — which is the same design +the old pin already runs in production. **So the hot path is not new code to +zlay.** For a process already running ~2,800 threads, that mattered. + +## what I verified + +- `zig build` and `zig build -Doptimize=ReleaseSafe`, both clean, native. +- Full suite against a real Postgres, on both the old and new pin: + **107 pass, 1 skip, identical.** The 10 DB-backed tests that skip without + `DATABASE_URL` were included. +- pg.zig's own suite on the fork: 119/119, plus all four + openssl x column_names build permutations. +- No exhaustive error switches in zlay, so the newly-propagated + `error.Canceled` cannot break error handling. + +## what I did NOT verify + +- **The `-Dtarget=x86_64-linux-gnu` production build.** I built natively + only. This is build-time risk, not silent runtime risk — it either + compiles or it does not, and you will know immediately. The openssl-stub + change should make this *more* robust, not less, but it is untested. +- Runtime behavior under production load. No canary was run. +- zlay's tangled CI reports "0 pipeline runs", so nothing was verified + independently of my machine. + +## rollback + +`git revert 03391c8` — one commit, dependency plus three lines. The previous +pin is unchanged and still fetchable. + +## what to watch after deploy + +- DB pool metrics: acquire timeouts / `PoolExhausted` should stay at zero. +- Thread count and RSS: should be indistinguishable from before. Any jump in + thread count is the signal that something regressed in the pool wait path. +- Reconnect behavior, since the reconnector double-unlock fix lands here.