jetstream client
atproto jetstream client
jetstream docs failover.md
2.9 kB
Markdown

multi-host failover #

Timestamp-based failover is available starting in v0.1.5.

Live failover reconnects with cursor = last_time_us - failover_rewind_us (default ten seconds). Jetstream v2 accepts unix-microsecond timestamps on subscribeEvents and resolves them into the receiving instance's sequence space. No archive lookup or API key is required. This is the server's documented timestamp cursor behavior.

Sequence deduplication resets on rotation. After receiving events, reconnects to that host use its last sequence. The overlap can redeliver events; consumers must be idempotent. Witness times differ between instances, so the margin tolerates some skew, not arbitrary coverage differences or resyncs.

rotation and failback #

failover_after_errors consecutive live failures without an event trigger rotation. An event resets the counter. A rotation after delivered events retries the primary first; unsuccessful attempts advance through hosts. Retries use bounded exponential backoff. A single host or failover_after_errors = 0 disables rotation.

A sequence supplied through after_seq or live_cursor belongs to the primary. If that host fails before delivering any event, there is no portable time anchor: the client returns HostUnhealthy. A fresh subscription with no cursor can try another host's tip before its first delivery. A caller-supplied timestamp is already portable and transfers unchanged until the first event; it is not rewound again on each failed connection.

retention and archive replay #

Timestamp resume is limited by the receiving server's live retention. A clamped timestamp emits OutdatedCursor, forwarded through onInfo or logged if the handler has no callback. This signals a possible gap; callers requiring complete history must handle it explicitly. A rejected live sequence returns CursorTooOld, never silently falling back to the tip. Rewind margins must be nonnegative and yield a valid timestamp cursor.

Explicit after_seq requests still replay the archive and cut over to live on the original host. If that live tail later fails over, the new host is entered through its live timestamp cursor. Archive failures themselves never rotate hosts: accepted recoverable errors continue on the same archive; plan failures remain terminal. Archive credentials are only needed for explicit archive work, including the same-host re-backfill path.

bounded live smoke #

After zig build, run zig-out/bin/example-failover-smoke <host> with an external 30-second timeout. It rotates from a refused local connection to the real host and stops after ten events. Add --rewind to resume from ten seconds ago, exercising timestamp transfer before the first event. Neither mode requests archive history or needs an API key.