jetstream v2 in zig stream.waow.tech
Something went wrong. Try again.
3.8 kB · 69 lines
YAML
at main
12345678910111213141516171819202122232425262728293031323334353637383940414243444546474849505152535455565758596061626364656667686970# Alert rules for the stream process. The production Prometheus lives in the# docker compose stack at /opt/stream-experiment/ on the box; copy this file# into its rule_files path. This is the in-repo source of truth for the rules.## Alertmanager on the box routes severity critical and warning to Discord and# drops info (alertmanager.yml, deploy/README.md "paging").groups: - name: stream rules: # A restart takes the target down for ~4.5 minutes of startup, so any # changes()/resets() window shorter than that sees one sample and # evaluates to 0 (docs/gotchas.md). Keying off the absolute start time # instead: this is true from the first scrape after startup, for the # following 15 minutes, then clears. Exactly one firing per restart, # independent of the scrape interval, never stuck pending. - alert: StreamProcessRestarted expr: process_start_time_seconds{job="stream"} > time() - 900 labels: severity: info annotations: summary: "stream process started within the last 15 minutes" description: "process_start_time_seconds is {{ $value | humanizeTimestamp }}. A single restart at a batch boundary is expected; see StreamCrashLoop for repeats."
# A single restart is by design (batch boundary). Two or more inside an # hour is not. changes() over 1h counts distinct start times minus one, # so >= 2 means at least two restarts landed within the window. - alert: StreamCrashLoop expr: changes(process_start_time_seconds{job="stream"}[1h]) >= 2 labels: severity: critical annotations: summary: "stream restarted at least twice in the last hour" description: "{{ $value }} start-time changes in 1h. Check the container logs for the exit reason before the replay window lapses."
# Live-tail liveness, distinct from "process up": the gauge is the wall # clock of the last upstream event the live consumer processed. It is 0 # until the first event of a run, hence the > 0 guard so a fresh process # does not trip it during bootstrap. - alert: StreamLiveTailStalled expr: | (time() - jetstream_livestream_last_seen_upstream_event_timestamp_seconds{job="stream"} > 600) and jetstream_livestream_last_seen_upstream_event_timestamp_seconds{job="stream"} > 0 for: 2m labels: severity: critical annotations: summary: "no upstream event processed for over 10 minutes" description: "last upstream event was {{ $value | humanizeDuration }} ago while the process is up. Check jetstream_livestream_reconnects_total and the relay."
# StreamLiveTailStalled cannot see a trickle: ~18 events/s keeps its # gauge fresh (2026-10-01, 80 minutes at 4% of normal). The quietest # 30-minute average in the week before was ~220/s. - alert: StreamUpstreamTrickle expr: sum(rate(jetstream_livestream_events_received_total{job="stream"}[5m])) < 100 for: 10m labels: severity: critical annotations: summary: "live ingest under 100 events/s for 10 minutes" description: "ingest is {{ $value | humanize }} events/s. Compare the relay on relay-eval and check stream_upstream_slow_reconnects_total; a reconnect usually lands on a healthy connection."
- alert: StreamTargetDown expr: up{job="stream"} == 0 for: 10m labels: severity: warning annotations: summary: "stream metrics target unreachable for 10 minutes" description: "longer than the ~4.5 minute startup window, so this is not a routine batch-boundary restart."