title: retention #
hydrant keeps two things on a clock: stream events, bounded by EPHEMERAL_TTL in ephemeral indexers and in relay mode, and in permanent databases the replaced record versions in history, bounded by HISTORY_TTL.
record history #
when a record is updated or deleted, its previous body moves to the history keyspace, keyed by the record and the time it was replaced, so replaying an older event can still show the version it was about. a delete of a record hydrant never saw leaves an empty entry, so "deleted" and "never existed" stay different.
without HISTORY_TTL history is kept forever. with it, entries older than the ttl are dropped by a compaction filter, so they go when compaction next rewrites their part of the keyspace, not right at the deadline. replaying an event whose version was trimmed gives the event without its record body, and /stats counts those as counts.history_trim_misses.
counts.history in /stats is approximate. it's read from the storage engine's table metadata rather than kept as a counter, so it does go down when retention or a repo erase removes entries, but only once compaction has rewritten the tables holding them. until then removed and overwritten entries still count. read it as about how many versions are on disk, not an exact number.
stream events #
an indexer only expires its events when EPHEMERAL is on, while a relay always expires its subscribeRepos events. both work the same way. events are numbered in commit order, and once a minute the ttl worker writes a watermark into the cursors keyspace saying "at this second, the next event id was N", then looks for the newest watermark at least EPHEMERAL_TTL old and drops every event below its id, deleting the watermarks it has used up. the watermarks only record time and position, not the ttl, so the ttl in effect is whatever hydrant was started with, applied fresh on every tick against whatever watermarks are on disk.
in ephemeral databases an event carries its record body inline, so that body lives exactly as long as its event: it replays with the event while the event is kept and goes in the same drop. nothing else (no record head, history entry or body archive) holds a copy.
changing EPHEMERAL_TTL across a restart #
- shrinking takes effect on the first tick, about a minute after startup. the newest watermark past the new, shorter cutoff picks where to cut, and everything older goes in one go.
- growing never brings anything back, since already pruned events are gone. what's still stored is kept until it's older than the new ttl, so pruning pauses until the oldest remaining watermark ages past it, and from then on the window is the new length.
- either way the cut is only as fine as the watermarks: one a minute while hydrant runs and none while it's down. if no watermark falls exactly at the cutoff the newest one before it is used, so hydrant keeps a little more, never less. after downtime that can be up to the length of the gap.
boundary lag #
events are dropped a whole table at a time (drop_range), so a storage table holding both events older than the cutoff and newer ones stays until a later cutoff passes its newest event. expect a bit more than EPHEMERAL_TTL worth of events on disk, by about one table, and expired events in that table still replay with their bodies until it goes. counts in /stats are approximate for the same reason. jetstream replay metadata is pruned on the same clock, by timestamp rather than watermark, with the same table granularity.