Stabilization detection in dltb — reverse convergence detection #
Reference documentation for dltb-stabilization (src/dltb/detect_stabilization.py):
the algorithm, its verified behavior, and its limits. The claims marked
verified below are pinned by the seeded stress battery in
tests/test_detect_stabilization.py (plain Python, CPU-only; a real-run
regression case skips itself when that run's CSV is not present).
Status at time of writing (2026-09-14): the tool is a faithful, pandas-free port of a reference implementation; its numerics reproduce that implementation's documented results exactly, and the battery below is the reconstructed validation of its robustness claims.
1. Purpose and I/O #
The analysis tools leave behind metric CSVs — distance_metrics.csv
(dltb-distance, dltb-distance-loop, dltb-distance-feedback) and
drift_metrics.csv (analyze_drift.py) — one row per frame, one column per
metric. dltb-stabilization reads such a CSV and answers, per column:
did the series stabilize, at which frame, and at what value — or why it
deserves human attention instead.
- First column = frame label; every other column is analysed independently;
non-numeric / non-finite cells are dropped per column, with labels kept
aligned to the surviving values (so
t_stabindexes and frame labels cannot drift apart). - Output: a verdict table on stdout,
stabilization_report.csvnext to the input (-ooverrides), optional--plotdiagnostic PNG (one panel per column). Exit code is always 0 — sweeps grep the table forneeds_attentionrather than checking status. - CPU-only: numpy + stdlib CSV; matplotlib is imported only when
--plotis given.
2. The conceptual frame #
The classic mistake is scanning forward and asking "has it settled yet?" at each sample — that forces a burn-in guess before you know the answer. This tool flips the question: assume the end of the run is stationary, characterize that regime, then ask how far back it extends.
One backward rule then handles every transient shape with no per-shape logic: rapid rise, spike-then-fall, and two-plateaus-with-a-jump are all "walk back to the last sample outside the final regime's band". The single real assumption is that the tail actually is stationary (§8.1).
3. The pipeline, per column #
Defaults in parentheses.
- Sample gate. Fewer than 20 valid samples →
too_few_samples, verdictneeds_attention, not_stab. - Smoothing. Centered rolling median, window
w = max(3, round(0.01·N))forced odd (≈1% of the run). A rolling median annihilates isolated spikes — any burst of ≤ ⌊w/2⌋ consecutive outlying samples cannot move the smoothed series at all; wider bursts survive and (correctly) pusht_stabpast them. - Tail reference. Over the last
max(10, round(tail_frac·N))samples (0.40): medianmedand robust scaleσ = 1.4826·MAD(the 1.4826 makes MAD consistent with standard deviation for Gaussian noise). A degenerate tail (σ ≤ 10⁻¹²·(1+|med|)) →constant_series, stabilized at the median from sample 0,stderr = 0. - Band.
[med − k_band·σ, med + k_band·σ](k = 5). This band is the practical tolerance — what "settled" means on the chart (§6). - Backward scan.
t_stab = (index of last smoothed sample outside the band) + 1, or 0 if none; clamped into[0, N−1]. The plateau is the raw (unsmoothed) samplesy[t_stab:]. If even the final sample is out of band (still moving at the end), the clamp leaves a ~1-sample plateau →plateau_too_short: "run ended before settling". - Diagnostics (§4), then the verdict:
stabilized⇔ no flags beyond the informationalflat_from_start(constant_seriesshort-circuits to stabilized in step 3).
4. The plateau diagnostics #
Any failure flags the column needs_attention. Thresholds for the drift and
shift tests are in band units (§6); the correlation and spectral tests
are scale-free.
| Flag | Test | Catches |
|---|---|---|
plateau_too_short |
P < max(10, 0.15·N) |
run ended before settling (still rising/drifting at the end) |
residual_drift |
Theil–Sen slope of the plateau's second half: ` | slope |
level_shift |
` | median(h1) − median(h2) |
wandering |
lag-1 autocorr r1 > 0.95, or r1 > 0.90 and variance ratio > 0.75 |
random-walk-like meandering that never converges |
alternating_oscillation |
r1 < −0.5 |
period-2 limit cycle |
periodic_oscillation |
largest single-frequency share of variance > 0.5 |
limit cycle at any period |
constant_series |
zero tail variance | informational; always stabilized |
r1 is the lag-1 autocorrelation of the centered plateau. The variance
ratio is VR(q) = Var(z_{t+q} − z_t) / (q·Var(z_{t+1} − z_t)) with
q = min(8, max(4, P/4)): ≈1 for a random walk (differences scale linearly),
≪1 for stationary series (mean reversion anti-correlates differences). The
spectral share is the max-bin fraction of the detrended, Hann-windowed
periodogram. The statistical tests require P ≥ 12; below that only
plateau_too_short can fire.
Theil–Sen (median of pairwise slopes, on a ≤200-point subsample) is used
for drift because it is robust both to noise and to oscillation — the
property that keeps residual_drift alive at all (§8.2).
5. Reported value and uncertainty #
stabilized_value = median of the plateau. stderr = robust plateau sigma
divided by √n_eff with n_eff = P·(1−ρ)/(1+ρ), ρ = clip(r1, −0.9, 0.9):
strongly autocorrelated samples carry less independent information, and the
uncertainty is widened accordingly. For a constant series both are the
median and 0.
6. Tolerance philosophy #
The band is deliberately the practical tolerance, not a statistical one. The failure this avoids: a very quiet metric column (tiny σ) makes any σ-scaled significance test fire on wiggles invisible on the chart — statistical significance ≠ practical significance. Expressing the drift and shift thresholds in band units (multiples of 5σ, i.e. of what you'd accept as "the same level") makes them fire only on departures a human would care about. The flip side: the correlation and spectral tests are scale-free by nature, so no band Units can protect a highly persistent series from them — that boundary is §8.4.
7. Verified behavior (stress battery) #
All numbers from tests/test_detect_stabilization.py (seeded → exact).
| Shape | Result |
|---|---|
| exp rise → plateau (N=600, τ=40) | stabilized, t_stab=125 (theory: 40·ln 20 ≈ 120) |
| baseline + 2-sample spike at t=50 | stabilized, t_stab=52 — spike located, not fooled by |
| two plateaus, jump at t=200 of 500 | stabilized, t_stab=200, value = second level |
| noisy linear ramp, N=400 (ends t=100) | stabilized, t_stab=89, value ≈ ramp top |
| constant series | stabilized + constant_series |
| monotone rise, never settles | needs_attention (level_shift, wandering) |
| flat, then ramp into the last frame | needs_attention (plateau_too_short) |
| slow in-band creep under oscillation | needs_attention (residual_drift + co-fires, §8.2) |
| sine limit cycle | needs_attention (wandering, periodic_oscillation) |
| alternating (± every sample) | needs_attention (alternating_oscillation, periodic_oscillation) |
| N=15 | too_few_samples |
| random walks ×50 | 49/50 flagged |
| stationary AR(1) ×50 | 0/50 false alarms at ρ=0.8 and ρ=0.85 |
| stationary AR(1), ρ=0.9 ×50 | 12/50 false alarms (§8.4) |
| real sd-turbo feedback run (201 samples, 3 DreamSim series) | all stabilized, t_stab = 6/13/5, values matching the values established when the tool was developed |
8. Limitations #
Each of these was established empirically, not guessed.
- The tail assumption is load-bearing.
tail_frac(0.40) must be shorter than the slowest expected transient. If the run is still in its transient at the end, the band is calibrated on moving data: the verdict becomesplateau_too_short(the good outcome — the designed behavior), but a transient that happens to move less than the widened band can pass as stabilized with an inflatedstderr. A crop ofplateau_too_shortverdicts means the runs are too short — extend the runs rather than loosen thresholds. residual_driftcannot fire for monotone drift. For any monotone creep the tail MAD inflates with the drift itself and always outpaces the slope test: for a linear creep at rater, the tail (lengthL = 0.4·N) hasMAD ≈ r·L/4, soband ≈ 5·1.4826·r·L/4 ≈ 1.85·r·L, while the drift accumulated over the plateau's second half is at mostr·N/2 < r·L·1.25— the ratio never reaches the threshold of 1 band. Verified by a 4000-shape random search: the only hits (2) were drift mixed with oscillation, where Theil–Sen's robustness to oscillation lets the slope test see what the MAD band absorbs — and therelevel_shift/wandering/periodic_oscillationalways co-fire. Consequence: treatresidual_driftas a redundant diagnostic at default thresholds; monotone drift is caught anyway — drift running into the end viaplateau_too_short, slow creep viawandering.- Random-walk misses are inherent, not fixable. A short random-walk segment is genuinely indistinguishable from stationarity (1/50 misses in the battery). The only real cure is longer runs: random walks grow like √t, stationary noise doesn't.
- The persistence boundary sits at AR(1) ρ ≈ 0.9. A stationary AR(1)
with ρ=0.9 has
r1 ≈ 0.9andVR(8) = (1−ρ⁸)/(8·(1−ρ)) ≈ 0.71— right on the joint gate (r1 > 0.90andVR > 0.75), so sampling noise rejects ~25% of genuinely stationary runs aswandering. Below ρ=0.85 the battery is false-alarm-free (0/50). If a metric is known to be legitimately that persistent, expect and tolerate rejections; do not loosenautocorr_hi/vr_threshwithout re-checking the random-walk catch rate, which trades off directly against this. - "Stabilized" means tail-stationary within a visual tolerance, not convergence to a true value: a drift smaller than the band passes forever, and the certified value is "the level the tail sits at", with no claim it's where the series should be.
- Small-sample edges. N < 20 is unjudgeable; a plateau of 12–19
samples gets no correlation/spectral diagnostics;
min_plateau_frac(0.15) bounds how short a certified plateau may be relative to the run. - Columns are judged independently — no cross-metric corroboration
(e.g. "
dreamsim_to_prevsettled butdreamsim_to_refstill moves" is two separate verdicts, which is by design: the columns answer different questions).
9. Tuning #
CLI exposes the three knobs worth calibrating against runs you have already eye-balled:
--tail-frac(0.40): must stay below1 − (longest acceptable transient fraction). The one real assumption (§8.1).--k-band(5.0): the tolerance width. Larger → latert_stab, more forgiving diagnostics; smaller → noisier t_stab and more attention flags.--min-plateau-frac(0.15): how much of the run must be settled.
The remaining thresholds (drift_bands, shift_bands, spectral_thresh,
autocorr_hi, autocorr_lo, vr_thresh, smooth_window) are deliberately
CLI-invisible: they interact (§8.2, §8.4) and are only meant to be changed
through the Python API (detect_stabilization(y, ...)) with the battery in
hand.
10. File map #
| File | Role |
|---|---|
src/dltb/detect_stabilization.py |
algorithm + dltb-stabilization CLI (detect_stabilization(), analyze_csv() are the API) |
tests/test_detect_stabilization.py |
the seeded stress battery of §7; also pins the pandas-compatible rolling-median convention and the CSV label-alignment behavior |
| this doc | algorithm reference + verified limitations |