# Stabilization detection in dltb — reverse convergence detection Reference documentation for `dltb-stabilization` (`src/dltb/detect_stabilization.py`): the algorithm, its verified behavior, and its limits. The claims marked *verified* below are pinned by the seeded stress battery in `tests/test_detect_stabilization.py` (plain Python, CPU-only; a real-run regression case skips itself when that run's CSV is not present). Status at time of writing (2026-09-14): the tool is a faithful, pandas-free port of a reference implementation; its numerics reproduce that implementation's documented results exactly, and the battery below is the reconstructed validation of its robustness claims. --- ## 1. Purpose and I/O The analysis tools leave behind metric CSVs — `distance_metrics.csv` (`dltb-distance`, `dltb-distance-loop`, `dltb-distance-feedback`) and `drift_metrics.csv` (`analyze_drift.py`) — one row per frame, one column per metric. `dltb-stabilization` reads such a CSV and answers, **per column**: did the series stabilize, at which frame, and at what value — or why it deserves human attention instead. - First column = frame label; every other column is analysed independently; non-numeric / non-finite cells are dropped per column, with labels kept aligned to the surviving values (so `t_stab` indexes and frame labels cannot drift apart). - Output: a verdict table on stdout, `stabilization_report.csv` next to the input (`-o` overrides), optional `--plot` diagnostic PNG (one panel per column). Exit code is always 0 — sweeps grep the table for `needs_attention` rather than checking status. - CPU-only: numpy + stdlib CSV; matplotlib is imported only when `--plot` is given. ## 2. The conceptual frame The classic mistake is scanning forward and asking "has it settled yet?" at each sample — that forces a burn-in guess before you know the answer. This tool flips the question: **assume the end of the run is stationary, characterize that regime, then ask how far back it extends.** One backward rule then handles every transient shape with no per-shape logic: rapid rise, spike-then-fall, and two-plateaus-with-a-jump are all "walk back to the last sample outside the final regime's band". The single real assumption is that the *tail* actually is stationary (§8.1). ## 3. The pipeline, per column Defaults in parentheses. 1. **Sample gate.** Fewer than 20 valid samples → `too_few_samples`, verdict `needs_attention`, no `t_stab`. 2. **Smoothing.** Centered rolling median, window `w = max(3, round(0.01·N))` forced odd (≈1% of the run). A rolling *median* annihilates isolated spikes — any burst of ≤ ⌊w/2⌋ consecutive outlying samples cannot move the smoothed series at all; wider bursts survive and (correctly) push `t_stab` past them. 3. **Tail reference.** Over the last `max(10, round(tail_frac·N))` samples (0.40): median `med` and robust scale `σ = 1.4826·MAD` (the 1.4826 makes MAD consistent with standard deviation for Gaussian noise). A degenerate tail (`σ ≤ 10⁻¹²·(1+|med|)`) → `constant_series`, stabilized at the median from sample 0, `stderr = 0`. 4. **Band.** `[med − k_band·σ, med + k_band·σ]` (k = 5). This band is the **practical tolerance** — what "settled" means on the chart (§6). 5. **Backward scan.** `t_stab = (index of last smoothed sample outside the band) + 1`, or 0 if none; clamped into `[0, N−1]`. The plateau is the **raw** (unsmoothed) samples `y[t_stab:]`. If even the final sample is out of band (still moving at the end), the clamp leaves a ~1-sample plateau → `plateau_too_short`: "run ended before settling". 6. **Diagnostics** (§4), then the verdict: `stabilized` ⇔ no flags beyond the informational `flat_from_start` (`constant_series` short-circuits to stabilized in step 3). ## 4. The plateau diagnostics Any failure flags the column `needs_attention`. Thresholds for the drift and shift tests are in **band units** (§6); the correlation and spectral tests are scale-free. | Flag | Test | Catches | |---|---|---| | `plateau_too_short` | `P < max(10, 0.15·N)` | run ended before settling (still rising/drifting at the end) | | `residual_drift` | Theil–Sen slope of the plateau's second half: `|slope|·len(h2) > 1·band` | ongoing trend inside the "settled" region | | `level_shift` | `|median(h1) − median(h2)| > 0.5·band` | slow drift / step across the plateau | | `wandering` | lag-1 autocorr `r1 > 0.95`, or `r1 > 0.90` and variance ratio `> 0.75` | random-walk-like meandering that never converges | | `alternating_oscillation` | `r1 < −0.5` | period-2 limit cycle | | `periodic_oscillation` | largest single-frequency share of variance `> 0.5` | limit cycle at any period | | `constant_series` | zero tail variance | informational; always stabilized | `r1` is the lag-1 autocorrelation of the centered plateau. The variance ratio is `VR(q) = Var(z_{t+q} − z_t) / (q·Var(z_{t+1} − z_t))` with `q = min(8, max(4, P/4))`: ≈1 for a random walk (differences scale linearly), ≪1 for stationary series (mean reversion anti-correlates differences). The spectral share is the max-bin fraction of the detrended, Hann-windowed periodogram. The statistical tests require `P ≥ 12`; below that only `plateau_too_short` can fire. Theil–Sen (median of pairwise slopes, on a ≤200-point subsample) is used for drift because it is robust both to noise and to oscillation — the property that keeps `residual_drift` alive at all (§8.2). ## 5. Reported value and uncertainty `stabilized_value` = median of the plateau. `stderr` = robust plateau sigma divided by `√n_eff` with `n_eff = P·(1−ρ)/(1+ρ)`, `ρ = clip(r1, −0.9, 0.9)`: strongly autocorrelated samples carry less independent information, and the uncertainty is widened accordingly. For a constant series both are the median and 0. ## 6. Tolerance philosophy The band is deliberately the *practical* tolerance, not a statistical one. The failure this avoids: a very quiet metric column (tiny σ) makes any σ-scaled significance test fire on wiggles invisible on the chart — statistical significance ≠ practical significance. Expressing the drift and shift thresholds in band units (multiples of 5σ, i.e. of what you'd accept as "the same level") makes them fire only on departures a human would care about. The flip side: the correlation and spectral tests are scale-free by nature, so no band Units can protect a highly persistent series from them — that boundary is §8.4. ## 7. Verified behavior (stress battery) All numbers from `tests/test_detect_stabilization.py` (seeded → exact). | Shape | Result | |---|---| | exp rise → plateau (N=600, τ=40) | stabilized, `t_stab`=125 (theory: 40·ln 20 ≈ 120) | | baseline + 2-sample spike at t=50 | stabilized, `t_stab`=52 — spike located, not fooled by | | two plateaus, jump at t=200 of 500 | stabilized, `t_stab`=200, value = **second** level | | noisy linear ramp, N=400 (ends t=100) | stabilized, `t_stab`=89, value ≈ ramp top | | constant series | stabilized + `constant_series` | | monotone rise, never settles | `needs_attention` (`level_shift`, `wandering`) | | flat, then ramp into the last frame | `needs_attention` (`plateau_too_short`) | | slow in-band creep under oscillation | `needs_attention` (`residual_drift` + co-fires, §8.2) | | sine limit cycle | `needs_attention` (`wandering`, `periodic_oscillation`) | | alternating (± every sample) | `needs_attention` (`alternating_oscillation`, `periodic_oscillation`) | | N=15 | `too_few_samples` | | random walks ×50 | **49/50** flagged | | stationary AR(1) ×50 | **0/50** false alarms at ρ=0.8 and ρ=0.85 | | stationary AR(1), ρ=0.9 ×50 | **12/50** false alarms (§8.4) | | real sd-turbo feedback run (201 samples, 3 DreamSim series) | all stabilized, `t_stab` = 6/13/5, values matching the values established when the tool was developed | ## 8. Limitations Each of these was established empirically, not guessed. 1. **The tail assumption is load-bearing.** `tail_frac` (0.40) must be shorter than the slowest expected transient. If the run is *still in its transient* at the end, the band is calibrated on moving data: the verdict becomes `plateau_too_short` (the good outcome — the designed behavior), but a transient that happens to move less than the widened band can pass as stabilized with an inflated `stderr`. A crop of `plateau_too_short` verdicts means the runs are too short — extend the runs rather than loosen thresholds. 2. **`residual_drift` cannot fire for monotone drift.** For any monotone creep the tail MAD inflates with the drift itself and always outpaces the slope test: for a linear creep at rate `r`, the tail (length `L = 0.4·N`) has `MAD ≈ r·L/4`, so `band ≈ 5·1.4826·r·L/4 ≈ 1.85·r·L`, while the drift accumulated over the plateau's second half is at most `r·N/2 < r·L·1.25` — the ratio never reaches the threshold of 1 band. Verified by a 4000-shape random search: the only hits (2) were drift mixed with *oscillation*, where Theil–Sen's robustness to oscillation lets the slope test see what the MAD band absorbs — and there `level_shift`/`wandering`/`periodic_oscillation` always co-fire. **Consequence:** treat `residual_drift` as a redundant diagnostic at default thresholds; monotone drift is caught anyway — drift running into the end via `plateau_too_short`, slow creep via `wandering`. 3. **Random-walk misses are inherent, not fixable.** A short random-walk segment is genuinely indistinguishable from stationarity (1/50 misses in the battery). The only real cure is longer runs: random walks grow like √t, stationary noise doesn't. 4. **The persistence boundary sits at AR(1) ρ ≈ 0.9.** A stationary AR(1) with ρ=0.9 has `r1 ≈ 0.9` and `VR(8) = (1−ρ⁸)/(8·(1−ρ)) ≈ 0.71` — right on the joint gate (`r1 > 0.90` and `VR > 0.75`), so sampling noise rejects ~25% of genuinely stationary runs as `wandering`. Below ρ=0.85 the battery is false-alarm-free (0/50). If a metric is *known* to be legitimately that persistent, expect and tolerate rejections; do not loosen `autocorr_hi`/`vr_thresh` without re-checking the random-walk catch rate, which trades off directly against this. 5. **"Stabilized" means tail-stationary within a visual tolerance**, not convergence to a true value: a drift smaller than the band passes forever, and the certified value is "the level the tail sits at", with no claim it's where the series *should* be. 6. **Small-sample edges.** N < 20 is unjudgeable; a plateau of 12–19 samples gets no correlation/spectral diagnostics; `min_plateau_frac` (0.15) bounds how short a *certified* plateau may be relative to the run. 7. **Columns are judged independently** — no cross-metric corroboration (e.g. "`dreamsim_to_prev` settled but `dreamsim_to_ref` still moves" is two separate verdicts, which is by design: the columns answer different questions). ## 9. Tuning CLI exposes the three knobs worth calibrating against runs you have already eye-balled: - `--tail-frac` (0.40): must stay below `1 − (longest acceptable transient fraction)`. The one real assumption (§8.1). - `--k-band` (5.0): the tolerance width. Larger → later `t_stab`, more forgiving diagnostics; smaller → noisier t_stab and more attention flags. - `--min-plateau-frac` (0.15): how much of the run must be settled. The remaining thresholds (`drift_bands`, `shift_bands`, `spectral_thresh`, `autocorr_hi`, `autocorr_lo`, `vr_thresh`, `smooth_window`) are deliberately CLI-invisible: they interact (§8.2, §8.4) and are only meant to be changed through the Python API (`detect_stabilization(y, ...)`) with the battery in hand. ## 10. File map | File | Role | |---|---| | `src/dltb/detect_stabilization.py` | algorithm + `dltb-stabilization` CLI (`detect_stabilization()`, `analyze_csv()` are the API) | | `tests/test_detect_stabilization.py` | the seeded stress battery of §7; also pins the pandas-compatible rolling-median convention and the CSV label-alignment behavior | | this doc | algorithm reference + verified limitations |