Deep Learning Tripping Balls
dltb doc algorithm stabilization.md
12 kB
Markdown
at main

Stabilization detection in dltb — reverse convergence detection #

Reference documentation for dltb-stabilization (src/dltb/detect_stabilization.py): the algorithm, its verified behavior, and its limits. The claims marked verified below are pinned by the seeded stress battery in tests/test_detect_stabilization.py (plain Python, CPU-only; a real-run regression case skips itself when that run's CSV is not present).

Status at time of writing (2026-09-14): the tool is a faithful, pandas-free port of a reference implementation; its numerics reproduce that implementation's documented results exactly, and the battery below is the reconstructed validation of its robustness claims.


1. Purpose and I/O #

The analysis tools leave behind metric CSVs — distance_metrics.csv (dltb-distance, dltb-distance-loop, dltb-distance-feedback) and drift_metrics.csv (analyze_drift.py) — one row per frame, one column per metric. dltb-stabilization reads such a CSV and answers, per column: did the series stabilize, at which frame, and at what value — or why it deserves human attention instead.

  • First column = frame label; every other column is analysed independently; non-numeric / non-finite cells are dropped per column, with labels kept aligned to the surviving values (so t_stab indexes and frame labels cannot drift apart).
  • Output: a verdict table on stdout, stabilization_report.csv next to the input (-o overrides), optional --plot diagnostic PNG (one panel per column). Exit code is always 0 — sweeps grep the table for needs_attention rather than checking status.
  • CPU-only: numpy + stdlib CSV; matplotlib is imported only when --plot is given.

2. The conceptual frame #

The classic mistake is scanning forward and asking "has it settled yet?" at each sample — that forces a burn-in guess before you know the answer. This tool flips the question: assume the end of the run is stationary, characterize that regime, then ask how far back it extends.

One backward rule then handles every transient shape with no per-shape logic: rapid rise, spike-then-fall, and two-plateaus-with-a-jump are all "walk back to the last sample outside the final regime's band". The single real assumption is that the tail actually is stationary (§8.1).

3. The pipeline, per column #

Defaults in parentheses.

  1. Sample gate. Fewer than 20 valid samples → too_few_samples, verdict needs_attention, no t_stab.
  2. Smoothing. Centered rolling median, window w = max(3, round(0.01·N)) forced odd (≈1% of the run). A rolling median annihilates isolated spikes — any burst of ≤ ⌊w/2⌋ consecutive outlying samples cannot move the smoothed series at all; wider bursts survive and (correctly) push t_stab past them.
  3. Tail reference. Over the last max(10, round(tail_frac·N)) samples (0.40): median med and robust scale σ = 1.4826·MAD (the 1.4826 makes MAD consistent with standard deviation for Gaussian noise). A degenerate tail (σ ≤ 10⁻¹²·(1+|med|)) → constant_series, stabilized at the median from sample 0, stderr = 0.
  4. Band. [med − k_band·σ, med + k_band·σ] (k = 5). This band is the practical tolerance — what "settled" means on the chart (§6).
  5. Backward scan. t_stab = (index of last smoothed sample outside the band) + 1, or 0 if none; clamped into [0, N−1]. The plateau is the raw (unsmoothed) samples y[t_stab:]. If even the final sample is out of band (still moving at the end), the clamp leaves a ~1-sample plateau → plateau_too_short: "run ended before settling".
  6. Diagnostics (§4), then the verdict: stabilized ⇔ no flags beyond the informational flat_from_start (constant_series short-circuits to stabilized in step 3).

4. The plateau diagnostics #

Any failure flags the column needs_attention. Thresholds for the drift and shift tests are in band units (§6); the correlation and spectral tests are scale-free.

Flag Test Catches
plateau_too_short P < max(10, 0.15·N) run ended before settling (still rising/drifting at the end)
residual_drift Theil–Sen slope of the plateau's second half: ` slope
level_shift ` median(h1) − median(h2)
wandering lag-1 autocorr r1 > 0.95, or r1 > 0.90 and variance ratio > 0.75 random-walk-like meandering that never converges
alternating_oscillation r1 < −0.5 period-2 limit cycle
periodic_oscillation largest single-frequency share of variance > 0.5 limit cycle at any period
constant_series zero tail variance informational; always stabilized

r1 is the lag-1 autocorrelation of the centered plateau. The variance ratio is VR(q) = Var(z_{t+q} − z_t) / (q·Var(z_{t+1} − z_t)) with q = min(8, max(4, P/4)): ≈1 for a random walk (differences scale linearly), ≪1 for stationary series (mean reversion anti-correlates differences). The spectral share is the max-bin fraction of the detrended, Hann-windowed periodogram. The statistical tests require P ≥ 12; below that only plateau_too_short can fire.

Theil–Sen (median of pairwise slopes, on a ≤200-point subsample) is used for drift because it is robust both to noise and to oscillation — the property that keeps residual_drift alive at all (§8.2).

5. Reported value and uncertainty #

stabilized_value = median of the plateau. stderr = robust plateau sigma divided by √n_eff with n_eff = P·(1−ρ)/(1+ρ), ρ = clip(r1, −0.9, 0.9): strongly autocorrelated samples carry less independent information, and the uncertainty is widened accordingly. For a constant series both are the median and 0.

6. Tolerance philosophy #

The band is deliberately the practical tolerance, not a statistical one. The failure this avoids: a very quiet metric column (tiny σ) makes any σ-scaled significance test fire on wiggles invisible on the chart — statistical significance ≠ practical significance. Expressing the drift and shift thresholds in band units (multiples of 5σ, i.e. of what you'd accept as "the same level") makes them fire only on departures a human would care about. The flip side: the correlation and spectral tests are scale-free by nature, so no band Units can protect a highly persistent series from them — that boundary is §8.4.

7. Verified behavior (stress battery) #

All numbers from tests/test_detect_stabilization.py (seeded → exact).

Shape Result
exp rise → plateau (N=600, τ=40) stabilized, t_stab=125 (theory: 40·ln 20 ≈ 120)
baseline + 2-sample spike at t=50 stabilized, t_stab=52 — spike located, not fooled by
two plateaus, jump at t=200 of 500 stabilized, t_stab=200, value = second level
noisy linear ramp, N=400 (ends t=100) stabilized, t_stab=89, value ≈ ramp top
constant series stabilized + constant_series
monotone rise, never settles needs_attention (level_shift, wandering)
flat, then ramp into the last frame needs_attention (plateau_too_short)
slow in-band creep under oscillation needs_attention (residual_drift + co-fires, §8.2)
sine limit cycle needs_attention (wandering, periodic_oscillation)
alternating (± every sample) needs_attention (alternating_oscillation, periodic_oscillation)
N=15 too_few_samples
random walks ×50 49/50 flagged
stationary AR(1) ×50 0/50 false alarms at ρ=0.8 and ρ=0.85
stationary AR(1), ρ=0.9 ×50 12/50 false alarms (§8.4)
real sd-turbo feedback run (201 samples, 3 DreamSim series) all stabilized, t_stab = 6/13/5, values matching the values established when the tool was developed

8. Limitations #

Each of these was established empirically, not guessed.

  1. The tail assumption is load-bearing. tail_frac (0.40) must be shorter than the slowest expected transient. If the run is still in its transient at the end, the band is calibrated on moving data: the verdict becomes plateau_too_short (the good outcome — the designed behavior), but a transient that happens to move less than the widened band can pass as stabilized with an inflated stderr. A crop of plateau_too_short verdicts means the runs are too short — extend the runs rather than loosen thresholds.
  2. residual_drift cannot fire for monotone drift. For any monotone creep the tail MAD inflates with the drift itself and always outpaces the slope test: for a linear creep at rate r, the tail (length L = 0.4·N) has MAD ≈ r·L/4, so band ≈ 5·1.4826·r·L/4 ≈ 1.85·r·L, while the drift accumulated over the plateau's second half is at most r·N/2 < r·L·1.25 — the ratio never reaches the threshold of 1 band. Verified by a 4000-shape random search: the only hits (2) were drift mixed with oscillation, where Theil–Sen's robustness to oscillation lets the slope test see what the MAD band absorbs — and there level_shift/wandering/periodic_oscillation always co-fire. Consequence: treat residual_drift as a redundant diagnostic at default thresholds; monotone drift is caught anyway — drift running into the end via plateau_too_short, slow creep via wandering.
  3. Random-walk misses are inherent, not fixable. A short random-walk segment is genuinely indistinguishable from stationarity (1/50 misses in the battery). The only real cure is longer runs: random walks grow like √t, stationary noise doesn't.
  4. The persistence boundary sits at AR(1) ρ ≈ 0.9. A stationary AR(1) with ρ=0.9 has r1 ≈ 0.9 and VR(8) = (1−ρ⁸)/(8·(1−ρ)) ≈ 0.71 — right on the joint gate (r1 > 0.90 and VR > 0.75), so sampling noise rejects ~25% of genuinely stationary runs as wandering. Below ρ=0.85 the battery is false-alarm-free (0/50). If a metric is known to be legitimately that persistent, expect and tolerate rejections; do not loosen autocorr_hi/vr_thresh without re-checking the random-walk catch rate, which trades off directly against this.
  5. "Stabilized" means tail-stationary within a visual tolerance, not convergence to a true value: a drift smaller than the band passes forever, and the certified value is "the level the tail sits at", with no claim it's where the series should be.
  6. Small-sample edges. N < 20 is unjudgeable; a plateau of 12–19 samples gets no correlation/spectral diagnostics; min_plateau_frac (0.15) bounds how short a certified plateau may be relative to the run.
  7. Columns are judged independently — no cross-metric corroboration (e.g. "dreamsim_to_prev settled but dreamsim_to_ref still moves" is two separate verdicts, which is by design: the columns answer different questions).

9. Tuning #

CLI exposes the three knobs worth calibrating against runs you have already eye-balled:

  • --tail-frac (0.40): must stay below 1 − (longest acceptable transient fraction). The one real assumption (§8.1).
  • --k-band (5.0): the tolerance width. Larger → later t_stab, more forgiving diagnostics; smaller → noisier t_stab and more attention flags.
  • --min-plateau-frac (0.15): how much of the run must be settled.

The remaining thresholds (drift_bands, shift_bands, spectral_thresh, autocorr_hi, autocorr_lo, vr_thresh, smooth_window) are deliberately CLI-invisible: they interact (§8.2, §8.4) and are only meant to be changed through the Python API (detect_stabilization(y, ...)) with the battery in hand.

10. File map #

File Role
src/dltb/detect_stabilization.py algorithm + dltb-stabilization CLI (detect_stabilization(), analyze_csv() are the API)
tests/test_detect_stabilization.py the seeded stress battery of §7; also pins the pandas-compatible rolling-median convention and the CSV label-alignment behavior
this doc algorithm reference + verified limitations