# The dltb-continuous video loop — carried temporal state around one model Reference documentation for `dltb-continuous` (`src/dltb/continuous.py`): the loop topology, its variants (modes, reprojection, tails, conditioning strategies), the exact per-frame pipeline, and the outputs it leaves behind. Reprojection internals live in `imaging.make_reprojector` and are described here to the level needed to reason about results. --- ## 1. Purpose and I/O `dltb-continuous` simulates a real-time video enhancement pipeline (TAA, DLSS 2–5) around a diffusion model: each **source** frame of an input video is processed exactly once, while the model's **own previous output** can be carried forward as temporal state — the feedback loop that DLSS 5 documents conceptually as P_n = f(state(P_{n-1}), N_n, motion_vectors_n, artistic_direction) where `P` is the processed (carried) frame and `N` the freshly rendered source frame. The interesting questions are what happens to that loop when the anchor is weakened, dropped, or fed garbage — the failure tails. - Input: one video (mp4/mov/mkv/webm/avi, read via imageio/ffmpeg). - Output: one run directory `untracked/output_/_/` (`--output-dir` overrides) containing the main video `processed_.mp4` at the source fps, optional per-mode tail videos `tail_.mp4`, `end_state.png`, `frame_0000_source.png`, and every Nth frame as PNG under `frames/` (§7). - The `run()` function is also the engine of `dltb-klein` (via the conditioning hook, §6) and the loop half of `dltb-distance-feedback`. ## 2. The two loop topologies (`--mode`) | mode | recursion | what it tests | |---|---|---| | `anchored` | `P_n = f(N_n)` | independent per-frame passes — the **boil test**: how much does the same input move under re-rendering noise alone (no state carried)? | | `stateful` (default) | `P_n = f(blend(R(P_{n-1}), N_n))` | a healthy pipeline: carried state, optionally reprojected `R`, blended with the fresh frame | Stateful details: - **Blend** (`--anchor-blend alpha`, default 0.3): `blend = (1−a)·carried + a·new_frame` (linear `Image.blend` in pixel space). `a = 0` is free-running (pure self-iteration), `a = 1` re-anchors completely every frame — which is `--mode anchored` **exactly**: `blend(P, N, 1) = N`, the carried state contributes zero, and the flow/warp is skipped at `a = 1` too (since 2026-09-15; its result would be blended away), so the two modes are pixel- and per-frame-cost identical. `anchored` remains a separate mode only for klein dual-ref (§6), where alpha is ignored and it is the only single-reference (boil-test) topology. The blend is the anchor strength of the loop; klein's active range sits far below the img2img models'. - **First frame**: `current` starts `None`, so frame 1 is always processed as a fresh source regardless of mode — the loop needs a seed state. - **Reprojection** `R` is stateful-only and defaults ON (§4); `--mode anchored` ignores it (and `--anchor-blend`), which the run tag reflects (§7). In stateful, `a = 1` skips the per-frame flow/warp as well (a NOTE is printed when that skip is active, since the header reports the `--reproject` *setting*, not the work). ## 3. The main loop, step by step Setup (all fail-fast before any accelerator work): 1. Resolve geometry (model-native size unless `--width/--height`, both multiples of 16) and check requirements (gated models need `HF_TOKEN`; `num_inference_steps · strength ≥ 1` for img2img models, else diffusers runs 0 denoise steps). 2. Validate `--anchor-blend ∈ [0, 1]`. 3. Build the run directory and tag (§7). 4. If reprojection is active, build the reprojector (`imaging.make_reprojector`, §4); otherwise `estimate_flow = warp = None`. 5. Obtain the conditioning pair `combine, tail_source` (§6) — the default is the pixel blend; `dltb-klein` substitutes dual-reference conditioning. 6. Resolve device (auto: CUDA → MPS → hard error; `--device` overrides), load the pipeline once, and open reader/writer (libx264 at source fps). Per source frame `n = 1, 2, …` (stopping after `--max-frames` if given): 1. `new_frame = prepare_frame(frame)`: convert RGB, center-crop to the target aspect ratio, LANCZOS-resize to width × height. Frame 1 is also saved as `frame_0000_source.png`; `last_source` tracks the newest prepared frame (used by tails). 2. Build the model input: - `anchored` or `current is None` → `source = new_frame`; - `stateful` → `source = combine(current, new_frame, prev_source)` (default: warp the carried state along the flow `prev_source → new_frame` when reprojection is on, then blend; §4–§6). At `a = 1` the warp is skipped — its result would be blended away — and `source = N_n` exactly, i.e. the anchored case. 3. `prev_source = new_frame` (flow is always estimated between consecutive *source* frames, never between processed ones). 4. `one_pass(source, n)`: - generator: `torch.Generator` seeded `seed` when `--fixed-seed`, else `seed + n` — fresh noise per pass by default; a fixed seed makes drift purely model bias (DLSS-5-like determinism); - `run_pass` = one img2img call with the shared `PassSettings` (prompt, steps, strength for `uses_strength` models, guidance — inert for the step-distilled klein models); SD/SDXL infer output size from the input, FLUX pipelines get width/height explicitly (`spec.pass_size`); - `current = result` (the carried state), appended to the mp4 writer, saved to `frames/frame_{n:04d}.png` every `--save-every` (default 10) frames; progress logged every 10th frame. After the loop, `processed_.mp4` is complete and the tail phases (§5) branch from the final `current`. ## 4. Reprojection (`--reproject`, default ON in stateful) Real temporal pipelines never blend raw history: the engine's motion vectors warp the carried state so stale content lands where the scene has moved to, and only then is it combined with the new frame. Engine motion vectors are unavailable here, so they are **estimated with dense Farneback optical flow between consecutive source frames** (`imaging.make_reprojector`), giving the effective main-loop recursion source_n = (1−a) · warp(P_{n−1}, flow N_{n−1} → N_n) + a · N_n At `a = 1` the warp term has weight zero, so the flow/warp step is skipped entirely (`blend(P, N, 1) = N` exactly in pixel space) — `stateful a=1` makes `--mode anchored` redundant at the same per-frame cost. The reprojector itself is still constructed (one lazy cv2 import — setup noise, not per-frame work), and a NOTE is printed because the run header reports the `--reproject` setting, not the work actually done. (At `a = 0` the opposite holds: the fresh frame drops out and the warp is the *entire* input, so it stays load-bearing.) Mechanics: - `estimate_flow(prev, next)` → forward flow (backward flow too, only when the mask is on), Farneback at fixed parameters (pyr_scale 0.5, 4 levels, winsize 15, 3 iterations, poly 5/1.2) on grayscale frames. - `warp(state, flow)` remaps each target pixel `(x, y)` from source coordinate `(x − dx, y − dy)` (backward mapping along the forward flow), bilinear interpolation, `BORDER_REPLICATE` at the edges. - **Naive blend** (`--no-reproject`) skips the warp and blends raw history: old- and new-position content superimpose → ghosting. Useful as the pathological A/B leg; under klein dual-ref, no reprojection outright *freezes motion* (NOTES.md, 2026-09-13 — the regime-mismatch finding that made reprojection load-bearing). **Disocclusion mask** (`--reproject-mask`, default OFF, tag suffix `-fbmask`; needs `--reproject`, ignored otherwise): the bare warp smears disoccluded regions (replicate borders + no flow confidence). With the mask, pixels are patched from the fresh frame — exactly the content the disocclusion just revealed — when either trigger fires: 1. forward–backward circularity: `‖fwd + bwd‖ > fb_tau` (1.5 px) at the target pixel (consistent flow should have `bwd ≈ −fwd`); 2. the warp samples outside the state frame (would-be replicate smear). The bad-pixel mask is dilated by a 5 px ellipse to cover the smear's penumbra before patching. Rationale (NOTES.md): klein faithfully preserves warp smears, so warp quality is the artifact ceiling. ## 5. Tail phases (`--tail-frames N --tail-modes freeze,free,black`) After the last source frame, the carried state `P_end` is kept in memory, saved as `end_state.png`, and **each requested tail mode branches from a copy of that same state** into its own video — the scenarios after the engine stops behaving: | mode | source per tail step | scenario | |---|---|---| | `freeze` | `blend(P_{n−1}, last_source, a)` | engine keeps re-submitting the last real frame (static menu with valid zero motion vectors — conceptually an identity warp; tails do not estimate flow) | | `free` | `P_{n−1}` itself | anchor dropped entirely — the buffer-echo bug; mathematically free-running at ANY blend, since `blend(P, P, a) = P` | | `black` | `blend(P_{n−1}, black, a)` | renderer submits black frames; a decay driver, NOT free-running | Tail frames run through the same `one_pass` (same prompt/steps/strength), indexed `t = 1…N` restarting per mode — so all tail videos of a run share the same per-frame noise sequence (seed + t), making them directly comparable. Each tail writes `tail_.mp4` at the source fps. ## 6. The conditioning hook `run(args, make_conditioning=None)` separates the *loop skeleton* (§3) from *how carried state and fresh frame become the model input*. The optional factory make_conditioning(args, estimate_flow, warp) -> (combine, tail_source) combine(carried, new_frame, prev_source) -> source # main loop tail_source(mode, current, last_source, black) -> source # tail phases receives the reprojector (or `None`s) and returns the two source-construction functions. The default (`_blend_conditioning`) implements the pixel blend of §2–§5: `combine` warps-then-blends, `tail_source` switches on the mode. `dltb-klein` uses the hook for **dual-reference conditioning** (`--conditioning dual-ref`): `[state, frame]` passed as two *clean reference images* instead of one pre-blended input, with tails `freeze = [P, last_source]`, `free = [P]` alone, `black = [P, black]` (NOTES.md, "klein dual-reference conditioning"). This is why the factory returns a *pair* of functions rather than one `condition(...)` callable — tails need their own source construction, and under dual-ref that construction differs structurally, not just parametrically. ## 7. Outputs and run tags Run directory: `untracked/output_/_/` processed_.mp4 # main loop (anchored | stateful) tail_.mp4 # one per requested tail mode end_state.png # P_end the tails branched from frame_0000_source.png # the (prepared) first source frame frames/frame_NNNN.png # every --save-every Nth processed frame frames/ frame_NNNN.png # idem, tail phases (note the space) Tag grammar (`mode_tag`): - `anchored` — no alpha/reprojection components (they don't apply); - stateful: `stateful-a` (e.g. `stateful-a0.3`), or `dualref` under dual-ref conditioning (no blend component — alpha is not in play); then `-norepro` if `--no-reproject`, else `-fbmask` if the mask is on; - `+ _tails` when tails are requested (e.g. `_tailsfreeze-free-black60`). **Warning (repository invariant):** the tag encodes only mode / blend-or- conditioning / tails. Runs differing in prompt, strength, steps, seed, or `--ref-order` collide silently — use `--output-dir` subtrees for those axes (`sweep.sh` does; `sweep-klein.sh` takes `OUT_PREFIX` for the same reason). Local (MPS) and pod (CUDA) outputs must not share a tree either. ## 8. Verified behavior (NOTES.md) Findings established on real runs, not guessed: - Under klein dual-ref, **reprojection is load-bearing**: without it motion freezes after the first frames (a regime mismatch — the reference images say "static"), replicated three times across A/B sweeps. - Warp quality is the artifact ceiling under klein: smears are preserved faithfully, which is what motivated `--reproject-mask`. - Under dual-ref, the **black tail ≈ freeze tail** (blending a black reference barely differs from re-submitting the last frame) — black is not a decay driver there, unlike the pixel-blend path. - The default freeze-tail preview recipe (stateful, `--reproject`, `--max-frames 30 --tail-frames 10 --tail-modes freeze`) is the minimal smoke configuration used throughout NOTES.md. ## 9. Limitations and caveats - **No flow in tail phases**: tails construct sources directly (§5); `freeze` relies on the conceptual identity warp of a zero-motion static scene. A hypothetical tail with *nonzero* synthetic motion has no path through the current code. - **Farneback is a coarse estimator** (grayscale, fixed params); it stands in for engine motion vectors it cannot match. The mask's two triggers catch gross disocclusions, not subtle flow errors. - The blend is **linear in pixel space** — a regime real pipelines don't use; it was chosen as the simplest composable anchor, with dual-ref as the alternative. - Tail noise indices restart at 1 per mode (by design, for comparability — but it means tail frame *t* uses the same noise as main-loop frame *t* would with a non-fixed seed). - `--save-every` writes PNGs at `idx % N == 0` (frame 10, 20, …), so frame 1 is only in the mp4 (and `frame_0000_source.png` is the *input*, not the output). - Mode `anchored` accepts (and ignores) `--anchor-blend`/`--reproject` after validation; the run tag honestly reflects only what ran, but the CLI does not warn for that direction. The converse — `stateful --anchor-blend 1.0` with reprojection on — prints a NOTE (§4), because the header's `reproject=on` reflects the setting while the flow/warp is skipped. - Everything else about a pass (guidance being inert for klein, size rules, offloading for the big models) is inherited from `models.py` / `imaging.run_pass` and documented in README.md. ## 10. File map | File | Role | |---|---| | `src/dltb/continuous.py` | loop skeleton, modes, tails, tag grammar, `run(args, make_conditioning=…)` API; `dltb-continuous` CLI | | `src/dltb/imaging.py` | `run_pass` (one img2img pass), `prepare_frame`, `make_reprojector` (Farneback flow + remap + disocclusion mask), `make_generator` (seed policy) | | `src/dltb/klein.py` | `dltb-klein`: the dual-ref `make_conditioning` passed into `continuous.run` | | `src/dltb/models.py` | `ModelSpec`/`load_pipeline` (device, dtype, offload, MPS slicing), geometry and requirement checks | | `src/dltb/output.py` | `run_dir` layout under `untracked/output_/` | | `scripts/sweep.sh` | video sweep driver exercising modes, blends, and tails | | this doc | algorithm reference for the loop |