Deep Learning Tripping Balls
dltb doc algorithm continuous.md
15 kB
Markdown
at main

The dltb-continuous video loop — carried temporal state around one model #

Reference documentation for dltb-continuous (src/dltb/continuous.py): the loop topology, its variants (modes, reprojection, tails, conditioning strategies), the exact per-frame pipeline, and the outputs it leaves behind. Reprojection internals live in imaging.make_reprojector and are described here to the level needed to reason about results.


1. Purpose and I/O #

dltb-continuous simulates a real-time video enhancement pipeline (TAA, DLSS 2–5) around a diffusion model: each source frame of an input video is processed exactly once, while the model's own previous output can be carried forward as temporal state — the feedback loop that DLSS 5 documents conceptually as

P_n = f(state(P_{n-1}), N_n, motion_vectors_n, artistic_direction)

where P is the processed (carried) frame and N the freshly rendered source frame. The interesting questions are what happens to that loop when the anchor is weakened, dropped, or fed garbage — the failure tails.

  • Input: one video (mp4/mov/mkv/webm/avi, read via imageio/ffmpeg).
  • Output: one run directory untracked/output_<model>/<stem>_<tag>/ (--output-dir overrides) containing the main video processed_<mode>.mp4 at the source fps, optional per-mode tail videos tail_<mode>.mp4, end_state.png, frame_0000_source.png, and every Nth frame as PNG under frames/ (§7).
  • The run() function is also the engine of dltb-klein (via the conditioning hook, §6) and the loop half of dltb-distance-feedback.

2. The two loop topologies (--mode) #

mode recursion what it tests
anchored P_n = f(N_n) independent per-frame passes — the boil test: how much does the same input move under re-rendering noise alone (no state carried)?
stateful (default) P_n = f(blend(R(P_{n-1}), N_n)) a healthy pipeline: carried state, optionally reprojected R, blended with the fresh frame

Stateful details:

  • Blend (--anchor-blend alpha, default 0.3): blend = (1−a)·carried + a·new_frame (linear Image.blend in pixel space). a = 0 is free-running (pure self-iteration), a = 1 re-anchors completely every frame — which is --mode anchored exactly: blend(P, N, 1) = N, the carried state contributes zero, and the flow/warp is skipped at a = 1 too (since 2026-09-15; its result would be blended away), so the two modes are pixel- and per-frame-cost identical. anchored remains a separate mode only for klein dual-ref (§6), where alpha is ignored and it is the only single-reference (boil-test) topology. The blend is the anchor strength of the loop; klein's active range sits far below the img2img models'.
  • First frame: current starts None, so frame 1 is always processed as a fresh source regardless of mode — the loop needs a seed state.
  • Reprojection R is stateful-only and defaults ON (§4); --mode anchored ignores it (and --anchor-blend), which the run tag reflects (§7). In stateful, a = 1 skips the per-frame flow/warp as well (a NOTE is printed when that skip is active, since the header reports the --reproject setting, not the work).

3. The main loop, step by step #

Setup (all fail-fast before any accelerator work):

  1. Resolve geometry (model-native size unless --width/--height, both multiples of 16) and check requirements (gated models need HF_TOKEN; num_inference_steps · strength ≥ 1 for img2img models, else diffusers runs 0 denoise steps).
  2. Validate --anchor-blend ∈ [0, 1].
  3. Build the run directory and tag (§7).
  4. If reprojection is active, build the reprojector (imaging.make_reprojector, §4); otherwise estimate_flow = warp = None.
  5. Obtain the conditioning pair combine, tail_source (§6) — the default is the pixel blend; dltb-klein substitutes dual-reference conditioning.
  6. Resolve device (auto: CUDA → MPS → hard error; --device overrides), load the pipeline once, and open reader/writer (libx264 at source fps).

Per source frame n = 1, 2, … (stopping after --max-frames if given):

  1. new_frame = prepare_frame(frame): convert RGB, center-crop to the target aspect ratio, LANCZOS-resize to width × height. Frame 1 is also saved as frame_0000_source.png; last_source tracks the newest prepared frame (used by tails).
  2. Build the model input:
    • anchored or current is None → source = new_frame;
    • stateful → source = combine(current, new_frame, prev_source) (default: warp the carried state along the flow prev_source → new_frame when reprojection is on, then blend; §4–§6). At a = 1 the warp is skipped — its result would be blended away — and source = N_n exactly, i.e. the anchored case.
  3. prev_source = new_frame (flow is always estimated between consecutive source frames, never between processed ones).
  4. one_pass(source, n):
    • generator: torch.Generator seeded seed when --fixed-seed, else seed + n — fresh noise per pass by default; a fixed seed makes drift purely model bias (DLSS-5-like determinism);
    • run_pass = one img2img call with the shared PassSettings (prompt, steps, strength for uses_strength models, guidance — inert for the step-distilled klein models); SD/SDXL infer output size from the input, FLUX pipelines get width/height explicitly (spec.pass_size);
    • current = result (the carried state), appended to the mp4 writer, saved to frames/frame_{n:04d}.png every --save-every (default 10) frames; progress logged every 10th frame.

After the loop, processed_<mode>.mp4 is complete and the tail phases (§5) branch from the final current.

4. Reprojection (--reproject, default ON in stateful) #

Real temporal pipelines never blend raw history: the engine's motion vectors warp the carried state so stale content lands where the scene has moved to, and only then is it combined with the new frame. Engine motion vectors are unavailable here, so they are estimated with dense Farneback optical flow between consecutive source frames (imaging.make_reprojector), giving the effective main-loop recursion

source_n = (1−a) · warp(P_{n−1}, flow N_{n−1} → N_n) + a · N_n

At a = 1 the warp term has weight zero, so the flow/warp step is skipped entirely (blend(P, N, 1) = N exactly in pixel space) — stateful a=1 makes --mode anchored redundant at the same per-frame cost. The reprojector itself is still constructed (one lazy cv2 import — setup noise, not per-frame work), and a NOTE is printed because the run header reports the --reproject setting, not the work actually done. (At a = 0 the opposite holds: the fresh frame drops out and the warp is the entire input, so it stays load-bearing.)

Mechanics:

  • estimate_flow(prev, next) → forward flow (backward flow too, only when the mask is on), Farneback at fixed parameters (pyr_scale 0.5, 4 levels, winsize 15, 3 iterations, poly 5/1.2) on grayscale frames.
  • warp(state, flow) remaps each target pixel (x, y) from source coordinate (x − dx, y − dy) (backward mapping along the forward flow), bilinear interpolation, BORDER_REPLICATE at the edges.
  • Naive blend (--no-reproject) skips the warp and blends raw history: old- and new-position content superimpose → ghosting. Useful as the pathological A/B leg; under klein dual-ref, no reprojection outright freezes motion (NOTES.md, 2026-09-13 — the regime-mismatch finding that made reprojection load-bearing).

Disocclusion mask (--reproject-mask, default OFF, tag suffix -fbmask; needs --reproject, ignored otherwise): the bare warp smears disoccluded regions (replicate borders + no flow confidence). With the mask, pixels are patched from the fresh frame — exactly the content the disocclusion just revealed — when either trigger fires:

  1. forward–backward circularity: ‖fwd + bwd‖ > fb_tau (1.5 px) at the target pixel (consistent flow should have bwd ≈ −fwd);
  2. the warp samples outside the state frame (would-be replicate smear).

The bad-pixel mask is dilated by a 5 px ellipse to cover the smear's penumbra before patching. Rationale (NOTES.md): klein faithfully preserves warp smears, so warp quality is the artifact ceiling.

5. Tail phases (--tail-frames N --tail-modes freeze,free,black) #

After the last source frame, the carried state P_end is kept in memory, saved as end_state.png, and each requested tail mode branches from a copy of that same state into its own video — the scenarios after the engine stops behaving:

mode source per tail step scenario
freeze blend(P_{n−1}, last_source, a) engine keeps re-submitting the last real frame (static menu with valid zero motion vectors — conceptually an identity warp; tails do not estimate flow)
free P_{n−1} itself anchor dropped entirely — the buffer-echo bug; mathematically free-running at ANY blend, since blend(P, P, a) = P
black blend(P_{n−1}, black, a) renderer submits black frames; a decay driver, NOT free-running

Tail frames run through the same one_pass (same prompt/steps/strength), indexed t = 1…N restarting per mode — so all tail videos of a run share the same per-frame noise sequence (seed + t), making them directly comparable. Each tail writes tail_<mode>.mp4 at the source fps.

6. The conditioning hook #

run(args, make_conditioning=None) separates the loop skeleton (§3) from how carried state and fresh frame become the model input. The optional factory

make_conditioning(args, estimate_flow, warp) -> (combine, tail_source)

combine(carried, new_frame, prev_source) -> source        # main loop
tail_source(mode, current, last_source, black) -> source  # tail phases

receives the reprojector (or Nones) and returns the two source-construction functions. The default (_blend_conditioning) implements the pixel blend of §2–§5: combine warps-then-blends, tail_source switches on the mode.

dltb-klein uses the hook for dual-reference conditioning (--conditioning dual-ref): [state, frame] passed as two clean reference images instead of one pre-blended input, with tails freeze = [P, last_source], free = [P] alone, black = [P, black] (NOTES.md, "klein dual-reference conditioning"). This is why the factory returns a pair of functions rather than one condition(...) callable — tails need their own source construction, and under dual-ref that construction differs structurally, not just parametrically.

7. Outputs and run tags #

Run directory: untracked/output_<model>/<input stem>_<tag>/

processed_<mode>.mp4        # main loop (anchored | stateful)
tail_<mode>.mp4             # one per requested tail mode
end_state.png               # P_end the tails branched from
frame_0000_source.png       # the (prepared) first source frame
frames/frame_NNNN.png       # every --save-every Nth processed frame
frames/<mode> frame_NNNN.png  # idem, tail phases (note the space)

Tag grammar (mode_tag):

  • anchored — no alpha/reprojection components (they don't apply);
  • stateful: stateful-a<alpha> (e.g. stateful-a0.3), or dualref under dual-ref conditioning (no blend component — alpha is not in play); then -norepro if --no-reproject, else -fbmask if the mask is on;
  • + _tails<modes><N> when tails are requested (e.g. _tailsfreeze-free-black60).

Warning (repository invariant): the tag encodes only mode / blend-or- conditioning / tails. Runs differing in prompt, strength, steps, seed, or --ref-order collide silently — use --output-dir subtrees for those axes (sweep.sh does; sweep-klein.sh takes OUT_PREFIX for the same reason). Local (MPS) and pod (CUDA) outputs must not share a tree either.

8. Verified behavior (NOTES.md) #

Findings established on real runs, not guessed:

  • Under klein dual-ref, reprojection is load-bearing: without it motion freezes after the first frames (a regime mismatch — the reference images say "static"), replicated three times across A/B sweeps.
  • Warp quality is the artifact ceiling under klein: smears are preserved faithfully, which is what motivated --reproject-mask.
  • Under dual-ref, the black tail ≈ freeze tail (blending a black reference barely differs from re-submitting the last frame) — black is not a decay driver there, unlike the pixel-blend path.
  • The default freeze-tail preview recipe (stateful, --reproject, --max-frames 30 --tail-frames 10 --tail-modes freeze) is the minimal smoke configuration used throughout NOTES.md.

9. Limitations and caveats #

  • No flow in tail phases: tails construct sources directly (§5); freeze relies on the conceptual identity warp of a zero-motion static scene. A hypothetical tail with nonzero synthetic motion has no path through the current code.
  • Farneback is a coarse estimator (grayscale, fixed params); it stands in for engine motion vectors it cannot match. The mask's two triggers catch gross disocclusions, not subtle flow errors.
  • The blend is linear in pixel space — a regime real pipelines don't use; it was chosen as the simplest composable anchor, with dual-ref as the alternative.
  • Tail noise indices restart at 1 per mode (by design, for comparability — but it means tail frame t uses the same noise as main-loop frame t would with a non-fixed seed).
  • --save-every writes PNGs at idx % N == 0 (frame 10, 20, …), so frame 1 is only in the mp4 (and frame_0000_source.png is the input, not the output).
  • Mode anchored accepts (and ignores) --anchor-blend/--reproject after validation; the run tag honestly reflects only what ran, but the CLI does not warn for that direction. The converse — stateful --anchor-blend 1.0 with reprojection on — prints a NOTE (§4), because the header's reproject=on reflects the setting while the flow/warp is skipped.
  • Everything else about a pass (guidance being inert for klein, size rules, offloading for the big models) is inherited from models.py / imaging.run_pass and documented in README.md.

10. File map #

File Role
src/dltb/continuous.py loop skeleton, modes, tails, tag grammar, run(args, make_conditioning=…) API; dltb-continuous CLI
src/dltb/imaging.py run_pass (one img2img pass), prepare_frame, make_reprojector (Farneback flow + remap + disocclusion mask), make_generator (seed policy)
src/dltb/klein.py dltb-klein: the dual-ref make_conditioning passed into continuous.run
src/dltb/models.py ModelSpec/load_pipeline (device, dtype, offload, MPS slicing), geometry and requirement checks
src/dltb/output.py run_dir layout under untracked/output_<model>/
scripts/sweep.sh video sweep driver exercising modes, blends, and tails
this doc algorithm reference for the loop