The dltb-continuous video loop — carried temporal state around one model #
Reference documentation for dltb-continuous (src/dltb/continuous.py): the
loop topology, its variants (modes, reprojection, tails, conditioning
strategies), the exact per-frame pipeline, and the outputs it leaves behind.
Reprojection internals live in imaging.make_reprojector and are described
here to the level needed to reason about results.
1. Purpose and I/O #
dltb-continuous simulates a real-time video enhancement pipeline (TAA,
DLSS 2–5) around a diffusion model: each source frame of an input video
is processed exactly once, while the model's own previous output can be
carried forward as temporal state — the feedback loop that DLSS 5 documents
conceptually as
P_n = f(state(P_{n-1}), N_n, motion_vectors_n, artistic_direction)
where P is the processed (carried) frame and N the freshly rendered
source frame. The interesting questions are what happens to that loop when
the anchor is weakened, dropped, or fed garbage — the failure tails.
- Input: one video (mp4/mov/mkv/webm/avi, read via imageio/ffmpeg).
- Output: one run directory
untracked/output_<model>/<stem>_<tag>/(--output-diroverrides) containing the main videoprocessed_<mode>.mp4at the source fps, optional per-mode tail videostail_<mode>.mp4,end_state.png,frame_0000_source.png, and every Nth frame as PNG underframes/(§7). - The
run()function is also the engine ofdltb-klein(via the conditioning hook, §6) and the loop half ofdltb-distance-feedback.
2. The two loop topologies (--mode) #
| mode | recursion | what it tests |
|---|---|---|
anchored |
P_n = f(N_n) |
independent per-frame passes — the boil test: how much does the same input move under re-rendering noise alone (no state carried)? |
stateful (default) |
P_n = f(blend(R(P_{n-1}), N_n)) |
a healthy pipeline: carried state, optionally reprojected R, blended with the fresh frame |
Stateful details:
- Blend (
--anchor-blend alpha, default 0.3):blend = (1−a)·carried + a·new_frame(linearImage.blendin pixel space).a = 0is free-running (pure self-iteration),a = 1re-anchors completely every frame — which is--mode anchoredexactly:blend(P, N, 1) = N, the carried state contributes zero, and the flow/warp is skipped ata = 1too (since 2026-09-15; its result would be blended away), so the two modes are pixel- and per-frame-cost identical.anchoredremains a separate mode only for klein dual-ref (§6), where alpha is ignored and it is the only single-reference (boil-test) topology. The blend is the anchor strength of the loop; klein's active range sits far below the img2img models'. - First frame:
currentstartsNone, so frame 1 is always processed as a fresh source regardless of mode — the loop needs a seed state. - Reprojection
Ris stateful-only and defaults ON (§4);--mode anchoredignores it (and--anchor-blend), which the run tag reflects (§7). In stateful,a = 1skips the per-frame flow/warp as well (a NOTE is printed when that skip is active, since the header reports the--reprojectsetting, not the work).
3. The main loop, step by step #
Setup (all fail-fast before any accelerator work):
- Resolve geometry (model-native size unless
--width/--height, both multiples of 16) and check requirements (gated models needHF_TOKEN;num_inference_steps · strength ≥ 1for img2img models, else diffusers runs 0 denoise steps). - Validate
--anchor-blend ∈ [0, 1]. - Build the run directory and tag (§7).
- If reprojection is active, build the reprojector
(
imaging.make_reprojector, §4); otherwiseestimate_flow = warp = None. - Obtain the conditioning pair
combine, tail_source(§6) — the default is the pixel blend;dltb-kleinsubstitutes dual-reference conditioning. - Resolve device (auto: CUDA → MPS → hard error;
--deviceoverrides), load the pipeline once, and open reader/writer (libx264 at source fps).
Per source frame n = 1, 2, … (stopping after --max-frames if given):
new_frame = prepare_frame(frame): convert RGB, center-crop to the target aspect ratio, LANCZOS-resize to width × height. Frame 1 is also saved asframe_0000_source.png;last_sourcetracks the newest prepared frame (used by tails).- Build the model input:
anchoredorcurrent is None→source = new_frame;stateful→source = combine(current, new_frame, prev_source)(default: warp the carried state along the flowprev_source → new_framewhen reprojection is on, then blend; §4–§6). Ata = 1the warp is skipped — its result would be blended away — andsource = N_nexactly, i.e. the anchored case.
prev_source = new_frame(flow is always estimated between consecutive source frames, never between processed ones).one_pass(source, n):- generator:
torch.Generatorseededseedwhen--fixed-seed, elseseed + n— fresh noise per pass by default; a fixed seed makes drift purely model bias (DLSS-5-like determinism); run_pass= one img2img call with the sharedPassSettings(prompt, steps, strength foruses_strengthmodels, guidance — inert for the step-distilled klein models); SD/SDXL infer output size from the input, FLUX pipelines get width/height explicitly (spec.pass_size);current = result(the carried state), appended to the mp4 writer, saved toframes/frame_{n:04d}.pngevery--save-every(default 10) frames; progress logged every 10th frame.
- generator:
After the loop, processed_<mode>.mp4 is complete and the tail phases (§5)
branch from the final current.
4. Reprojection (--reproject, default ON in stateful) #
Real temporal pipelines never blend raw history: the engine's motion vectors
warp the carried state so stale content lands where the scene has moved to,
and only then is it combined with the new frame. Engine motion vectors are
unavailable here, so they are estimated with dense Farneback optical flow
between consecutive source frames (imaging.make_reprojector), giving the
effective main-loop recursion
source_n = (1−a) · warp(P_{n−1}, flow N_{n−1} → N_n) + a · N_n
At a = 1 the warp term has weight zero, so the flow/warp step is skipped
entirely (blend(P, N, 1) = N exactly in pixel space) — stateful a=1
makes --mode anchored redundant at the same per-frame cost. The
reprojector itself is still constructed (one lazy cv2 import — setup noise,
not per-frame work), and a NOTE is printed because the run header reports
the --reproject setting, not the work actually done. (At a = 0 the
opposite holds: the fresh frame drops out and the warp is the entire
input, so it stays load-bearing.)
Mechanics:
estimate_flow(prev, next)→ forward flow (backward flow too, only when the mask is on), Farneback at fixed parameters (pyr_scale 0.5, 4 levels, winsize 15, 3 iterations, poly 5/1.2) on grayscale frames.warp(state, flow)remaps each target pixel(x, y)from source coordinate(x − dx, y − dy)(backward mapping along the forward flow), bilinear interpolation,BORDER_REPLICATEat the edges.- Naive blend (
--no-reproject) skips the warp and blends raw history: old- and new-position content superimpose → ghosting. Useful as the pathological A/B leg; under klein dual-ref, no reprojection outright freezes motion (NOTES.md, 2026-09-13 — the regime-mismatch finding that made reprojection load-bearing).
Disocclusion mask (--reproject-mask, default OFF, tag suffix
-fbmask; needs --reproject, ignored otherwise): the bare warp smears
disoccluded regions (replicate borders + no flow confidence). With the mask,
pixels are patched from the fresh frame — exactly the content the
disocclusion just revealed — when either trigger fires:
- forward–backward circularity:
‖fwd + bwd‖ > fb_tau(1.5 px) at the target pixel (consistent flow should havebwd ≈ −fwd); - the warp samples outside the state frame (would-be replicate smear).
The bad-pixel mask is dilated by a 5 px ellipse to cover the smear's penumbra before patching. Rationale (NOTES.md): klein faithfully preserves warp smears, so warp quality is the artifact ceiling.
5. Tail phases (--tail-frames N --tail-modes freeze,free,black) #
After the last source frame, the carried state P_end is kept in memory,
saved as end_state.png, and each requested tail mode branches from a
copy of that same state into its own video — the scenarios after the
engine stops behaving:
| mode | source per tail step | scenario |
|---|---|---|
freeze |
blend(P_{n−1}, last_source, a) |
engine keeps re-submitting the last real frame (static menu with valid zero motion vectors — conceptually an identity warp; tails do not estimate flow) |
free |
P_{n−1} itself |
anchor dropped entirely — the buffer-echo bug; mathematically free-running at ANY blend, since blend(P, P, a) = P |
black |
blend(P_{n−1}, black, a) |
renderer submits black frames; a decay driver, NOT free-running |
Tail frames run through the same one_pass (same prompt/steps/strength),
indexed t = 1…N restarting per mode — so all tail videos of a run share
the same per-frame noise sequence (seed + t), making them directly
comparable. Each tail writes tail_<mode>.mp4 at the source fps.
6. The conditioning hook #
run(args, make_conditioning=None) separates the loop skeleton (§3) from
how carried state and fresh frame become the model input. The optional
factory
make_conditioning(args, estimate_flow, warp) -> (combine, tail_source)
combine(carried, new_frame, prev_source) -> source # main loop
tail_source(mode, current, last_source, black) -> source # tail phases
receives the reprojector (or Nones) and returns the two source-construction
functions. The default (_blend_conditioning) implements the pixel blend of
§2–§5: combine warps-then-blends, tail_source switches on the mode.
dltb-klein uses the hook for dual-reference conditioning
(--conditioning dual-ref): [state, frame] passed as two clean reference
images instead of one pre-blended input, with tails freeze = [P, last_source], free = [P] alone, black = [P, black] (NOTES.md, "klein
dual-reference conditioning"). This is why the factory returns a pair of
functions rather than one condition(...) callable — tails need their own
source construction, and under dual-ref that construction differs
structurally, not just parametrically.
7. Outputs and run tags #
Run directory: untracked/output_<model>/<input stem>_<tag>/
processed_<mode>.mp4 # main loop (anchored | stateful)
tail_<mode>.mp4 # one per requested tail mode
end_state.png # P_end the tails branched from
frame_0000_source.png # the (prepared) first source frame
frames/frame_NNNN.png # every --save-every Nth processed frame
frames/<mode> frame_NNNN.png # idem, tail phases (note the space)
Tag grammar (mode_tag):
anchored— no alpha/reprojection components (they don't apply);- stateful:
stateful-a<alpha>(e.g.stateful-a0.3), ordualrefunder dual-ref conditioning (no blend component — alpha is not in play); then-noreproif--no-reproject, else-fbmaskif the mask is on; + _tails<modes><N>when tails are requested (e.g._tailsfreeze-free-black60).
Warning (repository invariant): the tag encodes only mode / blend-or-
conditioning / tails. Runs differing in prompt, strength, steps, seed, or
--ref-order collide silently — use --output-dir subtrees for those axes
(sweep.sh does; sweep-klein.sh takes OUT_PREFIX for the same reason).
Local (MPS) and pod (CUDA) outputs must not share a tree either.
8. Verified behavior (NOTES.md) #
Findings established on real runs, not guessed:
- Under klein dual-ref, reprojection is load-bearing: without it motion freezes after the first frames (a regime mismatch — the reference images say "static"), replicated three times across A/B sweeps.
- Warp quality is the artifact ceiling under klein: smears are preserved
faithfully, which is what motivated
--reproject-mask. - Under dual-ref, the black tail ≈ freeze tail (blending a black reference barely differs from re-submitting the last frame) — black is not a decay driver there, unlike the pixel-blend path.
- The default freeze-tail preview recipe (stateful,
--reproject,--max-frames 30 --tail-frames 10 --tail-modes freeze) is the minimal smoke configuration used throughout NOTES.md.
9. Limitations and caveats #
- No flow in tail phases: tails construct sources directly (§5);
freezerelies on the conceptual identity warp of a zero-motion static scene. A hypothetical tail with nonzero synthetic motion has no path through the current code. - Farneback is a coarse estimator (grayscale, fixed params); it stands in for engine motion vectors it cannot match. The mask's two triggers catch gross disocclusions, not subtle flow errors.
- The blend is linear in pixel space — a regime real pipelines don't use; it was chosen as the simplest composable anchor, with dual-ref as the alternative.
- Tail noise indices restart at 1 per mode (by design, for comparability — but it means tail frame t uses the same noise as main-loop frame t would with a non-fixed seed).
--save-everywrites PNGs atidx % N == 0(frame 10, 20, …), so frame 1 is only in the mp4 (andframe_0000_source.pngis the input, not the output).- Mode
anchoredaccepts (and ignores)--anchor-blend/--reprojectafter validation; the run tag honestly reflects only what ran, but the CLI does not warn for that direction. The converse —stateful --anchor-blend 1.0with reprojection on — prints a NOTE (§4), because the header'sreproject=onreflects the setting while the flow/warp is skipped. - Everything else about a pass (guidance being inert for klein, size rules,
offloading for the big models) is inherited from
models.py/imaging.run_passand documented in README.md.
10. File map #
| File | Role |
|---|---|
src/dltb/continuous.py |
loop skeleton, modes, tails, tag grammar, run(args, make_conditioning=…) API; dltb-continuous CLI |
src/dltb/imaging.py |
run_pass (one img2img pass), prepare_frame, make_reprojector (Farneback flow + remap + disocclusion mask), make_generator (seed policy) |
src/dltb/klein.py |
dltb-klein: the dual-ref make_conditioning passed into continuous.run |
src/dltb/models.py |
ModelSpec/load_pipeline (device, dtype, offload, MPS slicing), geometry and requirement checks |
src/dltb/output.py |
run_dir layout under untracked/output_<model>/ |
scripts/sweep.sh |
video sweep driver exercising modes, blends, and tails |
| this doc | algorithm reference for the loop |