FLUX.2 klein in dltb — the video pipeline under test #
Reference documentation for the klein video loop: how the current pipeline
(dltb-klein --conditioning dual-ref --reproject [--reproject-mask], driven
by scripts/sweep-klein.sh / scripts/sweep-klein-mask.sh) works end to
end, and how it got here. Investigation details live in NOTES.md (sections
linked below); this doc is the consolidated, current-state view.
Status at time of writing (2026-09-13, concluded): the reproject-mask A/B
was analyzed and positive on both klein models, but the neutral-prompt
control showed the per-pass compounding deformation is intrinsic to the
closed klein self-reference loop (it persists, weaker, with an empty
prompt). The experiment line is concluded and back on the drawing board;
--reproject-mask remains opt-in (the planned follow-through — mask as
dual-ref default, sweep axis — is shelved with the line).
1. The conceptual frame #
The experiment simulates a DLSS-5-style inference loop:
P_n = f(state(P_{n-1}), N_n, motion_vectors_n, artistic_direction)
where P is the model's own previous output ("carried temporal state") and
N the freshly rendered frame. FLUX.2 klein (4B / 9B) plays the role of
f. It is a reference-image editor, unlike the img2img models in this
repo:
- no partial noising, hence no
--strength— it regenerates from pure noise attending to reference tokens; --guidance-scale > 1is inert (CFG hard-disabled in step-wise distilled checkpoints, no guidance embedding; NOTES.md, "FLUX.2 klein:--guidance-scaleis inert").--negative-promptequally so;- the prompt is the de-facto per-pass edit-strength knob (empty/neutral
preserves; an edit instruction compounds every pass), and
--num-inference-steps(default 4, the card default) is the other direct intensity knob.
2. The loop, per frame (current pipeline) #
Tools: dltb-klein (arg surface) → continuous.run() (loop) with the
make_conditioning hook → klein._dual_ref_conditioning. Geometry 768×768
(klein ModelSpec defaults, pass_size=True), bf16, fixed seed 1234
(--fixed-seed default: same noise every pass — DLSS-5-like determinism
that holds only while conditioning is bit-identical, see §5).
- Ingest: each source frame →
prepare_frame: RGB, center-crop to the target aspect, LANCZOS resize to 768². - Frame 1: no state yet; single clean reference
N_1→P_1. - Frames n ≥ 2 (
combine(carried=P_{n-1}, new_frame=N_n, prev_source=N_{n-1})):- Flow:
estimate_flow(N_{n-1}, N_n)— two Farneback passes on grayscale (forward A→B, backward B→A when masking is on). Flows are estimated between source frames only; the model output never participates. - Warp (
--reproject, default on): the carried state is reprojected along the forward flow — target pixelpfetches state atp − fwd(p)(bilinear,BORDER_REPLICATE). With--reproject-mask, disoccluded pixels are patched fromN_nfirst (§4). This warp emulates engine motion vectors and is load-bearing (§5). - Pairing:
ordered(state', N_n)→[state', N_n]with--ref-order state-first(default) or[N_n, state']withframe-first.--anchor-blendis ignored in this mode — so--mode anchoredis the only single-reference control under dual-ref (under blend conditioning,a = 1.0would reproduce it exactly).
- Flow:
- Prompt: role-naming, indices derived from ref order —
"image {F} is the current frame; keep the appearance of image {S}, <edit instruction>". Note: composition comes from the references; the prompt modulates texture intensity only (§5). - Pass:
run_pass→Flux2KleinPipeline(prompt, image=[state', N_n], num_inference_steps=4, guidance_scale=1.0 (spec default; inert anyway), width=height=768, generator)— the list flows through untouched (§3). - Emit: result becomes
P_n(the new carried state); appended toprocessed_stateful.mp4; every 10th frame saved as PNG. - Tails (after the last source frame, branching from one shared
end_state.png):freeze = [P, last_source](static-menu case),free = [P]alone (buffer echo),black = [P, black](renderer crash). No flow/warp in tails — static input means zero motion vectors.
Run-directory tag: <stem>_dualref[-norepro|-fbmask]_tails<modes><N>
(-norepro = warp off, -fbmask = warp + mask on).
3. Inside Flux2KleinPipeline (diffusers 0.40.0) #
What actually happens to the reference list inside one pass:
-
Prompt → Qwen3 text encoder → text tokens.
-
Reference preprocessing — each list element independently: area check (downscale to ≤ 1 MP if needed; 768² = 0.59 MP passes untouched), dimensions snapped to multiples of
vae_scale_factor * 2, normalize, VAE-encode,_patchify_latents(2×2 patch packing), then batch-norm whitening with the VAE's running stats. Each 768² reference ≈ 2.3k tokens (dim 128). -
Token coordinates — the decisive detail (
_prepare_image_ids): every reference token gets a 4D coordinate(T, H, W, L); reference i sits atT = 10 + 10·i, the generation target's noise tokens atT = 0. One spatiotemporal sequence enters attention:[text] [target noise @ t=0] [state' tokens @ t=10] [N_n tokens @ t=20]The reference list is encoded as a temporally-indexed sequence of observations of one scene — restoration-style multi-frame conditioning, not "named subjects" an editor arbitrates between. Consequences in §5. Side effect: total reference token count (~4.6k for two refs) feeds
compute_empirical_mu, which shifts the timestep schedule — dual-ref subtly changes the denoising schedule vs single-ref. -
Denoising (4 steps):
torch.cat([latents, image_latents], dim=1)— target and reference tokens concatenated on the sequence axis; the transformer self-attends across all of them each step. References are clean (never noised); the target starts as pure noise and is denoised toward a consensus that preserves reference appearance. CFG branch unreachable (is_distilled). -
Decode: unpatchify → VAE decode → 768² PIL →
P_n.
The mask (§4) operates entirely upstream of all of this: it edits the pixel content of reference 1 before torch ever sees it.
4. The disocclusion mask (--reproject-mask) #
imaging.make_reprojector(mask_disocclusions=True, fb_tau=1.5, dilate_px=5):
- Forward–backward circularity: if target pixel
ptruly corresponds to something in the previous frame, the two independent flow estimates must agree on one track:fwd(p) ≈ −bwd(p). Round-trip residual|fwd(p) + bwd(p)| > fb_tau(1.5 px) → no trustworthy history atp(newly revealed content, newly occluded content, or flow hallucination — all three mean: don't use the warp). - Out-of-bounds sampling:
p − fwd(p)outside the frame would have beenBORDER_REPLICATE-smeared → masked unconditionally. - Dilation (5 px ellipse) grows the mask over the smear's penumbra (bilinear rim mixing, threshold misses).
- Composite: masked pixels take the fresh frame's pixel — a
disocclusion is by definition content just revealed in
N_n; the fresh frame is its only witness. Hard binary swap (feathered compositing is parked, NOTES.md). - In
blendconditioning the same warp runs pre-blend; at masked pixels the alpha blend then yields exactlyN_n. Ata = 1.0the whole warp — mask included — is skipped, since the blend discards it (stateful a=1equals--mode anchored; the skip lives incontinuous._blend_conditioning, not in the dual-ref hook, which is unaffected).
Cost: one extra CPU Farneback pass per frame. Measured on the tracked clip
(consecutive pairs, run-identical constants): fires on 1.9–7.3% of
pixels per pass (≈4% mean), OOB 0.2–0.6%; fires concentrate at genuine
motion (a flow-magnitude gate changes nothing); false positives are benign
(fallback is same-scene content). Because the patch becomes part of P_n,
the next pass carries it forward — TAA-style history rebuild, but performed
implicitly by the generative loop instead of an explicit confidence buffer.
Known limits: the forward flow is used as a per-target-pixel backward map
(small-motion approximation, fine at this clip's 1–5 px inter-frame
displacement); no temporal confidence weighting; fb_tau/dilate_px are
clip-tuned constants.
5. Established empirical properties #
The loop's "physics", each verified on-GPU (NOTES.md, "Klein dual-ref GPU validation" and "reproject-mask A/B"):
-
Reprojection is load-bearing. Without the warp, dual-ref freezes into a static consensus within a few frames — regardless of reference order, prompt, or model (4b and 9b; five-run freeze ledger). At 59.94 fps the fresh frame's per-pass pull (~1.6 MAD toward maximally different content) is far below inter-frame displacement, so only the warped state carries position. The warp also puts
[state', N_n]back inside the training distribution (two aligned observations of one scene). -
Restoration-consensus regime. The T-coordinate encoding (§3) makes the reference bundle "one scene observed repeatedly"; the output is a stable consensus of it. Corollaries: any second reference — even black, whose content is not reproduced — stabilizes against buffer-echo drift (
[P]alone drifts hard); reference order and role-naming prompts do not rescue motion (REF_ORDER deprioritized); composition comes from the references, the prompt modulates texture only. -
Same-seed ≠ same trajectory under perturbed conditioning. Klein redraws texture from noise every pass, so a ~4%-of-pixels conditioning perturbation (the mask) decorrelates the render within ~10 passes (frame-identical at pass 1, MAD 34 by pass 10, plateau ~90). A/Bs of this loop must be statistical, not frame-identity.
-
The mask works end-to-end (4b A/B, 300 frames + freeze tail, identical seed/prompt/steps). Output-vs-source inside vs outside disocclusion zones, mean over 30 saved frames:
leg MAD | disocc MAD | rest HF energy | disocc saturation | disocc plain (warp only) 92.5 89.0 26.8 146 fbmask (warp+mask) 31.0 78.2 10.8 105 source (ground truth) — — 5.5 98 Patched fresh-frame content demonstrably propagates into the output; the unmasked leg's streaks are high-frequency, oversaturated replicate smears (they read as "blurry" at video speed but are spectrally sharp). Residual ~2× HF in fbmask zones — split finding from the neutral-prompt control: the spectral/color vividness was prompt-driven (zone HF/saturation drop to near-source with
PROMPT=""on both legs), while the global compounding drift is prompt-independent (Δoriginal neutral 88.5/89.0 vs enhanced 78.4/92.0 — end frames similar either way). Under a neutral prompt the mask wins on every zone axis (MAD 20.9 vs 48.9, churn 22.6 vs 52.5). Global "heavy deformation" in both legs is this by-design compounding drift of a 300-pass run, orthogonal to the mask.9b confirmation (same script, both legs): verdict replicates — zones MAD 58.5→33.1, HF 21.8→11.9, saturation 215→123 (source 5.5/98). The twist: 9b's unmasked streaks are not raw smears but plausible pseudo-3D structure with extreme color (saturation 2.2× source) — a stronger editor resolves invalid history into coherent-looking artifacts. fbmask zones are nearly model-independent (correct fresh content is preserved either way). Runtime measured: ~6.2 s/pass (9b) / ~3.2 s/pass (4b) on RTX 6000 Ada, 768², dual-ref, 4 steps.
-
Blend conditioning (historical): klein under pixel-blend anchoring was too weak at
--anchor-blend 0.6–0.8and "melted" at 0.1 — the anomaly that started the dual-ref redesign.
6. Historical development #
- 2026-09-12 — calibration anomaly. Cross-model stateful sweep (blend
conditioning): sd/sdxl-turbo and flux-schnell showed the expected
cartoonish degradation;
flux2-klein-9bwas too weak at 0.6–0.8 and melted at 0.1. Experiments paused pending a topology rethink. - 2026-09-12 — dltb-klein redesign. New tool (
dltb-klein) +scripts/sweep-klein.sh: prompt-as-strength ladder, steps probes; klein framed as a reference editor (prompt = edit knob). - 2026-09-12 — guidance proved inert. Mid-sweep log flood; three-point source proof in diffusers 0.40.0; guidance probes removed, replaced by steps probes; one-time warning added to the tool.
- 2026-09-13 — dual-reference conditioning implemented.
continuous.run()generalized with themake_conditioningfactory hook;klein._dual_ref_conditioningpasses[R(P_{n-1}), N_n]as two clean references (verified: the pipeline natively accepts a list, concatenates per-reference latents on the sequence axis).--ref-order, role-naming prompts,REPROJECT=abA/B wired into the sweep. - 2026-09-13 — first dual-ref sweep, interrupted. Early legs looked
"similar to stateful blend"; suspected motion artifacts; sweep stopped
after
enhance-slight(partial legs preserved). Debug session: empty---inputhardening; norepro freeze observed; frame-first and role-naming probes (both freeze);analyze_driftshows locked composition + perpetual texture churn. - 2026-09-13 — mechanism found. Diffusers source read: T-coordinate reference encoding → restoration-consensus regime. Tail-based reference-gain readout (freeze/free/black from one shared state): second reference has real but low per-pass gain that compounds; any second reference stabilizes; black content is not reproduced. Verdict: reproject load-bearing, dual-ref-norepro structurally dead for motion.
- 2026-09-13 — reproject-mask A/B, then conclusion. FB-consistency
disocclusion mask implemented in
imaging.make_reprojector;--reproject-maskflag indltb-continuous/dltb-klein;scripts/sweep-klein-mask.shrunner. Results in §5.4 — confirmed on 4b and 9b. The neutral-prompt control (same day) showed the compounding deformation is intrinsic to the closed loop, not prompt-driven; the experiment line was concluded and shelved (mask stays opt-in), pending a new approach to the self-reference drift problem.
7. File map #
| File | Role |
|---|---|
src/dltb/klein.py |
arg surface, dual-ref conditioning factory, inert-guidance warning |
src/dltb/continuous.py |
the loop, tail phases, tag scheme, _blend_conditioning |
src/dltb/imaging.py |
run_pass, prepare_frame, make_reprojector (warp + mask) |
src/dltb/models.py |
klein ModelSpecs (768², bf16, guidance 1.0, no strength, pass_size) |
scripts/sweep-klein.sh |
prompt ladder + steps probes, both conditionings, reproject A/B |
scripts/sweep-klein-mask.sh |
the mask A/B runner (fbmask vs plain legs) |
NOTES.md |
dated investigation log (sections of 2026-09-12/13) |
8. Open questions (after the 2026-09-13 conclusion) #
The mask, reproject, and dual-ref characterization stand (§5); the open problem is the intrinsic per-pass compounding of a self-referencing regeneration loop — even an empty prompt deforms the output over ~100s of passes. Drawing-board candidates:
- Explicit history/confidence buffer OUTSIDE the model (true TAA-style accumulation, with klein only as the per-frame renderer).
- Periodic hard re-anchoring (state reset every k frames) to bound drift.
- A different conditioning topology entirely (klein's T-coordinate multi-reference layout invites feeding genuinely time-spaced frames — motion extrapolation rather than restoration).
- Or: accept the drift as the object of study (it is the drift experiment's subject) and scope klein out of the "healthy pipeline" simulation.
- Untouched: whether klein's multi-reference T-coordinates could be exploited directly (e.g., feeding genuinely time-spaced references for motion extrapolation rather than restoration).