# FLUX.2 klein in dltb — the video pipeline under test Reference documentation for the klein video loop: how the current pipeline (`dltb-klein --conditioning dual-ref --reproject [--reproject-mask]`, driven by `scripts/sweep-klein.sh` / `scripts/sweep-klein-mask.sh`) works end to end, and how it got here. Investigation details live in NOTES.md (sections linked below); this doc is the consolidated, current-state view. Status at time of writing (2026-09-13, concluded): the reproject-mask A/B was analyzed and positive on both klein models, but the neutral-prompt control showed the per-pass compounding deformation is intrinsic to the closed klein self-reference loop (it persists, weaker, with an empty prompt). The experiment line is concluded and back on the drawing board; `--reproject-mask` remains opt-in (the planned follow-through — mask as dual-ref default, sweep axis — is shelved with the line). --- ## 1. The conceptual frame The experiment simulates a DLSS-5-style inference loop: P_n = f(state(P_{n-1}), N_n, motion_vectors_n, artistic_direction) where `P` is the model's own previous output ("carried temporal state") and `N` the freshly rendered frame. FLUX.2 klein (4B / 9B) plays the role of `f`. It is a **reference-image editor**, unlike the img2img models in this repo: - no partial noising, hence **no `--strength`** — it regenerates from pure noise attending to reference tokens; - `--guidance-scale > 1` is **inert** (CFG hard-disabled in step-wise distilled checkpoints, no guidance embedding; NOTES.md, "FLUX.2 klein: `--guidance-scale` is inert"). `--negative-prompt` equally so; - the **prompt is the de-facto per-pass edit-strength knob** (empty/neutral preserves; an edit instruction compounds every pass), and `--num-inference-steps` (default 4, the card default) is the other direct intensity knob. ## 2. The loop, per frame (current pipeline) Tools: `dltb-klein` (arg surface) → `continuous.run()` (loop) with the `make_conditioning` hook → `klein._dual_ref_conditioning`. Geometry 768×768 (klein `ModelSpec` defaults, `pass_size=True`), bf16, fixed seed 1234 (`--fixed-seed` default: same noise every pass — DLSS-5-like determinism that holds only while conditioning is bit-identical, see §5). 1. **Ingest**: each source frame → `prepare_frame`: RGB, center-crop to the target aspect, LANCZOS resize to 768². 2. **Frame 1**: no state yet; single clean reference `N_1` → `P_1`. 3. **Frames n ≥ 2** (`combine(carried=P_{n-1}, new_frame=N_n, prev_source=N_{n-1})`): 1. **Flow**: `estimate_flow(N_{n-1}, N_n)` — two Farneback passes on grayscale (forward A→B, backward B→A when masking is on). Flows are estimated **between source frames only**; the model output never participates. 2. **Warp** (`--reproject`, default on): the carried state is reprojected along the forward flow — target pixel `p` fetches state at `p − fwd(p)` (bilinear, `BORDER_REPLICATE`). With `--reproject-mask`, disoccluded pixels are patched from `N_n` first (§4). This warp emulates engine motion vectors and is **load-bearing** (§5). 3. **Pairing**: `ordered(state', N_n)` → `[state', N_n]` with `--ref-order state-first` (default) or `[N_n, state']` with `frame-first`. `--anchor-blend` is ignored in this mode — so `--mode anchored` is the only single-reference control under dual-ref (under blend conditioning, `a = 1.0` would reproduce it exactly). 4. **Prompt**: role-naming, indices derived from ref order — `"image {F} is the current frame; keep the appearance of image {S}, "`. Note: composition comes from the references; the prompt modulates **texture intensity only** (§5). 5. **Pass**: `run_pass` → `Flux2KleinPipeline(prompt, image=[state', N_n], num_inference_steps=4, guidance_scale=1.0 (spec default; inert anyway), width=height=768, generator)` — the list flows through untouched (§3). 6. **Emit**: result becomes `P_n` (the new carried state); appended to `processed_stateful.mp4`; every 10th frame saved as PNG. 7. **Tails** (after the last source frame, branching from one shared `end_state.png`): `freeze = [P, last_source]` (static-menu case), `free = [P]` alone (buffer echo), `black = [P, black]` (renderer crash). **No flow/warp in tails** — static input means zero motion vectors. Run-directory tag: `_dualref[-norepro|-fbmask]_tails` (`-norepro` = warp off, `-fbmask` = warp + mask on). ## 3. Inside `Flux2KleinPipeline` (diffusers 0.40.0) What actually happens to the reference list inside one pass: 1. **Prompt** → Qwen3 text encoder → text tokens. 2. **Reference preprocessing — each list element independently**: area check (downscale to ≤ 1 MP if needed; 768² = 0.59 MP passes untouched), dimensions snapped to multiples of `vae_scale_factor * 2`, normalize, VAE-encode, `_patchify_latents` (2×2 patch packing), then batch-norm whitening with the VAE's running stats. Each 768² reference ≈ 2.3k tokens (dim 128). 3. **Token coordinates — the decisive detail** (`_prepare_image_ids`): every reference token gets a 4D coordinate `(T, H, W, L)`; reference *i* sits at `T = 10 + 10·i`, the generation target's noise tokens at `T = 0`. One spatiotemporal sequence enters attention: [text] [target noise @ t=0] [state' tokens @ t=10] [N_n tokens @ t=20] The reference list is encoded as a **temporally-indexed sequence of observations of one scene** — restoration-style multi-frame conditioning, not "named subjects" an editor arbitrates between. Consequences in §5. Side effect: total reference token count (~4.6k for two refs) feeds `compute_empirical_mu`, which shifts the timestep schedule — dual-ref subtly changes the denoising schedule vs single-ref. 4. **Denoising (4 steps)**: `torch.cat([latents, image_latents], dim=1)` — target and reference tokens concatenated on the **sequence axis**; the transformer self-attends across all of them each step. References are clean (never noised); the target starts as pure noise and is denoised toward a **consensus that preserves reference appearance**. CFG branch unreachable (`is_distilled`). 5. **Decode**: unpatchify → VAE decode → 768² PIL → `P_n`. The mask (§4) operates entirely **upstream** of all of this: it edits the pixel content of reference 1 before torch ever sees it. ## 4. The disocclusion mask (`--reproject-mask`) `imaging.make_reprojector(mask_disocclusions=True, fb_tau=1.5, dilate_px=5)`: - **Forward–backward circularity**: if target pixel `p` truly corresponds to something in the previous frame, the two independent flow estimates must agree on one track: `fwd(p) ≈ −bwd(p)`. Round-trip residual `|fwd(p) + bwd(p)| > fb_tau` (1.5 px) → no trustworthy history at `p` (newly revealed content, newly occluded content, or flow hallucination — all three mean: don't use the warp). - **Out-of-bounds sampling**: `p − fwd(p)` outside the frame would have been `BORDER_REPLICATE`-smeared → masked unconditionally. - **Dilation** (5 px ellipse) grows the mask over the smear's penumbra (bilinear rim mixing, threshold misses). - **Composite**: masked pixels take the **fresh frame's** pixel — a disocclusion is by definition content just revealed in `N_n`; the fresh frame is its only witness. Hard binary swap (feathered compositing is parked, NOTES.md). - In `blend` conditioning the same warp runs pre-blend; at masked pixels the alpha blend then yields exactly `N_n`. At `a = 1.0` the whole warp — mask included — is skipped, since the blend discards it (`stateful a=1` equals `--mode anchored`; the skip lives in `continuous._blend_conditioning`, not in the dual-ref hook, which is unaffected). Cost: one extra CPU Farneback pass per frame. Measured on the tracked clip (consecutive pairs, run-identical constants): fires on **1.9–7.3%** of pixels per pass (≈4% mean), OOB 0.2–0.6%; fires concentrate at genuine motion (a flow-magnitude gate changes nothing); false positives are benign (fallback is same-scene content). Because the patch becomes part of `P_n`, the next pass carries it forward — TAA-style history rebuild, but performed implicitly by the generative loop instead of an explicit confidence buffer. Known limits: the forward flow is used as a per-target-pixel backward map (small-motion approximation, fine at this clip's 1–5 px inter-frame displacement); no temporal confidence weighting; `fb_tau`/`dilate_px` are clip-tuned constants. ## 5. Established empirical properties The loop's "physics", each verified on-GPU (NOTES.md, "Klein dual-ref GPU validation" and "reproject-mask A/B"): 1. **Reprojection is load-bearing.** Without the warp, dual-ref freezes into a static consensus within a few frames — regardless of reference order, prompt, or model (4b and 9b; five-run freeze ledger). At 59.94 fps the fresh frame's per-pass pull (~1.6 MAD toward maximally different content) is far below inter-frame displacement, so only the warped state carries position. The warp also puts `[state', N_n]` back inside the training distribution (two aligned observations of one scene). 2. **Restoration-consensus regime.** The T-coordinate encoding (§3) makes the reference bundle "one scene observed repeatedly"; the output is a stable consensus of it. Corollaries: any second reference — even black, whose content is not reproduced — stabilizes against buffer-echo drift (`[P]` alone drifts hard); reference order and role-naming prompts do not rescue motion (REF_ORDER deprioritized); composition comes from the references, the prompt modulates texture only. 3. **Same-seed ≠ same trajectory under perturbed conditioning.** Klein redraws texture from noise every pass, so a ~4%-of-pixels conditioning perturbation (the mask) decorrelates the render within ~10 passes (frame-identical at pass 1, MAD 34 by pass 10, plateau ~90). A/Bs of this loop must be statistical, not frame-identity. 4. **The mask works end-to-end** (4b A/B, 300 frames + freeze tail, identical seed/prompt/steps). Output-vs-source inside vs outside disocclusion zones, mean over 30 saved frames: | leg | MAD \| disocc | MAD \| rest | HF energy \| disocc | saturation \| disocc | |---|---|---|---|---| | plain (warp only) | 92.5 | 89.0 | 26.8 | 146 | | fbmask (warp+mask) | **31.0** | 78.2 | **10.8** | **105** | | source (ground truth) | — | — | 5.5 | 98 | Patched fresh-frame content demonstrably propagates into the output; the unmasked leg's streaks are high-frequency, oversaturated replicate smears (they read as "blurry" at video speed but are spectrally sharp). Residual ~2× HF in fbmask zones — split finding from the neutral-prompt control: the **spectral/color vividness was prompt-driven** (zone HF/saturation drop to near-source with `PROMPT=""` on both legs), while the **global compounding drift is prompt-independent** (Δoriginal neutral 88.5/89.0 vs enhanced 78.4/92.0 — end frames similar either way). Under a neutral prompt the mask wins on every zone axis (MAD 20.9 vs 48.9, churn 22.6 vs 52.5). Global "heavy deformation" in both legs is this by-design compounding drift of a 300-pass run, orthogonal to the mask. **9b confirmation** (same script, both legs): verdict replicates — zones MAD 58.5→33.1, HF 21.8→11.9, saturation 215→123 (source 5.5/98). The twist: 9b's unmasked streaks are not raw smears but *plausible pseudo-3D structure* with extreme color (saturation 2.2× source) — a stronger editor resolves invalid history into coherent-looking artifacts. fbmask zones are nearly model-independent (correct fresh content is preserved either way). Runtime measured: ~6.2 s/pass (9b) / ~3.2 s/pass (4b) on RTX 6000 Ada, 768², dual-ref, 4 steps. 5. **Blend conditioning (historical)**: klein under pixel-blend anchoring was too weak at `--anchor-blend 0.6–0.8` and "melted" at 0.1 — the anomaly that started the dual-ref redesign. ## 6. Historical development - **2026-09-12 — calibration anomaly.** Cross-model stateful sweep (blend conditioning): sd/sdxl-turbo and flux-schnell showed the expected cartoonish degradation; `flux2-klein-9b` was too weak at 0.6–0.8 and melted at 0.1. Experiments paused pending a topology rethink. - **2026-09-12 — dltb-klein redesign.** New tool (`dltb-klein`) + `scripts/sweep-klein.sh`: prompt-as-strength ladder, steps probes; klein framed as a reference editor (prompt = edit knob). - **2026-09-12 — guidance proved inert.** Mid-sweep log flood; three-point source proof in diffusers 0.40.0; guidance probes removed, replaced by steps probes; one-time warning added to the tool. - **2026-09-13 — dual-reference conditioning implemented.** `continuous.run()` generalized with the `make_conditioning` factory hook; `klein._dual_ref_conditioning` passes `[R(P_{n-1}), N_n]` as two clean references (verified: the pipeline natively accepts a list, concatenates per-reference latents on the sequence axis). `--ref-order`, role-naming prompts, `REPROJECT=ab` A/B wired into the sweep. - **2026-09-13 — first dual-ref sweep, interrupted.** Early legs looked "similar to stateful blend"; suspected motion artifacts; sweep stopped after `enhance-slight` (partial legs preserved). Debug session: empty- `--input` hardening; norepro freeze observed; frame-first and role-naming probes (both freeze); `analyze_drift` shows locked composition + perpetual texture churn. - **2026-09-13 — mechanism found.** Diffusers source read: T-coordinate reference encoding → restoration-consensus regime. Tail-based reference-gain readout (freeze/free/black from one shared state): second reference has real but low per-pass gain that compounds; any second reference stabilizes; black content is not reproduced. Verdict: reproject load-bearing, dual-ref-norepro structurally dead for motion. - **2026-09-13 — reproject-mask A/B, then conclusion.** FB-consistency disocclusion mask implemented in `imaging.make_reprojector`; `--reproject-mask` flag in `dltb-continuous`/`dltb-klein`; `scripts/sweep-klein-mask.sh` runner. Results in §5.4 — confirmed on 4b and 9b. The neutral-prompt control (same day) showed the compounding deformation is intrinsic to the closed loop, not prompt-driven; the experiment line was concluded and shelved (mask stays opt-in), pending a new approach to the self-reference drift problem. ## 7. File map | File | Role | |---|---| | `src/dltb/klein.py` | arg surface, dual-ref conditioning factory, inert-guidance warning | | `src/dltb/continuous.py` | the loop, tail phases, tag scheme, `_blend_conditioning` | | `src/dltb/imaging.py` | `run_pass`, `prepare_frame`, `make_reprojector` (warp + mask) | | `src/dltb/models.py` | klein `ModelSpec`s (768², bf16, guidance 1.0, no strength, `pass_size`) | | `scripts/sweep-klein.sh` | prompt ladder + steps probes, both conditionings, reproject A/B | | `scripts/sweep-klein-mask.sh` | the mask A/B runner (fbmask vs plain legs) | | `NOTES.md` | dated investigation log (sections of 2026-09-12/13) | ## 8. Open questions (after the 2026-09-13 conclusion) The mask, reproject, and dual-ref characterization stand (§5); the open problem is the **intrinsic per-pass compounding of a self-referencing regeneration loop** — even an empty prompt deforms the output over ~100s of passes. Drawing-board candidates: - Explicit history/confidence buffer OUTSIDE the model (true TAA-style accumulation, with klein only as the per-frame renderer). - Periodic hard re-anchoring (state reset every k frames) to bound drift. - A different conditioning topology entirely (klein's T-coordinate multi-reference layout invites feeding genuinely time-spaced frames — motion extrapolation rather than restoration). - Or: accept the drift as the object of study (it is the drift experiment's subject) and scope klein out of the "healthy pipeline" simulation. - Untouched: whether klein's multi-reference T-coordinates could be exploited directly (e.g., feeding genuinely time-spaced references for motion extrapolation rather than restoration).