Deep Learning Tripping Balls
dltb doc algorithm flux2klein.md
16 kB
Markdown
at main

FLUX.2 klein in dltb — the video pipeline under test #

Reference documentation for the klein video loop: how the current pipeline (dltb-klein --conditioning dual-ref --reproject [--reproject-mask], driven by scripts/sweep-klein.sh / scripts/sweep-klein-mask.sh) works end to end, and how it got here. Investigation details live in NOTES.md (sections linked below); this doc is the consolidated, current-state view.

Status at time of writing (2026-09-13, concluded): the reproject-mask A/B was analyzed and positive on both klein models, but the neutral-prompt control showed the per-pass compounding deformation is intrinsic to the closed klein self-reference loop (it persists, weaker, with an empty prompt). The experiment line is concluded and back on the drawing board; --reproject-mask remains opt-in (the planned follow-through — mask as dual-ref default, sweep axis — is shelved with the line).


1. The conceptual frame #

The experiment simulates a DLSS-5-style inference loop:

P_n = f(state(P_{n-1}), N_n, motion_vectors_n, artistic_direction)

where P is the model's own previous output ("carried temporal state") and N the freshly rendered frame. FLUX.2 klein (4B / 9B) plays the role of f. It is a reference-image editor, unlike the img2img models in this repo:

  • no partial noising, hence no --strength — it regenerates from pure noise attending to reference tokens;
  • --guidance-scale > 1 is inert (CFG hard-disabled in step-wise distilled checkpoints, no guidance embedding; NOTES.md, "FLUX.2 klein: --guidance-scale is inert"). --negative-prompt equally so;
  • the prompt is the de-facto per-pass edit-strength knob (empty/neutral preserves; an edit instruction compounds every pass), and --num-inference-steps (default 4, the card default) is the other direct intensity knob.

2. The loop, per frame (current pipeline) #

Tools: dltb-klein (arg surface) → continuous.run() (loop) with the make_conditioning hook → klein._dual_ref_conditioning. Geometry 768×768 (klein ModelSpec defaults, pass_size=True), bf16, fixed seed 1234 (--fixed-seed default: same noise every pass — DLSS-5-like determinism that holds only while conditioning is bit-identical, see §5).

  1. Ingest: each source frame → prepare_frame: RGB, center-crop to the target aspect, LANCZOS resize to 768².
  2. Frame 1: no state yet; single clean reference N_1 → P_1.
  3. Frames n ≥ 2 (combine(carried=P_{n-1}, new_frame=N_n, prev_source=N_{n-1})):
    1. Flow: estimate_flow(N_{n-1}, N_n) — two Farneback passes on grayscale (forward A→B, backward B→A when masking is on). Flows are estimated between source frames only; the model output never participates.
    2. Warp (--reproject, default on): the carried state is reprojected along the forward flow — target pixel p fetches state at p − fwd(p) (bilinear, BORDER_REPLICATE). With --reproject-mask, disoccluded pixels are patched from N_n first (§4). This warp emulates engine motion vectors and is load-bearing (§5).
    3. Pairing: ordered(state', N_n) → [state', N_n] with --ref-order state-first (default) or [N_n, state'] with frame-first. --anchor-blend is ignored in this mode — so --mode anchored is the only single-reference control under dual-ref (under blend conditioning, a = 1.0 would reproduce it exactly).
  4. Prompt: role-naming, indices derived from ref order — "image {F} is the current frame; keep the appearance of image {S}, <edit instruction>". Note: composition comes from the references; the prompt modulates texture intensity only (§5).
  5. Pass: run_pass → Flux2KleinPipeline(prompt, image=[state', N_n], num_inference_steps=4, guidance_scale=1.0 (spec default; inert anyway), width=height=768, generator) — the list flows through untouched (§3).
  6. Emit: result becomes P_n (the new carried state); appended to processed_stateful.mp4; every 10th frame saved as PNG.
  7. Tails (after the last source frame, branching from one shared end_state.png): freeze = [P, last_source] (static-menu case), free = [P] alone (buffer echo), black = [P, black] (renderer crash). No flow/warp in tails — static input means zero motion vectors.

Run-directory tag: <stem>_dualref[-norepro|-fbmask]_tails<modes><N> (-norepro = warp off, -fbmask = warp + mask on).

3. Inside Flux2KleinPipeline (diffusers 0.40.0) #

What actually happens to the reference list inside one pass:

  1. Prompt → Qwen3 text encoder → text tokens.

  2. Reference preprocessing — each list element independently: area check (downscale to ≤ 1 MP if needed; 768² = 0.59 MP passes untouched), dimensions snapped to multiples of vae_scale_factor * 2, normalize, VAE-encode, _patchify_latents (2×2 patch packing), then batch-norm whitening with the VAE's running stats. Each 768² reference ≈ 2.3k tokens (dim 128).

  3. Token coordinates — the decisive detail (_prepare_image_ids): every reference token gets a 4D coordinate (T, H, W, L); reference i sits at T = 10 + 10·i, the generation target's noise tokens at T = 0. One spatiotemporal sequence enters attention:

    [text] [target noise @ t=0] [state' tokens @ t=10] [N_n tokens @ t=20]
    

    The reference list is encoded as a temporally-indexed sequence of observations of one scene — restoration-style multi-frame conditioning, not "named subjects" an editor arbitrates between. Consequences in §5. Side effect: total reference token count (~4.6k for two refs) feeds compute_empirical_mu, which shifts the timestep schedule — dual-ref subtly changes the denoising schedule vs single-ref.

  4. Denoising (4 steps): torch.cat([latents, image_latents], dim=1) — target and reference tokens concatenated on the sequence axis; the transformer self-attends across all of them each step. References are clean (never noised); the target starts as pure noise and is denoised toward a consensus that preserves reference appearance. CFG branch unreachable (is_distilled).

  5. Decode: unpatchify → VAE decode → 768² PIL → P_n.

The mask (§4) operates entirely upstream of all of this: it edits the pixel content of reference 1 before torch ever sees it.

4. The disocclusion mask (--reproject-mask) #

imaging.make_reprojector(mask_disocclusions=True, fb_tau=1.5, dilate_px=5):

  • Forward–backward circularity: if target pixel p truly corresponds to something in the previous frame, the two independent flow estimates must agree on one track: fwd(p) ≈ −bwd(p). Round-trip residual |fwd(p) + bwd(p)| > fb_tau (1.5 px) → no trustworthy history at p (newly revealed content, newly occluded content, or flow hallucination — all three mean: don't use the warp).
  • Out-of-bounds sampling: p − fwd(p) outside the frame would have been BORDER_REPLICATE-smeared → masked unconditionally.
  • Dilation (5 px ellipse) grows the mask over the smear's penumbra (bilinear rim mixing, threshold misses).
  • Composite: masked pixels take the fresh frame's pixel — a disocclusion is by definition content just revealed in N_n; the fresh frame is its only witness. Hard binary swap (feathered compositing is parked, NOTES.md).
  • In blend conditioning the same warp runs pre-blend; at masked pixels the alpha blend then yields exactly N_n. At a = 1.0 the whole warp — mask included — is skipped, since the blend discards it (stateful a=1 equals --mode anchored; the skip lives in continuous._blend_conditioning, not in the dual-ref hook, which is unaffected).

Cost: one extra CPU Farneback pass per frame. Measured on the tracked clip (consecutive pairs, run-identical constants): fires on 1.9–7.3% of pixels per pass (≈4% mean), OOB 0.2–0.6%; fires concentrate at genuine motion (a flow-magnitude gate changes nothing); false positives are benign (fallback is same-scene content). Because the patch becomes part of P_n, the next pass carries it forward — TAA-style history rebuild, but performed implicitly by the generative loop instead of an explicit confidence buffer.

Known limits: the forward flow is used as a per-target-pixel backward map (small-motion approximation, fine at this clip's 1–5 px inter-frame displacement); no temporal confidence weighting; fb_tau/dilate_px are clip-tuned constants.

5. Established empirical properties #

The loop's "physics", each verified on-GPU (NOTES.md, "Klein dual-ref GPU validation" and "reproject-mask A/B"):

  1. Reprojection is load-bearing. Without the warp, dual-ref freezes into a static consensus within a few frames — regardless of reference order, prompt, or model (4b and 9b; five-run freeze ledger). At 59.94 fps the fresh frame's per-pass pull (~1.6 MAD toward maximally different content) is far below inter-frame displacement, so only the warped state carries position. The warp also puts [state', N_n] back inside the training distribution (two aligned observations of one scene).

  2. Restoration-consensus regime. The T-coordinate encoding (§3) makes the reference bundle "one scene observed repeatedly"; the output is a stable consensus of it. Corollaries: any second reference — even black, whose content is not reproduced — stabilizes against buffer-echo drift ([P] alone drifts hard); reference order and role-naming prompts do not rescue motion (REF_ORDER deprioritized); composition comes from the references, the prompt modulates texture only.

  3. Same-seed ≠ same trajectory under perturbed conditioning. Klein redraws texture from noise every pass, so a ~4%-of-pixels conditioning perturbation (the mask) decorrelates the render within ~10 passes (frame-identical at pass 1, MAD 34 by pass 10, plateau ~90). A/Bs of this loop must be statistical, not frame-identity.

  4. The mask works end-to-end (4b A/B, 300 frames + freeze tail, identical seed/prompt/steps). Output-vs-source inside vs outside disocclusion zones, mean over 30 saved frames:

    leg MAD | disocc MAD | rest HF energy | disocc saturation | disocc
    plain (warp only) 92.5 89.0 26.8 146
    fbmask (warp+mask) 31.0 78.2 10.8 105
    source (ground truth) — — 5.5 98

    Patched fresh-frame content demonstrably propagates into the output; the unmasked leg's streaks are high-frequency, oversaturated replicate smears (they read as "blurry" at video speed but are spectrally sharp). Residual ~2× HF in fbmask zones — split finding from the neutral-prompt control: the spectral/color vividness was prompt-driven (zone HF/saturation drop to near-source with PROMPT="" on both legs), while the global compounding drift is prompt-independent (Δoriginal neutral 88.5/89.0 vs enhanced 78.4/92.0 — end frames similar either way). Under a neutral prompt the mask wins on every zone axis (MAD 20.9 vs 48.9, churn 22.6 vs 52.5). Global "heavy deformation" in both legs is this by-design compounding drift of a 300-pass run, orthogonal to the mask.

    9b confirmation (same script, both legs): verdict replicates — zones MAD 58.5→33.1, HF 21.8→11.9, saturation 215→123 (source 5.5/98). The twist: 9b's unmasked streaks are not raw smears but plausible pseudo-3D structure with extreme color (saturation 2.2× source) — a stronger editor resolves invalid history into coherent-looking artifacts. fbmask zones are nearly model-independent (correct fresh content is preserved either way). Runtime measured: ~6.2 s/pass (9b) / ~3.2 s/pass (4b) on RTX 6000 Ada, 768², dual-ref, 4 steps.

  5. Blend conditioning (historical): klein under pixel-blend anchoring was too weak at --anchor-blend 0.6–0.8 and "melted" at 0.1 — the anomaly that started the dual-ref redesign.

6. Historical development #

  • 2026-09-12 — calibration anomaly. Cross-model stateful sweep (blend conditioning): sd/sdxl-turbo and flux-schnell showed the expected cartoonish degradation; flux2-klein-9b was too weak at 0.6–0.8 and melted at 0.1. Experiments paused pending a topology rethink.
  • 2026-09-12 — dltb-klein redesign. New tool (dltb-klein) + scripts/sweep-klein.sh: prompt-as-strength ladder, steps probes; klein framed as a reference editor (prompt = edit knob).
  • 2026-09-12 — guidance proved inert. Mid-sweep log flood; three-point source proof in diffusers 0.40.0; guidance probes removed, replaced by steps probes; one-time warning added to the tool.
  • 2026-09-13 — dual-reference conditioning implemented. continuous.run() generalized with the make_conditioning factory hook; klein._dual_ref_conditioning passes [R(P_{n-1}), N_n] as two clean references (verified: the pipeline natively accepts a list, concatenates per-reference latents on the sequence axis). --ref-order, role-naming prompts, REPROJECT=ab A/B wired into the sweep.
  • 2026-09-13 — first dual-ref sweep, interrupted. Early legs looked "similar to stateful blend"; suspected motion artifacts; sweep stopped after enhance-slight (partial legs preserved). Debug session: empty- --input hardening; norepro freeze observed; frame-first and role-naming probes (both freeze); analyze_drift shows locked composition + perpetual texture churn.
  • 2026-09-13 — mechanism found. Diffusers source read: T-coordinate reference encoding → restoration-consensus regime. Tail-based reference-gain readout (freeze/free/black from one shared state): second reference has real but low per-pass gain that compounds; any second reference stabilizes; black content is not reproduced. Verdict: reproject load-bearing, dual-ref-norepro structurally dead for motion.
  • 2026-09-13 — reproject-mask A/B, then conclusion. FB-consistency disocclusion mask implemented in imaging.make_reprojector; --reproject-mask flag in dltb-continuous/dltb-klein; scripts/sweep-klein-mask.sh runner. Results in §5.4 — confirmed on 4b and 9b. The neutral-prompt control (same day) showed the compounding deformation is intrinsic to the closed loop, not prompt-driven; the experiment line was concluded and shelved (mask stays opt-in), pending a new approach to the self-reference drift problem.

7. File map #

File Role
src/dltb/klein.py arg surface, dual-ref conditioning factory, inert-guidance warning
src/dltb/continuous.py the loop, tail phases, tag scheme, _blend_conditioning
src/dltb/imaging.py run_pass, prepare_frame, make_reprojector (warp + mask)
src/dltb/models.py klein ModelSpecs (768², bf16, guidance 1.0, no strength, pass_size)
scripts/sweep-klein.sh prompt ladder + steps probes, both conditionings, reproject A/B
scripts/sweep-klein-mask.sh the mask A/B runner (fbmask vs plain legs)
NOTES.md dated investigation log (sections of 2026-09-12/13)

8. Open questions (after the 2026-09-13 conclusion) #

The mask, reproject, and dual-ref characterization stand (§5); the open problem is the intrinsic per-pass compounding of a self-referencing regeneration loop — even an empty prompt deforms the output over ~100s of passes. Drawing-board candidates:

  • Explicit history/confidence buffer OUTSIDE the model (true TAA-style accumulation, with klein only as the per-frame renderer).
  • Periodic hard re-anchoring (state reset every k frames) to bound drift.
  • A different conditioning topology entirely (klein's T-coordinate multi-reference layout invites feeding genuinely time-spaced frames — motion extrapolation rather than restoration).
  • Or: accept the drift as the object of study (it is the drift experiment's subject) and scope klein out of the "healthy pipeline" simulation.
  • Untouched: whether klein's multi-reference T-coordinates could be exploited directly (e.g., feeding genuinely time-spaced references for motion extrapolation rather than restoration).