Deep Learning Tripping Balls
dltb NOTES.md
44 kB
Markdown
at main

Notes #

Loose ends and follow-ups for imgiter.

Pod SSH: sshd not running by default (2026-09-12; root cause found & solved 2026-09-12) #

Symptom: every non-interactive access path fails while the pod itself is healthy — runpodctl ssh info ip:port gives connection refused, exec python tunnels over that same mapping, and croc relays are independently flaky (see 2026-09-12 transfer session: two dead relays, one Go panic in runpodctl's croc client, then a 10-min receive timeout on a third code). Only the console web terminal and the ssh <pod-id>-<token>@ssh.runpod.io gateway (PTY required — plain ssh host cmd is rejected with "Your SSH client doesn't support PTY"; scripted use needs a pty wrapper + fed commands) work.

Cause (corrected — the original "the pod image does not start sshd" was wrong): the base image's /start.sh does set up sshd — host keys, authorized_keys, service ssh start — but its setup_ssh block is gated on $PUBLIC_KEY, which Runpod injects from the account's registered SSH keys at pod start. This pod had booted with none registered, so PUBLIC_KEY was empty and the whole block no-opped. Not an image defect; baking keys into the image was considered and rejected (README_DEFERRED.md).

Fix (permanent): register a key once, then restart the pod. Every subsequent start brings sshd up automatically — nothing to redo manually (procedure: README_RUNPOD.md §3, "SSH access"):

runpodctl ssh add-key --key-file ~/.ssh/id_ed25519.pub
runpodctl ssh list-keys                     # confirm it's on the account
# then pod stop && pod start — keys added after boot are ignored until restart

Manual fallback (still valid, for a pod you cannot restart mid-run; ephemeral — redo after every pod start):

ssh-keygen -A          # generates /etc/ssh/ssh_host_* keys
service ssh start      # "no hostkeys available -- exiting" without the line above
ss -tlnp | grep :22
echo '<pubkey line>' >> /root/.ssh/authorized_keys   # else publickey auth fails

Troubleshooting notes worth keeping:

  • The gateway authenticates against the Runpod account; the in-container sshd against /root/.ssh/authorized_keys — two separate trust paths, each can break independently. A pod reachable via gateway but refused on the ip:port mapping means in-container sshd is down (this incident); the reverse (ip:port OK, gateway rejected) would point at the account side.
  • ssh <pod-id>-<token>@ssh.runpod.io needs a PTY — scripted use requires a pty wrapper with fed commands; prefer the ip:port mapping for automation.
  • Host keys, authorized_keys and the running sshd live on the ephemeral container disk and vanish on stop/restart/recreate. With the account key registered, /start.sh regenerates all of it at every boot.
  • After a stop→start the external port is reassigned, and the first runpodctl ssh info can report a stale port for ~90 s — retry until a connection succeeds before assuming sshd is down.
  • With sshd reachable, rsync -rtP --no-owner --no-group -e "ssh -i <key> -p <port>" is the preferred bulk-transfer path (measured ~14 MB/s). -r/-t are load-bearing: plain -P on a directory prints skipping directory . and transfers nothing, and without -t the quick-check has no preserved mtime to compare. Verified 2026-09-13 (macOS openrsync, protocol 29 ↔ pod rsync 3.2.7, protocol 31): incremental re-runs skip unchanged files, -P resumes an interrupted transfer from the partial, and a file pulled while it was still being written was re-synced correctly by a later run (final sha256 match). Add -c to verify by checksum instead.

Key choice for the template's PUBLIC_KEY (2026-09-13): the imgiter template sets PUBLIC_KEY explicitly (currently the mm-mbp key), and the explicit env value appears to replace Runpod's account-key injection — the 2026-09-13 pod's /root/.ssh/authorized_keys held only mm-mbp, not the other account-registered keys (runpodctl-ssh-key, affine@lat5400). The visible failure: runpodctl ssh info prints its suggested command with -i ~/.runpod/ssh/runpodctl-ssh-key, but that key gets Permission denied; direct SSH works only with -i ~/.ssh/id_ed25519 (mm-mbp). Next time set the template's PUBLIC_KEY to the public half of ~/.runpod/ssh/runpodctl-ssh-key (~/.runpod/ssh/runpodctl-ssh-key.pub, account-registered via runpodctl ssh add-key and cloud-synced), so the ssh info command works as printed. The gateway (ssh <pod-id>-<token>@ssh.runpod.io) is unaffected either way — it accepts the key as an argument.

Jupyter Lab: auto-starts, but needs a published port (2026-09-13): the base image's /start.sh starts Jupyter Lab on 0.0.0.0:8888 (preferred dir /workspace, login token = $JUPYTER_PASSWORD) whenever that variable is set — verified on the 2026-09-13 pod, so there is nothing to set up pod-side. What is missing is a mapping: the imgiter template publishes only 22/tcp, Runpod proxies only ports declared at pod creation, and a running pod cannot gain one — so https://<pod-id>-8888.proxy.runpod.net is dead for this pod. Interim access is an SSH tunnel (verified working, login page HTTP 200):

ssh -N -i ~/.ssh/id_ed25519 -p <port> -L 8888:127.0.0.1:8888 root@<ip>
# then http://localhost:8888; password = echo "$JUPYTER_PASSWORD" on the pod

Follow-up: add 8888/http to the template's ports (mark it secure in the console) so future pods can use the proxy URL directly; template edits apply only to newly created/recreated pods. The README §3 "Jupyter Lab (optional, untested)" heading is stale either way — it was exercised 2026-09-13.

macOS ._* AppleDouble files in pod bundles #

Status: fixed (2026-09-12) in scripts/bundle.sh; .gitignore updated. scripts/image-build.sh is not affected — it stages the same tree but never creates a tar archive.

Symptom: extracting a bundle on a pod lists ._Justfile, ._.gitignore, .___init__.py, ._imgiter-<stamp>, etc. next to the real files.

Investigation (all verified on the Mac against bundle/imgiter-202609121612.tar.gz):

  1. Not committed: git ls-files | grep -c '\._' → 0.
  2. Not the staging step: after running bundle.sh's exact staging pipeline, find "$stage" -name '._*' → 0 files.
  3. The staged files carry com.apple.provenance; xattr -l on a staged file shows it. macOS attaches this to files created during the extraction.
  4. The final tar --no-xattrs -czf turns that xattr into AppleDouble members: Python tarfile counted 26 ._* members out of 52.
  5. --no-xattrs does not prevent this. COPYFILE_DISABLE=1 → 0 members; --no-mac-metadata → 0 members.
  6. macOS tar -tzf hides ._* members when listing, which is why the first "the archive is clean" check was wrong.

Impact: inert on Linux (ignored by uv, Python imports, and git), but they bloat the archive and clutter extracted trees. They must not be committed.

Fix applied:

  • scripts/bundle.sh: export COPYFILE_DISABLE=1, plus a post-build guard that verifies the finished archive with python3's tarfile, and on failure deletes it and exits non-zero. The guard skips with a warning if python3 is unavailable.
  • .gitignore: added ._*.

Lessons:

  • Never verify AppleDouble members on macOS with tar -t/tar -tzf; use Python tarfile.
  • Any future tar creation run on macOS needs COPYFILE_DISABLE=1.

Runtime log noise (benign, optional cleanup) #

Observed during the first full sweep on the RTX 6000 Ada. None of these affect results; they are candidates for a cleanup pass on the next CLI/bundle change. Note that suppressing any of them requires a new bundle and a sweep restart.

1. There are modules in AutoencoderKL that should be kept in float32: [] #

Verdict: harmless diffusers false positive. Fires roughly twice per decoded frame for the SD/SDXL models (thousands of lines per run); does not appear for the FLUX/FLUX-2 models.

Mechanism:

  • diffusers/models/modeling_utils.py (~line 1523 in the installed version) has a buggy guard:
    fp32_modules = self._keep_in_fp32_modules or []      # never None
    if dtype_present_in_args and fp32_modules is not None:  # always true
        logger.warning(f"... should be kept in float32: {fp32_modules} ...")
    
    AutoencoderKL does not define _keep_in_fp32_modules, so the message prints [] — nothing actually needs special handling.
  • The warnings come from the SD/SDXL pipeline's intentional VAE upcast around decode: pipeline_stable_diffusion_xl_img2img.py (~lines 1447–1475) does self.upcast_vae() (vae.to(dtype=torch.float32)) before vae.decode(...) and self.vae.to(dtype=torch.float16) after, because the VAE overflows in fp16. Both .to(dtype=...) calls hit the buggy guard.
  • sd-turbo and sdxl-turbo both have "force_upcast": true in their VAE config.json.

Outputs were verified healthy (frame mean/stddev in normal ranges, no black or NaN frames). To silence later, add a targeted filter in dltb/models.py (inside load_pipeline, before the pipeline loads):

import logging

class _DropFp32FalsePositive(logging.Filter):
    def filter(self, record):
        return "should be kept in float32" not in record.getMessage()

logging.getLogger("diffusers.models.modeling_utils").addFilter(_DropFp32FalsePositive())

2. torch.jit.script is deprecated (FutureWarning) #

From diffusers internals (torch/jit/_script.py triggered inside the pipelines). Benign, no action planned beyond upstream updates.

3. requires torchvision (not installed); falling back to CLIPImageProcessorPil (resolved 2026-09-14) #

torchvision was not in the image, so transformers fell back to the PIL image processors (CLIPImageProcessorPil, SiglipImageProcessorPil). Adding dreamsim pulled torchvision into uv.lock (2026-09-14), so the default CLIPImageProcessor / Siglip2ImageProcessor names now resolve to the torchvision-backed classes and the warning is gone.

Verified safe for pixel output: in the pipelines dltb uses, feature_extractor is only called from run_safety_checker (dead here — safety_checker: null in the SD/SDXL repos) and from encode_image (dead — no IP-Adapter/image encoder). Input frames go through diffusers' own VaeImageProcessor / Flux2ImageProcessor, which never touch torchvision, and flux2-klein uses no transformers image processor at all. Re-check both dead branches before enabling a safety checker or an image encoder, and if that ever happens, compare deliberately against pre-2026-09-14 runs instead of mixing them.

4. upcast_vae deprecation #

The SDXL pipeline calls the deprecated upcast_vae() helper internally; if the deprecate line shows up it is internal diffusers churn, not our call site.

5. Siglip2ImageProcessorFast is deprecated #

Transformers deprecation of the Fast suffix on image processors, emitted while loading pipelines that use Siglip/Siglip2 encoders (the FLUX.2 klein models). Internal, benign; fixed by a transformers upgrade.

UserWarning from huggingface_hub/utils/_validators.py. The argument is a no-op in current huggingface_hub; something further up the pipeline stack still passes it. Benign; disappears when that caller is updated.

7. You have disabled the safety checker ... safety_checker=None #

Printed once per SD/SDXL pipeline load: those model repos ship safety_checker: null in model_index.json and the CLI never requests one, so diffusers emits the license reminder. Benign, expected for sd-turbo and sdxl-turbo; it is not something the CLI can (or should) silence.

8. Guidance scale 2.0 is ignored for step-wise distilled models. #

Not benign — it means the run is a no-op duplicate. Emitted once per pass by Flux2KleinPipeline.check_inputs whenever guidance_scale > 1.0 with a klein checkpoint (360 lines per probe run in the sweep log). The value is dropped on the floor: see FLUX.2 klein: --guidance-scale is inert below.

flux2-klein-9b anchor-blend calibration (superseded by the dltb-klein redesign) #

Context: for sd-turbo / sdxl-turbo / flux-schnell, the stateful sweep at --anchor-blend 0.1/0.3/0.5 produced an effect judged too strong (a "cartoonish" degradation), so the re-sweep grid was set to BLENDS="0.6 0.7 0.8" (BASELINE=0.7) for every model.

flux2-klein-9b did not follow that pattern:

  • preview at 0.6/0.7/0.8: too weak to be useful;
  • rerun at 0.1: a noticeable effect but qualitatively different — "melting" rather than the cartoonish degradation the other models show — and it drops off.

Status: experiments with this model are paused. The stateful-loop topology needs a return to the drawing board for flux2-klein-9b before it can go into the cross-model comparison. No conclusion yet on whether the model is suitable at all, or whether a different conditioning/topology is needed (flux2-klein-9b is a reference-image editor: no --strength, full 4-step regeneration).

2026-09-12: that topology redesign is now in the tree as dltb-klein + scripts/sweep-klein.sh (prompt-as-strength ladder, guidance probes). Later the same day the guidance-probe leg turned out to be inert for klein — see the next section.

Preview settings for reference: stateful, --reproject, --max-frames 30 --tail-frames 10 --tail-modes freeze.

FLUX.2 klein: --guidance-scale is inert (step-wise distilled) #

Found 2026-09-12, mid klein sweep: scripts/sweep-klein.sh reached its guidance probes (enhance-slight at --guidance-scale 2.0 / 4.0) and the log flooded with one Guidance scale 2.0 is ignored for step-wise distilled models. warning per pass. Verified against the deployed diffusers (0.40.0, diffusers/pipelines/flux2/pipeline_flux2_klein.py) — the value provably never reaches the model, via three independent points:

  1. check_inputs warns exactly when guidance_scale > 1.0 and self.config.is_distilled — klein checkpoints ship is_distilled: true.
  2. do_classifier_free_guidance is self._guidance_scale > 1 and not self.config.is_distilled — always False for klein, and the CFG branch (noise_pred + scale * (noise_pred - neg_noise_pred)) is the only consumer of guidance_scale in the pipeline.
  3. Unlike FLUX.1-dev there is no guidance-embedding fallback: the transformer is called with guidance=None unconditionally.

Consequences:

  • A --guidance-scale 2.0/4.0 run is bit-identical to the same-prompt default-guidance run (fixed seed) — the probes were duplicates of the prompt-enhance-slight run and measured nothing (~360 passes each).
  • --negative-prompt is equally inert (negative embeddings are only computed under CFG).
  • The sweep was killed mid-probe; the meaningful legs (prompt ladder, weathering attractor) were already on disk. (A partial guidance2.0/ duplicate of prompt-enhance-slight/ sat in the pod's output tree — moot since the pod's container disk is ephemeral.)

Follow-ups applied 2026-09-13: the guidance leg of sweep-klein.sh is replaced by a --num-inference-steps probe (STEPS="2 8", bracketing the default 4 that the ladder legs already run) — steps are the one remaining direct per-pass edit-intensity knob that actually reaches klein. dltb-klein now prints a one-time warning when --guidance-scale > 1 is passed (warn, not refuse, so a future diffusers that implements a real guidance path does not break the tool). The 2-frame hash A/B remains optional and only worth doing after a diffusers upgrade. Next-experiment sketch for klein's control problem: dual-reference conditioning, see the last section.

Klein dual-reference conditioning (--conditioning dual-ref) #

Status: IMPLEMENTED 2026-09-13. The candidate fix for klein's control problem, replacing the pixel-blend proxy. The design below is what landed (one deviation: continuous.run takes a make_conditioning factory that returns (combine, tail_source) function pairs, rather than a single condition(...) callable, because tails need their own source construction). Runs validating it (prompt ladder + order A/B) are still pending.

Why: klein's calibration trouble (too weak at blend 0.6–0.8, "melting" at 0.1 — see the calibration section) is plausibly an artifact of pixel-blending two frames into ONE reference image. Klein is trained as a (multi-)reference editor, and the pipeline natively accepts a LIST of reference images: verified in diffusers 0.40.0 Flux2KleinPipeline.__call__ (step 4 — each image is preprocessed, downscaled to ≤ 1 MP if needed, packed, and the packed latents are torch.cat([latents, image_latents], dim=1)-ed on the SEQUENCE axis; batch size comes from the prompt, not the image count). So the carried state and the fresh frame can both be conditioning inputs:

blend  (today) : P_n = f(image = (1-a)*R(P_{n-1}) + a*N_n)
dual-ref (new) : P_n = f(image = [R(P_{n-1}), N_n])    # two references

Design:

  • CLI in dltb-klein: --conditioning {blend,dual-ref} (default blend until validated). --anchor-blend applies to blend only; dual-ref run tag: <stem>_dualref[-norepro]_tails… (no blend component).
  • imaging.run_pass needs NO change — it forwards image=source, and a list of two PIL images flows straight through. --width/--height still set the output canvas; references are resized/packed per-image by the pipeline (our 768² frames are under the 1 MP auto-resize cap).
  • continuous.run was generalized (no fork): it accepts make_conditioning(args, estimate_flow, warp) -> (combine, tail_source), with the former blend/reproject logic as the default (_blend_conditioning); klein.py supplies _dual_ref_conditioning.
  • Reprojection: probably UNNECESSARY in dual-ref (the fresh frame is an explicit reference; the model aligns content, not pixel coordinates) — but keep it probeable: warping the carried reference may still help temporal stability. A/B --reproject / --no-reproject.
  • Reference order is a real variable: [P, N] vs [N, P] — likely encodes "primary vs target"; cheap 2-frame A/Bs answer it empirically.
  • Tails: freeze = [P, last_source]; free = [P] alone (single-reference regeneration from state — the pure buffer-echo case); black = [P, black].

Open questions / risks:

  • Token budget: each 768² reference packs to ~2.3k sequence tokens (2×2-packed VAE latents), so two references + text is a modest sequence — but measure the real VRAM/speed delta with the README_RUNPOD VRAM-probe pattern before scheduling runs.
  • Does klein weight multiple references equally, or is there an implicit "first = primary" convention? (The order A/B above answers this.)
  • Interaction with the steps probe: re-run the steps axis under dual-ref — intensity may interact with conditioning strength.

Suggested first runs:

uv run dltb-klein --model flux2-klein-4b --input untracked/input/video_cropped.mp4 \
    --conditioning dual-ref --prompt "slightly enhance the fine details" \
    --max-frames 30 --tail-frames 10 --save-every 1

# order A/B (2 frames each, compare): EXTRA_ARGS='--reproject' etc.
uv run dltb-klein --model flux2-klein-4b --input untracked/input/video_cropped.mp4 \
    --conditioning dual-ref --ref-order state-first --max-frames 2

Remaining follow-ups: none in-tree — the SMOKE_KLEIN=1 smoke leg, the reproject A/B (REPROJECT=1|0|ab, default ab under dual-ref), and REF_ORDER-aware role-naming prompts are all in scripts/smoke.sh / scripts/sweep-klein.sh. What is still pending is the GPU validation itself (prompt ladder + order A/B under dual-ref, then the reproject A/B).

Update, later the same day: the validation ran — see the next section. norepro freezes motion; the order/prompt axes are dead; reproject is load-bearing.

Klein dual-ref GPU validation: norepro freezes motion; reference gain measured (2026-09-13) #

Result: dual-ref WITHOUT reprojection cannot carry motion — a regime mismatch, not a bug. Reprojection is load-bearing. Found during the first sweep-klein.sh dual-ref run (stopped after the enhance-slight leg, so partial legs could be analyzed), then pinned down with 4b debug probes and a tail-based reference-gain readout.

Freeze ledger (all --mode stateful; motion stops within a few frames of the start — composition locks while texture keeps chattering):

  • 9b prompt-neutral …_dualref-norepro — freezes.
  • 9b prompt-enhance-slight …_dualref-norepro (role-naming prompt) — freezes.
  • 4b frame-first, empty prompt, norepro, 30 frames — freezes.
  • 4b frame-first, role-naming prompt, norepro, 30 frames — freezes.
  • 9b …_dualref (reproject ON) — motion continues. This also explains why early dual-ref results "looked like" blend mode: both carried motion via the warp.

Neither reference order nor a role-naming prompt rescues motion, so REF_ORDER is deprioritized as an axis (kept in the CLI for completeness). analyze_drift on the 4b frame-first pair: no fixed point — Δprev ≈ 9.6 (neutral) / 15.3 (prompt) MAD at save-every 5; Δoriginal 23.0 vs 35.7 — frozen composition + perpetual texture churn; the prompt escalates texture only (its mp4 is ~3× the neutral one's).

Not a pipeline-list bug. In diffusers 0.40.0 Flux2KleinPipeline each reference is separately preprocessed, VAE-encoded and packed, then concatenated on the sequence axis (nothing dropped; mu shifts with total reference token count via compute_empirical_mu). Key detail: _prepare_image_ids assigns reference i the time coordinate T = 10 + 10*i — the reference LIST is encoded as a temporally-indexed sequence of observations of ONE scene (restoration-style multi-frame conditioning), not "named subjects" an editor arbitrates between. The model card documents no multi-reference prompt convention, so role-naming phrasing is a guess — and it only modulates texture anyway.

Reference-gain readout (4b, 5 source frames then 10-frame tails freeze/free/black branching from ONE shared end state, empty prompt, fixed seed; pod untracked/output/debug/refweight). MAD between tail videos, t=1 → t=10:

frz-free   4.21 → 33.71   ([P, last_source] vs [P] alone)
frz-black  1.65 → 13.01   ([P, last_source] vs [P, black])
free-black 4.92 → 31.82

per-pass chatter (Δprev): freeze ≈ 2.5, black ≈ 3.0, free ≈ 4.0
mean luma t1→t10:         freeze 114→106, black 113→108, free 118→139

Reading: the second reference has real but LOW per-pass gain (~1.6 MAD vs a maximally different ref2 after one pass) that compounds over passes; ANY second reference — even black, whose content is not reproduced (no darkening: black-run luma tracks freeze-run) — stabilizes the consensus against buffer-echo drift ([P] alone drifts +21 luma in 10 passes and churns most). Motion death in the main loop follows: at 59.94 fps the inter-frame displacement is far below the frame reference's per-pass pull, so position is carried ONLY by the self-reinforcing state — and only the flow warp moves the state. Bonus finding: under dual-ref the black tail ≈ freeze tail (no decay driver) — the black scenario barely differs from freeze for klein.

Implications / follow-ups:

  • dual-ref + reproject is the viable klein video regime; the artifact ceiling is warp quality (grayscale Farneback flow, BORDER_REPLICATE smear, no disocclusion rejection — see imaging.make_reprojector). Next experiment: mask disocclusions via forward-backward flow consistency and patch them from the fresh frame before pairing.
  • The neutral sweep leg under dual-ref is effectively "freeze from frame ~5" — not a useful preservation baseline as-is.
  • The anchored control (--mode anchored) was not needed for this verdict; optional.

Klein dual-ref reproject-mask A/B (4b, 2026-09-13) #

Outcome: the disocclusion mask behaves as designed end-to-end. Run via scripts/sweep-klein-mask.sh (4b, 300 frames + 60-frame freeze tail, role-naming enhance-slight prompt): legs --reproject vs --reproject --reproject-mask, identical seed/steps/geometry.

  • Both legs keep motion (3rd replication: reproject is the motion carrier).
  • Mask firing on the actual clip (consecutive source pairs, run-identical constants): 1.9–7.3% of pixels per pass (mean ~4%), OOB 0.2–0.6% — fires at real motion; a flow-magnitude gate would change nothing. The carried state stays ~96% warped history per pass, so dual-ref semantics survive.
  • Same-seed legs DECORRELATE fully within ~10 passes (frame 1 identical → MAD 34 at frame 10 → plateau ~85–100; 75% of pixels differ >20): klein's per-pass texture redraw amplifies small conditioning deltas, so fixed-seed determinism holds only for identical conditioning. A/Bs of this loop must be statistical, not frame-identity.
  • Causal artifact metric (output-vs-source MAD inside vs outside the disocclusion mask, 30 saved frames): plain excess +3.5 (smears don't dominate pixel error), fbmask excess −47 (inside-mask output nearly matches source: 31 vs 78 MAD elsewhere) — patched fresh-frame content demonstrably propagates into the output; klein preserves the patch over stale state. 30/30 frames consistent.
  • Zone CHARACTER vs source (same frames, disocclusion zones): HF energy (|Laplacian|) plain 26.8 / fbmask 10.8 / source 5.5; saturation plain 146 / fbmask 105 / source 98. The unmasked streaks are high-frequency, oversaturated replicate smears (they read as "blurry" at video speed but are spectrally sharp); fbmask zones sit near source character, with a residual ~2× HF from the per-pass enhance prompt. All three metrics (pixel error, spectrum, color) favor the mask.

9b confirmation (2026-09-13, same script, MODEL=flux2-klein-9b, both legs complete): the 4b verdict replicates, with a twist that matches eyeball impressions ("streaks filled with 3D-looking shapes" on plain):

zones MAD vs src HF saturation
9b plain 58.5 21.8 215
9b fbmask 33.1 11.9 123
4b plain / fbmask 92.5 / 31.0 26.8 / 10.8 146 / 105
source — 5.5 98

The stronger editor doesn't produce raw smears — it resolves invalid history into plausible coherent structure (lower MAD than 4b plain) while cranking color/contrast (saturation 215 = 2.2× source): more convincing- LOOKING artifacts, still 4× source HF. The mask pins zones near source on BOTH models (fbmask rows are nearly model-independent — correct fresh content is preserved either way). Leg decorrelation as on 4b (t10 MAD 18, plateau ~50). fbmask cuts zone error −44% and halves the saturation excursion on 9b.

Neutral-prompt control & line concluded (2026-09-13; quantified from untracked/output/pod-20260913_debug-flux2_4). Two separable findings:

  • GLOBAL compounding drift is prompt-INDEPENDENT: analyze_drift delta_original (last-10 mean) neutral 88.5/89.0 (plain/fbmask) vs enhanced 78.4/92.0 — with an empty prompt the loop departs from the source just as far and the end frames look similar. The per-pass self-reference is the driver.
  • The ZONE vividness WAS prompt-driven: disocclusion-zone HF/saturation drop to near-source under the neutral prompt on BOTH legs (plain 21.8→8.6 / 215→92; fbmask 11.9→5.9 / 123→97; source 5.5/98) — the enhance prompt was rendering stale/smeared zones vividly. The mask still wins on every axis under the neutral prompt: zone MAD 20.9 vs 48.9 (−57%), and churn delta_prev 22.6 vs 52.5.

With that, the experiment line is concluded: dual-ref + reproject + mask is the best-characterized klein regime (mask verified near-source in disocclusion zones on both models, with and without prompt), but the fundamental per-pass compounding of a self-referencing regeneration loop remains unsolved. Follow-through (mask as dual-ref default, sweep axis) is SHELVED with the line — --reproject-mask stays opt-in; back to the drawing board (candidates: explicit history/confidence buffer outside the model, periodic hard re-anchoring, different conditioning entirely).

(Ops note: PROMPT is not in the run tag — the neutral sweep OVERWROTE the enhanced A/B outputs pod-side, same mask-ab/ tags; the enhanced data survives only in the earlier rsync. Differing-knob runs need OUT_PREFIX, same rule as REF_ORDER in sweep-klein.sh.)

GPU selection: re-check the whole Runpod catalog (action item, opened 2026-09-13) #

Runtime re-calibration (2026-09-13, RTX 6000 Ada, measured — scoped): the klein mask A/B ran NATIVELY (no --offload) with per leg 2 videos (processed_stateful + one freeze tail = 300 + 60 = 360 passes): 9b at 6.24 / 6.14 s/pass (fbmask/plain legs; whole sweep 74 min), 4b at ~3.2 s/pass. Scope when reusing these numbers: passes = frames summed over ALL videos a run emits (main + every tail mode), and note that the earlier 20–25 s/frame figure was measured WITH --offload on a 32 GB RTX 5090 (~10× native, see below) — NOT a native-vs-native comparison; plan against throughput-per-dollar using the matching mode.

Context: the 2026-09-13 klein-validation run is on an L40S (48 GB, Ada, sm_89). At commissioning time every 48 GB option showed Low stock on both clouds (runpodctl gpu list: RTX 6000 Ada, L40S, A6000, A40); A100 SXM was the only 80 GB card at Medium. An L40S was picked over a plain L40 despite the L40 being cheaper — see the reasoning below, which is exactly what this action item exists to verify rather than assume.

Reasoning to record (L40 vs L40S, and cheap-vs-fitting in general): raw $/hr is the wrong metric for this workload. A cheaper-but-slower 48 GB card is only a win if the price ratio beats the runtime ratio — compare price per pass (throughput per dollar), not price per hour. And either way, a slower 48 GB card that fits natively beats an OOM-ing faster card: --offload moves a whole pipeline component to the GPU per pipeline call and measured ~20–25 s per flux2-klein-9b frame on a 32 GB RTX 5090 (~10× native). That number is the floor for any "just rent a cheaper 32 GB card" argument.

Action: re-check every GPU rentable on Runpod (secure AND community) against its specs and rebuild the README_RUNPOD.md §1 table. For each candidate:

  • VRAM — native fit for all five models (48 GB remains the working recommendation; per-model peaks in the README table),
  • compute capability / sm_ version vs the CUDA 13 wheels in uv.lock (host driver ≥ 580; sm_89/90/120 verified supported),
  • $/hr secure vs community (stock changes hourly — re-run runpodctl gpu list and note the date),
  • throughput: record a real number where a run exists. Right now the tree has no measured L40S figure at all; the 2026-09-13 run should note frames/minute for flux2-klein-9b (plus sd-turbo/sdxl-turbo), so the next comparison works from data instead of a guess.

Rows the current README table is missing: L40 (48 GB, non-S — lower clocks/memory bandwidth than the L40S, often cheaper), and the non-48 GB classes already listed (32 GB consumer, 80 GB A100/H100) should stay in the table so the price/compute trade is explicit rather than implicit. The appendix VRAM probe is the cheap way to get peak VRAM; a fixed --max-frames run gives the throughput denominator.

Pod image retired: stock runpod/base + scripts/setup-pod.sh (2026-09-13) #

Decision: stop building refinementsystems/imgiter. Pods run the stock runpod/base:1.3.0-rc.164-ubuntu2404, pinned by the same digest the custom image was built FROM; scripts/setup-pod.sh (new, ships in every bundle) installs uv 0.12.13 into /usr/local/bin and runs uv sync --frozen.

Why: Runpod starts billing when the container image pull starts. The ~12 GB custom image was justified as skipping the multi-GB uv sync on boot, but it instead added a billed ~12 GB Docker Hub pull (often throttled) — the ~6 GB PyPI sync it skipped is cheaper, faster, and paid only when the lock actually changes. Secondary: every uv.lock change forced an emulated linux/amd64 rebuild + Docker Hub push + template digest re-pin; now lock changes ride the normal bundle workflow. For scale, both sides are dwarfed by the ~87.5 GB of HF model downloads every fresh pod pays anyway (container disk is wiped on stop and restart), so the whole optimization was noise.

Mechanics that replaced the baked venv:

  • Bundles extract into a fixed dir (/workspace/imgiter, --strip-components=1), so the project .venv (uv's default location, created by uv sync) and untracked/output/ survive new bundle extracts; re-running scripts/setup-pod.sh after each extract re-points the editable install (seconds when uv.lock is unchanged). This replaces the image's UV_PROJECT_ENVIRONMENT=/opt/imgiter/.venv env var, which cannot be provided container-wide without a custom image — and without it, stamp-dir extraction would re-download the stack once per bundle.
  • scripts/inputs.sh (sourced by every driver) now fails fast when .venv is missing, so uv run can never silently sync a fresh multi-GB venv mid-sweep. DRY_RUN=1 previews bypass the guard.
  • uv 0.12.13 is pinned to match the lockfile producer and the uv_build backend constraint (>=0.12.7,<0.13.0); installed from the GitHub release tarball (no curl | sh).

Not changed: the pod template (id 04u1mmp8nf) keeps its disk/env/ports; its image reference needs a one-time runpodctl template update --image re-point (README_RUNPOD.md §3). The old image tags stay on Docker Hub. hf-cache.sh's HF_HOME guard and the SSH/Jupyter behavior are base-image features, unaffected.

MPS (Apple Silicon) backend, local validation (2026-09-13) #

What already worked (no lockfile change needed). uv sync on this dev Mac installs torch 2.14.0 with a real MPS backend (is_available() == True); uv.lock already carries the macosx_14_0_arm64 wheels next to the Linux CUDA ones. torch.Generator("mps") constructs and seeds; fp16/bf16/fp32 matmuls run; the only CUDA hardcodes were models.load_pipeline() and imaging.make_generator().

Design. Device is a host property: models.resolve_device() picks CUDA -> MPS -> hard error (CPU only via explicit --device cpu), resolved once per tool and threaded into load_pipeline(spec, offload, device) and make_generator(..., device=...). ModelSpec stays device-free. New --device {cuda,mps,cpu} lives in args.add_output_args (so dltb-klein gets it too). On MPS load_pipeline enables attention slicing automatically. --offload now calls enable_model_cpu_offload(device=device) (device strings for CUDA are unchanged); MPS offload WORKS (verified) though it is slower than native (klein-4b 1 step: 2m29s offloaded vs ~75 s/step native). The five bash drivers' preflights accept CUDA or MPS (SKIP_GPU_CHECK=1 still bypasses); scripts/smoke-local.sh (new) runs one single-frame dltb-oneshot pass per locally-viable model with a non-uniformity check on the produced frame.

Measured on M1/16 GB, macOS 27.0, models cached (wall clock, attention slicing on, 1 denoise step where applicable):

model setting result
sd-turbo 512², default 4x0.4 = 1 step ~2 s/pass (10 passes in 28.6 s; 3 in 14.4 s)
sdxl-turbo 768², 1 step ~30-50 s/pass (10 passes in 8m7s; 3 in 1m54s; drifts with thermal/memory pressure)
flux2-klein-4b 768², 4 steps (no strength) warm single frame 5m18s (~75 s/step), ~11 GB swap touched

Checks that passed: scripts/smoke.sh PASSes locally on sd-turbo; two fixed-seed MPS runs produce identical frame hashes (determinism within one MPS build); analyze_drift.py on the local frames gives normal numbers (delta_prev ~9.4, delta_original ~20.7 over 3 frames). dltb-klein is the same loop (delegates to continuous.run), so it inherits --device.

Correction to the port plan: flux2-klein-4b DOES fit for single frames. It loads and generates on 16 GB unified memory, but it swaps hard (11 GB swap touched during 4 steps) and is far too slow for video; keep it in the single-frame smoke list, out of video runs. flux-schnell (~34 GB) and flux2-klein-9b (gated, ~20-29 GB) remain out of scope for local runs.

Caveats.

  • MPS is not bit-identical to CUDA (different kernels/reductions): never compare pixels across devices; analyze_drift.py is same-device only. On MPS, upcast_vae/fp32-VAE-upcast still works (SDXL prints the known benign diffusers false positive, see the log-noise section).
  • If an op errors with "not implemented for the mps backend", PYTORCH_ENABLE_MPS_FALLBACK=1 runs it on CPU silently (slow; escape hatch, not a default). No fp16-VAE black-frame artifacts were seen on any of the three models.
  • Local model cache is the default ~/.cache/huggingface; hf-cache.sh keep/clean still refuse to run without HF_HOME set — do not relax that guard just because the local cache now matters.
  • Local and pod outputs must not share a tree: run tags do not encode the device, so use --output-dir subtrees (untracked/output_<model> locally vs untracked/output/pod-* on the pod).
  • Fixed a pre-existing scripts/smoke.sh bug found during local validation: the dltb-iterate artifact check used frame_${ITERATIONS}.png but frames are written zero-padded (frame_0003.png), so the step could never pass; it now uses printf '%04d'. Also, python3 src/dltb/analyze_drift.py needs numpy/PIL, which the macOS system python does not have — use uv run python src/dltb/analyze_drift.py (docs updated).
  • Pod CUDA regression (bundle -> scripts/smoke.sh + --offload check) is still pending; nothing in the changed code path is CUDA-specific, but re-run it on the next pod session before trusting a cross-device comparison.

Watch: re-validate after torch/macOS upgrades (MPS op coverage and performance move quickly), and keep uv.lock's arm64 wheels in mind when bumping torch.

DreamSim perceptual metrics: dltb-distance (2026-09-14) #

Why: analyze_drift.py only sees pixels. A loop that has settled into a perceptual fixed point still chatters in pixel space (its delta_prev plateaus above the 0.5 threshold), while a loop can be pixel-stable and yet look nothing like its source. dltb-distance (src/dltb/analyze_distance.py) embeds frames with DreamSim and reports two distances per frame:

  • dreamsim_to_ref — to the run's source image (frame_0000_original.png / frame_0000_source.png, else the first frame): perceptual drift.
  • dreamsim_to_prev — to the previous analyzed frame: perceptual fixed-point detection. ~0 means the two frames are indistinguishable.

Validated on untracked/output/computer-enhance_sd-turbo/free-running/frames (200 frames, every 5th, ensemble): analyze_drift reported CONVERGED (delta_prev 0.119) with delta_original 97.3/255; DreamSim reported dreamsim_to_ref 0.789 (a stable image that is perceptually nothing like the source) and dreamsim_to_prev 0.0006, minimum 7.4e-05 at frame 0126. Both lenses agree on "settled" and differ, as designed, on "how far it went".

Facts verified for dreamsim 0.2.1 (do not re-derive):

  • uv add dreamsim adds exactly dreamsim, ftfy, open-clip-torch 3.3.0, peft 0.20.0, scipy 1.18.1, timm 1.0.29, torchvision 0.29.0, wcwidth — no existing pin changes (torch 2.14.0, transformers 5.17.0 stay). Linux-safe: only scipy/torchvision are compiled; the rest are py3-none-any wheels.
  • Side effect: torchvision is now in the lock — see log noise §3 above for why that is pixel-safe.
  • All six variants (ensemble, dino_vitb16, clip_vitb32, open_clip_vitb32, dinov2_vitb14, synclr_vitb16) load and score; --patch works for dino_vitb16 (verified) and dinov2_vitb14 (the only two the package ships patch checkpoints for).
  • MPS == CPU to float noise (~1e-7) on the same frames; a batch of N images is one model call (that is what --batch-size uses).
  • Identical images score exactly 0.0, but 1 - cos() undershoots on CPU (-2.4e-07 for a self-comparison), so the tool clamps at 0. Calibration: test_512.png vs contrast x1.15 + brightness x1.05 scores ~0.0086 (ensemble) — that is what the default --converged-below 0.01 is tuned to.
  • Weights come from GitHub releases via torch.hub, and that CDN intermittently answers HTTP 504 mid-transfer (open_clip_vitb32 and dinov2_vitb14 failed 1-3 attempts during validation, then succeeded). Hence --retries with 2/4/8 s backoff; a failed download leaves no partial file.
  • dreamsim's cache handling is crude in two ways: download_weights uses os.mkdir (single level, so the tool does mkdir -p itself), and every variant downloads to the same <cache_dir>/pretrained.zip while the cache check only looks at extracted checkpoints — so switching --dreamsim-type re-downloads the ~1.2 GB zip. Full ensemble cache: 3.8 GB.
  • Cache default is --cache-dir untracked/models (gitignored, cwd-relative). untracked/ is not packed into pod bundles (tracked files + untracked/input/ only), so a fresh pod re-downloads ~2.7 GB — that is why the smoke leg is gated behind SMOKE_DISTANCE=1 and runs on the frames step 2 already produced (no extra model pass).
  • Benign load noise, both on every load: the torch.nn.utils.weight_norm FutureWarning and peft's "Already found a peft_config attribute in the model. This will lead to having multiple adapters" UserWarning. Results are unaffected; do not add suppression without re-checking numbers.
  • Not encoded in the CSV/JSON names: --every, --dreamsim-type and --patch. Pass --out/--json when comparing variants, same rule as the run tags (AGENTS.md).

--mode anchored is the --anchor-blend 1.0 endpoint (2026-09-15) #

Analysis, then a small change. For the default pixel-blend conditioning,

blend((1-a)*R(P_{n-1}) + a*N_n)  --a=1.0-->  f(N_n)

because Image.blend(c, n, 1.0) = n exactly (uint8 -> float -> *1.0 -> round is the identity), so the carried state — and any optical-flow warp of it — contributes exactly zero. --mode stateful --anchor-blend 1.0 and --mode anchored produce pixel-identical frames, same seeds, same tails; dltb-distance-feedback's docstring already leaned on this endpoint as a correctness check (a=1.0 there = repeated independent passes = boil test).

The only cost difference was reprojection: stateful defaults to --reproject ON, and at a=1.0 the per-frame Farneback flow + warp was computed and then blended away — pure waste. continuous._blend_conditioning now skips flow/warp when anchor_blend >= 1.0, and run() prints a NOTE when that skip is active (the header still says reproject=on because the setting, not the work, drives the run tag). So stateful a=1.0 is now pixel- AND cost-identical to anchored.

Why keep the flag at all:

  • klein dual-ref. --conditioning dual-ref ignores --anchor-blend entirely (weighting is the model's job via attention over two clean references), so the blend endpoint does not exist there; --mode anchored is the only single-reference (boil-test) topology for klein.
  • Self-documenting CLI, distinct run tag (<stem>_anchored, processed_anchored.mp4 — smoke.sh asserts the latter), and it costs one or clause in the main loop.

Docs updated to say the flag is redundant except under klein dual-ref (continuous.py + klein.py docstrings/help, README, AGENTS.md gotchas).