# Notes Loose ends and follow-ups for imgiter. ## Pod SSH: sshd not running by default (2026-09-12; root cause found & solved 2026-09-12) **Symptom:** every non-interactive access path fails while the pod itself is healthy — `runpodctl ssh info` ip:port gives *connection refused*, `exec` python tunnels over that same mapping, and croc relays are independently flaky (see 2026-09-12 transfer session: two dead relays, one Go panic in runpodctl's croc client, then a 10-min receive timeout on a third code). Only the console **web terminal** and the `ssh -@ssh.runpod.io` gateway (PTY required — plain `ssh host cmd` is rejected with "Your SSH client doesn't support PTY"; scripted use needs a pty wrapper + fed commands) work. **Cause (corrected — the original "the pod image does not start sshd" was wrong):** the base image's `/start.sh` *does* set up sshd — host keys, `authorized_keys`, `service ssh start` — but its `setup_ssh` block is gated on `$PUBLIC_KEY`, which Runpod injects from the **account's registered SSH keys at pod start**. This pod had booted with none registered, so `PUBLIC_KEY` was empty and the whole block no-opped. Not an image defect; baking keys into the image was considered and rejected (README_DEFERRED.md). **Fix (permanent):** register a key once, then restart the pod. Every subsequent start brings sshd up automatically — nothing to redo manually (procedure: README_RUNPOD.md §3, "SSH access"): ```bash runpodctl ssh add-key --key-file ~/.ssh/id_ed25519.pub runpodctl ssh list-keys # confirm it's on the account # then pod stop && pod start — keys added after boot are ignored until restart ``` **Manual fallback** (still valid, for a pod you cannot restart mid-run; ephemeral — redo after every pod start): ```bash ssh-keygen -A # generates /etc/ssh/ssh_host_* keys service ssh start # "no hostkeys available -- exiting" without the line above ss -tlnp | grep :22 echo '' >> /root/.ssh/authorized_keys # else publickey auth fails ``` **Troubleshooting notes worth keeping:** - The **gateway** authenticates against the *Runpod account*; the in-container sshd against `/root/.ssh/authorized_keys` — two separate trust paths, each can break independently. A pod reachable via gateway but refused on the ip:port mapping means in-container sshd is down (this incident); the reverse (ip:port OK, gateway rejected) would point at the account side. - `ssh -@ssh.runpod.io` needs a PTY — scripted use requires a pty wrapper with fed commands; prefer the ip:port mapping for automation. - Host keys, authorized_keys and the running sshd live on the ephemeral container disk and vanish on stop/restart/recreate. With the account key registered, `/start.sh` regenerates all of it at every boot. - After a stop→start the external port is **reassigned**, and the first `runpodctl ssh info` can report a stale port for ~90 s — retry until a connection succeeds before assuming sshd is down. - With sshd reachable, `rsync -rtP --no-owner --no-group -e "ssh -i -p "` is the preferred bulk-transfer path (measured ~14 MB/s). `-r`/`-t` are load-bearing: plain `-P` on a directory prints `skipping directory .` and transfers nothing, and without `-t` the quick-check has no preserved mtime to compare. Verified 2026-09-13 (macOS openrsync, protocol 29 ↔ pod rsync 3.2.7, protocol 31): incremental re-runs skip unchanged files, `-P` resumes an interrupted transfer from the partial, and a file pulled while it was still being written was re-synced correctly by a later run (final sha256 match). Add `-c` to verify by checksum instead. **Key choice for the template's `PUBLIC_KEY` (2026-09-13):** the `imgiter` template sets `PUBLIC_KEY` explicitly (currently the `mm-mbp` key), and the explicit env value appears to replace Runpod's account-key injection — the 2026-09-13 pod's `/root/.ssh/authorized_keys` held only `mm-mbp`, not the other account-registered keys (`runpodctl-ssh-key`, `affine@lat5400`). The visible failure: `runpodctl ssh info` prints its suggested command with `-i ~/.runpod/ssh/runpodctl-ssh-key`, but that key gets *Permission denied*; direct SSH works only with `-i ~/.ssh/id_ed25519` (mm-mbp). **Next time set the template's `PUBLIC_KEY` to the public half of `~/.runpod/ssh/runpodctl-ssh-key`** (`~/.runpod/ssh/runpodctl-ssh-key.pub`, account-registered via `runpodctl ssh add-key` and cloud-synced), so the `ssh info` command works as printed. The gateway (`ssh -@ssh.runpod.io`) is unaffected either way — it accepts the key as an argument. **Jupyter Lab: auto-starts, but needs a published port (2026-09-13):** the base image's `/start.sh` starts Jupyter Lab on `0.0.0.0:8888` (preferred dir `/workspace`, login token = `$JUPYTER_PASSWORD`) whenever that variable is set — verified on the 2026-09-13 pod, so there is nothing to set up pod-side. What is missing is a mapping: the `imgiter` template publishes only `22/tcp`, Runpod proxies only ports declared at pod creation, and a running pod cannot gain one — so `https://-8888.proxy.runpod.net` is dead for this pod. Interim access is an SSH tunnel (verified working, login page HTTP 200): ```bash ssh -N -i ~/.ssh/id_ed25519 -p -L 8888:127.0.0.1:8888 root@ # then http://localhost:8888; password = echo "$JUPYTER_PASSWORD" on the pod ``` **Follow-up: add `8888/http` to the template's ports** (mark it *secure* in the console) so future pods can use the proxy URL directly; template edits apply only to newly created/recreated pods. The README §3 "Jupyter Lab (optional, untested)" heading is stale either way — it was exercised 2026-09-13. ## macOS `._*` AppleDouble files in pod bundles **Status: fixed** (2026-09-12) in `scripts/bundle.sh`; `.gitignore` updated. `scripts/image-build.sh` is **not** affected — it stages the same tree but never creates a tar archive. **Symptom:** extracting a bundle on a pod lists `._Justfile`, `._.gitignore`, `.___init__.py`, `._imgiter-`, etc. next to the real files. **Investigation** (all verified on the Mac against `bundle/imgiter-202609121612.tar.gz`): 1. Not committed: `git ls-files | grep -c '\._'` → `0`. 2. Not the staging step: after running bundle.sh's exact staging pipeline, `find "$stage" -name '._*'` → `0` files. 3. The staged files carry `com.apple.provenance`; `xattr -l` on a staged file shows it. macOS attaches this to files created during the extraction. 4. The **final `tar --no-xattrs -czf`** turns that xattr into AppleDouble members: Python `tarfile` counted **26 `._*` members out of 52**. 5. `--no-xattrs` does not prevent this. `COPYFILE_DISABLE=1` → `0` members; `--no-mac-metadata` → `0` members. 6. macOS `tar -tzf` **hides** `._*` members when listing, which is why the first "the archive is clean" check was wrong. **Impact:** inert on Linux (ignored by `uv`, Python imports, and git), but they bloat the archive and clutter extracted trees. They must not be committed. **Fix applied:** - `scripts/bundle.sh`: `export COPYFILE_DISABLE=1`, plus a post-build guard that verifies the finished archive with `python3`'s `tarfile`, and on failure deletes it and exits non-zero. The guard skips with a warning if `python3` is unavailable. - `.gitignore`: added `._*`. **Lessons:** - Never verify AppleDouble members on macOS with `tar -t`/`tar -tzf`; use Python `tarfile`. - Any future tar creation run on macOS needs `COPYFILE_DISABLE=1`. ## Runtime log noise (benign, optional cleanup) Observed during the first full sweep on the RTX 6000 Ada. None of these affect results; they are candidates for a cleanup pass on the next CLI/bundle change. Note that suppressing any of them requires a new bundle and a sweep restart. ### 1. `There are modules in AutoencoderKL that should be kept in float32: []` **Verdict: harmless diffusers false positive.** Fires roughly twice per decoded frame for the SD/SDXL models (thousands of lines per run); does not appear for the FLUX/FLUX-2 models. Mechanism: - `diffusers/models/modeling_utils.py` (~line 1523 in the installed version) has a buggy guard: ```python fp32_modules = self._keep_in_fp32_modules or [] # never None if dtype_present_in_args and fp32_modules is not None: # always true logger.warning(f"... should be kept in float32: {fp32_modules} ...") ``` `AutoencoderKL` does not define `_keep_in_fp32_modules`, so the message prints `[]` — nothing actually needs special handling. - The warnings come from the SD/SDXL pipeline's intentional VAE upcast around decode: `pipeline_stable_diffusion_xl_img2img.py` (~lines 1447–1475) does `self.upcast_vae()` (`vae.to(dtype=torch.float32)`) before `vae.decode(...)` and `self.vae.to(dtype=torch.float16)` after, because the VAE overflows in fp16. Both `.to(dtype=...)` calls hit the buggy guard. - `sd-turbo` and `sdxl-turbo` both have `"force_upcast": true` in their VAE `config.json`. Outputs were verified healthy (frame mean/stddev in normal ranges, no black or NaN frames). To silence later, add a targeted filter in `dltb/models.py` (inside `load_pipeline`, before the pipeline loads): ```python import logging class _DropFp32FalsePositive(logging.Filter): def filter(self, record): return "should be kept in float32" not in record.getMessage() logging.getLogger("diffusers.models.modeling_utils").addFilter(_DropFp32FalsePositive()) ``` ### 2. `torch.jit.script` is deprecated (FutureWarning) From diffusers internals (`torch/jit/_script.py` triggered inside the pipelines). Benign, no action planned beyond upstream updates. ### 3. `requires torchvision (not installed); falling back to CLIPImageProcessorPil` (resolved 2026-09-14) `torchvision` was not in the image, so transformers fell back to the PIL image processors (`CLIPImageProcessorPil`, `SiglipImageProcessorPil`). **Adding `dreamsim` pulled torchvision into `uv.lock` (2026-09-14)**, so the default `CLIPImageProcessor` / `Siglip2ImageProcessor` names now resolve to the torchvision-backed classes and the warning is gone. Verified safe for pixel output: in the pipelines dltb uses, `feature_extractor` is only called from `run_safety_checker` (dead here — `safety_checker: null` in the SD/SDXL repos) and from `encode_image` (dead — no IP-Adapter/image encoder). Input frames go through diffusers' own `VaeImageProcessor` / `Flux2ImageProcessor`, which never touch torchvision, and `flux2-klein` uses no transformers image processor at all. Re-check both dead branches before enabling a safety checker or an image encoder, and if that ever happens, compare deliberately against pre-2026-09-14 runs instead of mixing them. ### 4. `upcast_vae` deprecation The SDXL pipeline calls the deprecated `upcast_vae()` helper internally; if the `deprecate` line shows up it is internal diffusers churn, not our call site. ### 5. `Siglip2ImageProcessorFast is deprecated` Transformers deprecation of the `Fast` suffix on image processors, emitted while loading pipelines that use Siglip/Siglip2 encoders (the FLUX.2 klein models). Internal, benign; fixed by a transformers upgrade. ### 6. `hf_hub_download ... local_dir_use_symlinks is deprecated and ignored` `UserWarning` from `huggingface_hub/utils/_validators.py`. The argument is a no-op in current huggingface_hub; something further up the pipeline stack still passes it. Benign; disappears when that caller is updated. ### 7. `You have disabled the safety checker ... safety_checker=None` Printed once per SD/SDXL pipeline load: those model repos ship `safety_checker: null` in `model_index.json` and the CLI never requests one, so diffusers emits the license reminder. Benign, expected for `sd-turbo` and `sdxl-turbo`; it is not something the CLI can (or should) silence. ### 8. `Guidance scale 2.0 is ignored for step-wise distilled models.` **Not benign — it means the run is a no-op duplicate.** Emitted once per pass by `Flux2KleinPipeline.check_inputs` whenever `guidance_scale > 1.0` with a klein checkpoint (360 lines per probe run in the sweep log). The value is dropped on the floor: see [FLUX.2 klein: --guidance-scale is inert](#flux2-klein---guidance-scale-is-inert-step-wise-distilled) below. ## flux2-klein-9b anchor-blend calibration (superseded by the dltb-klein redesign) Context: for `sd-turbo` / `sdxl-turbo` / `flux-schnell`, the stateful sweep at `--anchor-blend 0.1/0.3/0.5` produced an effect judged too strong (a "cartoonish" degradation), so the re-sweep grid was set to `BLENDS="0.6 0.7 0.8"` (`BASELINE=0.7`) for every model. `flux2-klein-9b` did not follow that pattern: - preview at 0.6/0.7/0.8: **too weak** to be useful; - rerun at 0.1: a **noticeable effect but qualitatively different** — "melting" rather than the cartoonish degradation the other models show — and it drops off. Status: experiments with this model are paused. The stateful-loop topology needs a return to the drawing board for `flux2-klein-9b` before it can go into the cross-model comparison. No conclusion yet on whether the model is suitable at all, or whether a different conditioning/topology is needed (`flux2-klein-9b` is a reference-image editor: no `--strength`, full 4-step regeneration). 2026-09-12: that topology redesign is now in the tree as `dltb-klein` + `scripts/sweep-klein.sh` (prompt-as-strength ladder, guidance probes). Later the same day the guidance-probe leg turned out to be inert for klein — see the next section. Preview settings for reference: stateful, `--reproject`, `--max-frames 30 --tail-frames 10 --tail-modes freeze`. ## FLUX.2 klein: `--guidance-scale` is inert (step-wise distilled) **Found 2026-09-12, mid klein sweep:** `scripts/sweep-klein.sh` reached its guidance probes (`enhance-slight` at `--guidance-scale 2.0` / `4.0`) and the log flooded with one `Guidance scale 2.0 is ignored for step-wise distilled models.` warning per pass. Verified against the deployed diffusers (0.40.0, `diffusers/pipelines/flux2/pipeline_flux2_klein.py`) — the value provably never reaches the model, via three independent points: 1. `check_inputs` warns exactly when `guidance_scale > 1.0 and self.config.is_distilled` — klein checkpoints ship `is_distilled: true`. 2. `do_classifier_free_guidance` is `self._guidance_scale > 1 and not self.config.is_distilled` — always `False` for klein, and the CFG branch (`noise_pred + scale * (noise_pred - neg_noise_pred)`) is the **only** consumer of `guidance_scale` in the pipeline. 3. Unlike FLUX.1-dev there is no guidance-embedding fallback: the transformer is called with `guidance=None` unconditionally. **Consequences:** - A `--guidance-scale 2.0`/`4.0` run is bit-identical to the same-prompt default-guidance run (fixed seed) — the probes were duplicates of the `prompt-enhance-slight` run and measured nothing (~360 passes each). - `--negative-prompt` is equally inert (negative embeddings are only computed under CFG). - The sweep was killed mid-probe; the meaningful legs (prompt ladder, weathering attractor) were already on disk. (A partial `guidance2.0/` duplicate of `prompt-enhance-slight/` sat in the pod's output tree — moot since the pod's container disk is ephemeral.) **Follow-ups applied 2026-09-13:** the guidance leg of `sweep-klein.sh` is replaced by a `--num-inference-steps` probe (`STEPS="2 8"`, bracketing the default 4 that the ladder legs already run) — steps are the one remaining direct per-pass edit-intensity knob that actually reaches klein. `dltb-klein` now prints a one-time warning when `--guidance-scale > 1` is passed (warn, not refuse, so a future diffusers that implements a real guidance path does not break the tool). The 2-frame hash A/B remains optional and only worth doing after a diffusers upgrade. Next-experiment sketch for klein's control problem: dual-reference conditioning, see the last section. ## Klein dual-reference conditioning (`--conditioning dual-ref`) **Status: IMPLEMENTED 2026-09-13.** The candidate fix for klein's control problem, replacing the pixel-blend proxy. The design below is what landed (one deviation: `continuous.run` takes a `make_conditioning` factory that returns `(combine, tail_source)` function pairs, rather than a single `condition(...)` callable, because tails need their own source construction). Runs validating it (prompt ladder + order A/B) are still pending. **Why:** klein's calibration trouble (too weak at blend 0.6–0.8, "melting" at 0.1 — see the calibration section) is plausibly an artifact of pixel-blending two frames into ONE reference image. Klein is trained as a (multi-)reference editor, and the pipeline natively accepts a LIST of reference images: verified in diffusers 0.40.0 `Flux2KleinPipeline.__call__` (step 4 — each image is preprocessed, downscaled to ≤ 1 MP if needed, packed, and the packed latents are `torch.cat([latents, image_latents], dim=1)`-ed on the SEQUENCE axis; batch size comes from the prompt, not the image count). So the carried state and the fresh frame can both be conditioning inputs: blend (today) : P_n = f(image = (1-a)*R(P_{n-1}) + a*N_n) dual-ref (new) : P_n = f(image = [R(P_{n-1}), N_n]) # two references **Design:** - CLI in `dltb-klein`: `--conditioning {blend,dual-ref}` (default `blend` until validated). `--anchor-blend` applies to `blend` only; dual-ref run tag: `_dualref[-norepro]_tails…` (no blend component). - `imaging.run_pass` needs NO change — it forwards `image=source`, and a list of two PIL images flows straight through. `--width/--height` still set the output canvas; references are resized/packed per-image by the pipeline (our 768² frames are under the 1 MP auto-resize cap). - `continuous.run` was generalized (no fork): it accepts `make_conditioning(args, estimate_flow, warp) -> (combine, tail_source)`, with the former blend/reproject logic as the default (`_blend_conditioning`); `klein.py` supplies `_dual_ref_conditioning`. - Reprojection: probably UNNECESSARY in dual-ref (the fresh frame is an explicit reference; the model aligns content, not pixel coordinates) — but keep it probeable: warping the carried reference may still help temporal stability. A/B `--reproject` / `--no-reproject`. - Reference order is a real variable: `[P, N]` vs `[N, P]` — likely encodes "primary vs target"; cheap 2-frame A/Bs answer it empirically. - Tails: freeze = `[P, last_source]`; free = `[P]` alone (single-reference regeneration from state — the pure buffer-echo case); black = `[P, black]`. **Open questions / risks:** - Token budget: each 768² reference packs to ~2.3k sequence tokens (2×2-packed VAE latents), so two references + text is a modest sequence — but measure the real VRAM/speed delta with the README_RUNPOD VRAM-probe pattern before scheduling runs. - Does klein weight multiple references equally, or is there an implicit "first = primary" convention? (The order A/B above answers this.) - Interaction with the steps probe: re-run the steps axis under dual-ref — intensity may interact with conditioning strength. **Suggested first runs:** uv run dltb-klein --model flux2-klein-4b --input untracked/input/video_cropped.mp4 \ --conditioning dual-ref --prompt "slightly enhance the fine details" \ --max-frames 30 --tail-frames 10 --save-every 1 # order A/B (2 frames each, compare): EXTRA_ARGS='--reproject' etc. uv run dltb-klein --model flux2-klein-4b --input untracked/input/video_cropped.mp4 \ --conditioning dual-ref --ref-order state-first --max-frames 2 Remaining follow-ups: none in-tree — the `SMOKE_KLEIN=1` smoke leg, the reproject A/B (`REPROJECT=1|0|ab`, default `ab` under dual-ref), and REF_ORDER-aware role-naming prompts are all in `scripts/smoke.sh` / `scripts/sweep-klein.sh`. What is still pending is the GPU validation itself (prompt ladder + order A/B under dual-ref, then the reproject A/B). **Update, later the same day: the validation ran — see the next section. norepro freezes motion; the order/prompt axes are dead; reproject is load-bearing.** ## Klein dual-ref GPU validation: norepro freezes motion; reference gain measured (2026-09-13) **Result: dual-ref WITHOUT reprojection cannot carry motion — a regime mismatch, not a bug. Reprojection is load-bearing.** Found during the first `sweep-klein.sh` dual-ref run (stopped after the `enhance-slight` leg, so partial legs could be analyzed), then pinned down with 4b debug probes and a tail-based reference-gain readout. **Freeze ledger** (all `--mode stateful`; motion stops within a few frames of the start — composition locks while texture keeps chattering): - 9b `prompt-neutral` `…_dualref-norepro` — freezes. - 9b `prompt-enhance-slight` `…_dualref-norepro` (role-naming prompt) — freezes. - 4b frame-first, empty prompt, norepro, 30 frames — freezes. - 4b frame-first, role-naming prompt, norepro, 30 frames — freezes. - 9b `…_dualref` (reproject ON) — motion continues. This also explains why early dual-ref results "looked like" blend mode: both carried motion via the warp. Neither reference order nor a role-naming prompt rescues motion, so `REF_ORDER` is deprioritized as an axis (kept in the CLI for completeness). analyze_drift on the 4b frame-first pair: no fixed point — Δprev ≈ 9.6 (neutral) / 15.3 (prompt) MAD at save-every 5; Δoriginal 23.0 vs 35.7 — frozen composition + perpetual texture churn; the prompt escalates texture only (its mp4 is ~3× the neutral one's). **Not a pipeline-list bug.** In diffusers 0.40.0 `Flux2KleinPipeline` each reference is separately preprocessed, VAE-encoded and packed, then concatenated on the sequence axis (nothing dropped; `mu` shifts with total reference token count via `compute_empirical_mu`). Key detail: `_prepare_image_ids` assigns reference *i* the time coordinate T = 10 + 10*i — the reference LIST is encoded as a temporally-indexed sequence of observations of ONE scene (restoration-style multi-frame conditioning), not "named subjects" an editor arbitrates between. The model card documents no multi-reference prompt convention, so role-naming phrasing is a guess — and it only modulates texture anyway. **Reference-gain readout** (4b, 5 source frames then 10-frame tails freeze/free/black branching from ONE shared end state, empty prompt, fixed seed; pod `untracked/output/debug/refweight`). MAD between tail videos, t=1 → t=10: frz-free 4.21 → 33.71 ([P, last_source] vs [P] alone) frz-black 1.65 → 13.01 ([P, last_source] vs [P, black]) free-black 4.92 → 31.82 per-pass chatter (Δprev): freeze ≈ 2.5, black ≈ 3.0, free ≈ 4.0 mean luma t1→t10: freeze 114→106, black 113→108, free 118→139 Reading: the second reference has real but LOW per-pass gain (~1.6 MAD vs a maximally different ref2 after one pass) that compounds over passes; ANY second reference — even black, whose content is not reproduced (no darkening: black-run luma tracks freeze-run) — stabilizes the consensus against buffer-echo drift ([P] alone drifts +21 luma in 10 passes and churns most). Motion death in the main loop follows: at 59.94 fps the inter-frame displacement is far below the frame reference's per-pass pull, so position is carried ONLY by the self-reinforcing state — and only the flow warp moves the state. Bonus finding: under dual-ref the black tail ≈ freeze tail (no decay driver) — the black scenario barely differs from freeze for klein. **Implications / follow-ups:** - dual-ref + reproject is the viable klein video regime; the artifact ceiling is warp quality (grayscale Farneback flow, BORDER_REPLICATE smear, no disocclusion rejection — see `imaging.make_reprojector`). Next experiment: mask disocclusions via forward-backward flow consistency and patch them from the fresh frame before pairing. - The `neutral` sweep leg under dual-ref is effectively "freeze from frame ~5" — not a useful preservation baseline as-is. - The anchored control (`--mode anchored`) was not needed for this verdict; optional. ## Klein dual-ref reproject-mask A/B (4b, 2026-09-13) **Outcome: the disocclusion mask behaves as designed end-to-end.** Run via `scripts/sweep-klein-mask.sh` (4b, 300 frames + 60-frame freeze tail, role-naming enhance-slight prompt): legs `--reproject` vs `--reproject --reproject-mask`, identical seed/steps/geometry. - Both legs keep motion (3rd replication: reproject is the motion carrier). - Mask firing on the actual clip (consecutive source pairs, run-identical constants): 1.9–7.3% of pixels per pass (mean ~4%), OOB 0.2–0.6% — fires at real motion; a flow-magnitude gate would change nothing. The carried state stays ~96% warped history per pass, so dual-ref semantics survive. - Same-seed legs DECORRELATE fully within ~10 passes (frame 1 identical → MAD 34 at frame 10 → plateau ~85–100; 75% of pixels differ >20): klein's per-pass texture redraw amplifies small conditioning deltas, so fixed-seed determinism holds only for identical conditioning. A/Bs of this loop must be statistical, not frame-identity. - Causal artifact metric (output-vs-source MAD inside vs outside the disocclusion mask, 30 saved frames): plain excess +3.5 (smears don't dominate pixel error), fbmask excess −47 (inside-mask output nearly matches source: 31 vs 78 MAD elsewhere) — patched fresh-frame content demonstrably propagates into the output; klein preserves the patch over stale state. 30/30 frames consistent. - Zone CHARACTER vs source (same frames, disocclusion zones): HF energy (|Laplacian|) plain 26.8 / fbmask 10.8 / source 5.5; saturation plain 146 / fbmask 105 / source 98. The unmasked streaks are high-frequency, oversaturated replicate smears (they read as "blurry" at video speed but are spectrally sharp); fbmask zones sit near source character, with a residual ~2× HF from the per-pass enhance prompt. All three metrics (pixel error, spectrum, color) favor the mask. **9b confirmation (2026-09-13, same script, `MODEL=flux2-klein-9b`, both legs complete):** the 4b verdict replicates, with a twist that matches eyeball impressions ("streaks filled with 3D-looking shapes" on plain): | zones | MAD vs src | HF | saturation | |---|---|---|---| | 9b plain | 58.5 | 21.8 | **215** | | 9b fbmask | **33.1** | **11.9** | **123** | | 4b plain / fbmask | 92.5 / 31.0 | 26.8 / 10.8 | 146 / 105 | | source | — | 5.5 | 98 | The stronger editor doesn't produce raw smears — it *resolves* invalid history into plausible coherent structure (lower MAD than 4b plain) while cranking color/contrast (saturation 215 = 2.2× source): more convincing- LOOKING artifacts, still 4× source HF. The mask pins zones near source on BOTH models (fbmask rows are nearly model-independent — correct fresh content is preserved either way). Leg decorrelation as on 4b (t10 MAD 18, plateau ~50). fbmask cuts zone error −44% and halves the saturation excursion on 9b. **Neutral-prompt control & line concluded (2026-09-13; quantified from `untracked/output/pod-20260913_debug-flux2_4`).** Two separable findings: - GLOBAL compounding drift is prompt-INDEPENDENT: analyze_drift delta_original (last-10 mean) neutral 88.5/89.0 (plain/fbmask) vs enhanced 78.4/92.0 — with an empty prompt the loop departs from the source just as far and the end frames look similar. The per-pass self-reference is the driver. - The ZONE vividness WAS prompt-driven: disocclusion-zone HF/saturation drop to near-source under the neutral prompt on BOTH legs (plain 21.8→8.6 / 215→92; fbmask 11.9→5.9 / 123→97; source 5.5/98) — the enhance prompt was rendering stale/smeared zones vividly. The mask still wins on every axis under the neutral prompt: zone MAD 20.9 vs 48.9 (−57%), and churn delta_prev 22.6 vs 52.5. With that, the experiment line is concluded: dual-ref + reproject + mask is the best-characterized klein regime (mask verified near-source in disocclusion zones on both models, with and without prompt), but the fundamental per-pass compounding of a self-referencing regeneration loop remains unsolved. Follow-through (mask as dual-ref default, sweep axis) is SHELVED with the line — `--reproject-mask` stays opt-in; back to the drawing board (candidates: explicit history/confidence buffer outside the model, periodic hard re-anchoring, different conditioning entirely). (Ops note: PROMPT is not in the run tag — the neutral sweep OVERWROTE the enhanced A/B outputs pod-side, same `mask-ab/` tags; the enhanced data survives only in the earlier rsync. Differing-knob runs need OUT_PREFIX, same rule as REF_ORDER in `sweep-klein.sh`.) ## GPU selection: re-check the whole Runpod catalog (action item, opened 2026-09-13) **Runtime re-calibration (2026-09-13, RTX 6000 Ada, measured — scoped):** the klein mask A/B ran NATIVELY (no `--offload`) with per leg **2 videos** (`processed_stateful` + one freeze tail = 300 + 60 = 360 passes): 9b at **6.24 / 6.14 s/pass** (fbmask/plain legs; whole sweep 74 min), 4b at **~3.2 s/pass**. Scope when reusing these numbers: passes = frames summed over ALL videos a run emits (main + every tail mode), and note that the earlier 20–25 s/frame figure was measured WITH `--offload` on a 32 GB RTX 5090 (~10× native, see below) — NOT a native-vs-native comparison; plan against throughput-per-dollar using the matching mode. **Context:** the 2026-09-13 klein-validation run is on an **L40S** (48 GB, Ada, sm_89). At commissioning time every 48 GB option showed `Low` stock on both clouds (`runpodctl gpu list`: RTX 6000 Ada, L40S, A6000, A40); A100 SXM was the only 80 GB card at `Medium`. An L40S was picked over a plain **L40** despite the L40 being cheaper — see the reasoning below, which is exactly what this action item exists to verify rather than assume. **Reasoning to record (L40 vs L40S, and cheap-vs-fitting in general):** raw `$/hr` is the wrong metric for this workload. A cheaper-but-slower 48 GB card is only a win if the price ratio beats the runtime ratio — compare **price per pass (throughput per dollar)**, not price per hour. And either way, a slower 48 GB card that fits natively beats an OOM-ing faster card: `--offload` moves a whole pipeline component to the GPU per pipeline call and measured ~20–25 s per `flux2-klein-9b` frame on a 32 GB RTX 5090 (~10× native). That number is the floor for any "just rent a cheaper 32 GB card" argument. **Action: re-check every GPU rentable on Runpod (secure AND community) against its specs and rebuild the `README_RUNPOD.md` §1 table.** For each candidate: - VRAM — native fit for all five models (48 GB remains the working recommendation; per-model peaks in the README table), - compute capability / sm_ version vs the CUDA 13 wheels in `uv.lock` (host driver ≥ 580; sm_89/90/120 verified supported), - `$/hr` secure vs community (stock changes hourly — re-run `runpodctl gpu list` and note the date), - throughput: record a real number where a run exists. Right now the tree has no measured L40S figure at all; the 2026-09-13 run should note frames/minute for `flux2-klein-9b` (plus sd-turbo/sdxl-turbo), so the next comparison works from data instead of a guess. **Rows the current README table is missing:** **L40** (48 GB, non-S — lower clocks/memory bandwidth than the L40S, often cheaper), and the non-48 GB classes already listed (32 GB consumer, 80 GB A100/H100) should stay in the table so the price/compute trade is explicit rather than implicit. The appendix VRAM probe is the cheap way to get peak VRAM; a fixed `--max-frames` run gives the throughput denominator. ## Pod image retired: stock runpod/base + `scripts/setup-pod.sh` (2026-09-13) **Decision:** stop building `refinementsystems/imgiter`. Pods run the stock `runpod/base:1.3.0-rc.164-ubuntu2404`, pinned by the same digest the custom image was built `FROM`; `scripts/setup-pod.sh` (new, ships in every bundle) installs uv 0.12.13 into `/usr/local/bin` and runs `uv sync --frozen`. **Why:** Runpod starts billing when the container image pull starts. The ~12 GB custom image was justified as skipping the multi-GB `uv sync` on boot, but it instead added a billed ~12 GB Docker Hub pull (often throttled) — the ~6 GB PyPI sync it skipped is cheaper, faster, and paid only when the lock actually changes. Secondary: every `uv.lock` change forced an emulated linux/amd64 rebuild + Docker Hub push + template digest re-pin; now lock changes ride the normal bundle workflow. For scale, both sides are dwarfed by the ~87.5 GB of HF model downloads every fresh pod pays anyway (container disk is wiped on stop and restart), so the whole optimization was noise. **Mechanics that replaced the baked venv:** - Bundles extract into a **fixed dir** (`/workspace/imgiter`, `--strip-components=1`), so the project `.venv` (uv's default location, created by `uv sync`) and `untracked/output/` survive new bundle extracts; re-running `scripts/setup-pod.sh` after each extract re-points the editable install (seconds when `uv.lock` is unchanged). This replaces the image's `UV_PROJECT_ENVIRONMENT=/opt/imgiter/.venv` env var, which cannot be provided container-wide without a custom image — and without it, stamp-dir extraction would re-download the stack once per bundle. - `scripts/inputs.sh` (sourced by every driver) now fails fast when `.venv` is missing, so `uv run` can never silently sync a fresh multi-GB venv mid-sweep. `DRY_RUN=1` previews bypass the guard. - uv 0.12.13 is pinned to match the lockfile producer and the `uv_build` backend constraint (`>=0.12.7,<0.13.0`); installed from the GitHub release tarball (no `curl | sh`). **Not changed:** the pod template (id `04u1mmp8nf`) keeps its disk/env/ports; its image reference needs a one-time `runpodctl template update --image` re-point (README_RUNPOD.md §3). The old image tags stay on Docker Hub. `hf-cache.sh`'s `HF_HOME` guard and the SSH/Jupyter behavior are base-image features, unaffected. ## MPS (Apple Silicon) backend, local validation (2026-09-13) **What already worked (no lockfile change needed).** `uv sync` on this dev Mac installs torch 2.14.0 with a real MPS backend (`is_available() == True`); `uv.lock` already carries the `macosx_14_0_arm64` wheels next to the Linux CUDA ones. `torch.Generator("mps")` constructs and seeds; fp16/bf16/fp32 matmuls run; the only CUDA hardcodes were `models.load_pipeline()` and `imaging.make_generator()`. **Design.** Device is a host property: `models.resolve_device()` picks CUDA -> MPS -> hard error (CPU only via explicit `--device cpu`), resolved once per tool and threaded into `load_pipeline(spec, offload, device)` and `make_generator(..., device=...)`. `ModelSpec` stays device-free. New `--device {cuda,mps,cpu}` lives in `args.add_output_args` (so dltb-klein gets it too). On MPS `load_pipeline` enables attention slicing automatically. `--offload` now calls `enable_model_cpu_offload(device=device)` (device strings for CUDA are unchanged); MPS offload WORKS (verified) though it is slower than native (klein-4b 1 step: 2m29s offloaded vs ~75 s/step native). The five bash drivers' preflights accept CUDA or MPS (`SKIP_GPU_CHECK=1` still bypasses); `scripts/smoke-local.sh` (new) runs one single-frame `dltb-oneshot` pass per locally-viable model with a non-uniformity check on the produced frame. **Measured on M1/16 GB, macOS 27.0, models cached** (wall clock, attention slicing on, 1 denoise step where applicable): | model | setting | result | |---|---|---| | `sd-turbo` | 512², default 4x0.4 = 1 step | ~2 s/pass (10 passes in 28.6 s; 3 in 14.4 s) | | `sdxl-turbo` | 768², 1 step | ~30-50 s/pass (10 passes in 8m7s; 3 in 1m54s; drifts with thermal/memory pressure) | | `flux2-klein-4b` | 768², 4 steps (no strength) | warm single frame 5m18s (~75 s/step), ~11 GB swap touched | Checks that passed: `scripts/smoke.sh` PASSes locally on sd-turbo; two fixed-seed MPS runs produce identical frame hashes (determinism within one MPS build); `analyze_drift.py` on the local frames gives normal numbers (delta_prev ~9.4, delta_original ~20.7 over 3 frames). `dltb-klein` is the same loop (delegates to `continuous.run`), so it inherits `--device`. **Correction to the port plan: `flux2-klein-4b` DOES fit for single frames.** It loads and generates on 16 GB unified memory, but it swaps hard (11 GB swap touched during 4 steps) and is far too slow for video; keep it in the single-frame smoke list, out of video runs. `flux-schnell` (~34 GB) and `flux2-klein-9b` (gated, ~20-29 GB) remain out of scope for local runs. **Caveats.** - MPS is not bit-identical to CUDA (different kernels/reductions): never compare pixels across devices; `analyze_drift.py` is same-device only. On MPS, `upcast_vae`/fp32-VAE-upcast still works (SDXL prints the known benign diffusers false positive, see the log-noise section). - If an op errors with "not implemented for the mps backend", `PYTORCH_ENABLE_MPS_FALLBACK=1` runs it on CPU silently (slow; escape hatch, not a default). No fp16-VAE black-frame artifacts were seen on any of the three models. - Local model cache is the default `~/.cache/huggingface`; `hf-cache.sh keep/clean` still refuse to run without `HF_HOME` set — do not relax that guard just because the local cache now matters. - Local and pod outputs must not share a tree: run tags do not encode the device, so use `--output-dir` subtrees (`untracked/output_` locally vs `untracked/output/pod-*` on the pod). - Fixed a pre-existing `scripts/smoke.sh` bug found during local validation: the dltb-iterate artifact check used `frame_${ITERATIONS}.png` but frames are written zero-padded (`frame_0003.png`), so the step could never pass; it now uses `printf '%04d'`. Also, `python3 src/dltb/analyze_drift.py` needs numpy/PIL, which the macOS system python does not have — use `uv run python src/dltb/analyze_drift.py` (docs updated). - Pod CUDA regression (bundle -> `scripts/smoke.sh` + `--offload` check) is still pending; nothing in the changed code path is CUDA-specific, but re-run it on the next pod session before trusting a cross-device comparison. **Watch:** re-validate after torch/macOS upgrades (MPS op coverage and performance move quickly), and keep `uv.lock`'s arm64 wheels in mind when bumping torch. ## DreamSim perceptual metrics: `dltb-distance` (2026-09-14) **Why:** `analyze_drift.py` only sees pixels. A loop that has settled into a perceptual fixed point still chatters in pixel space (its `delta_prev` plateaus above the 0.5 threshold), while a loop can be pixel-stable and yet look nothing like its source. `dltb-distance` (`src/dltb/analyze_distance.py`) embeds frames with DreamSim and reports two distances per frame: - `dreamsim_to_ref` — to the run's source image (`frame_0000_original.png` / `frame_0000_source.png`, else the first frame): perceptual drift. - `dreamsim_to_prev` — to the previous analyzed frame: perceptual fixed-point detection. ~0 means the two frames are indistinguishable. Validated on `untracked/output/computer-enhance_sd-turbo/free-running/frames` (200 frames, every 5th, ensemble): `analyze_drift` reported CONVERGED (delta_prev 0.119) with delta_original 97.3/255; DreamSim reported `dreamsim_to_ref` 0.789 (a stable image that is perceptually nothing like the source) and `dreamsim_to_prev` 0.0006, minimum 7.4e-05 at frame 0126. Both lenses agree on "settled" and differ, as designed, on "how far it went". **Facts verified for dreamsim 0.2.1 (do not re-derive):** - `uv add dreamsim` adds exactly dreamsim, ftfy, open-clip-torch 3.3.0, peft 0.20.0, scipy 1.18.1, timm 1.0.29, torchvision 0.29.0, wcwidth — **no existing pin changes** (torch 2.14.0, transformers 5.17.0 stay). Linux-safe: only scipy/torchvision are compiled; the rest are `py3-none-any` wheels. - **Side effect: torchvision is now in the lock** — see log noise §3 above for why that is pixel-safe. - All six variants (`ensemble`, `dino_vitb16`, `clip_vitb32`, `open_clip_vitb32`, `dinov2_vitb14`, `synclr_vitb16`) load and score; `--patch` works for `dino_vitb16` (verified) and `dinov2_vitb14` (the only two the package ships patch checkpoints for). - **MPS == CPU to float noise (~1e-7)** on the same frames; a batch of N images is one model call (that is what `--batch-size` uses). - Identical images score exactly 0.0, but `1 - cos()` undershoots on CPU (-2.4e-07 for a self-comparison), so the tool clamps at 0. Calibration: `test_512.png` vs contrast x1.15 + brightness x1.05 scores ~0.0086 (ensemble) — that is what the default `--converged-below 0.01` is tuned to. - Weights come from GitHub releases via `torch.hub`, and that CDN intermittently answers **HTTP 504 mid-transfer** (`open_clip_vitb32` and `dinov2_vitb14` failed 1-3 attempts during validation, then succeeded). Hence `--retries` with 2/4/8 s backoff; a failed download leaves no partial file. - dreamsim's cache handling is crude in two ways: `download_weights` uses `os.mkdir` (**single level**, so the tool does `mkdir -p` itself), and every variant downloads to the same `/pretrained.zip` while the cache check only looks at extracted checkpoints — so switching `--dreamsim-type` re-downloads the ~1.2 GB zip. Full ensemble cache: 3.8 GB. - Cache default is `--cache-dir untracked/models` (gitignored, cwd-relative). `untracked/` is **not** packed into pod bundles (tracked files + `untracked/input/` only), so a fresh pod re-downloads ~2.7 GB — that is why the smoke leg is gated behind `SMOKE_DISTANCE=1` and runs on the frames step 2 already produced (no extra model pass). - Benign load noise, both on every load: the `torch.nn.utils.weight_norm` FutureWarning and peft's "Already found a `peft_config` attribute in the model. This will lead to having multiple adapters" UserWarning. Results are unaffected; do not add suppression without re-checking numbers. - Not encoded in the CSV/JSON names: `--every`, `--dreamsim-type` and `--patch`. Pass `--out`/`--json` when comparing variants, same rule as the run tags (AGENTS.md). ## `--mode anchored` is the `--anchor-blend 1.0` endpoint (2026-09-15) Analysis, then a small change. For the default pixel-blend conditioning, blend((1-a)*R(P_{n-1}) + a*N_n) --a=1.0--> f(N_n) because `Image.blend(c, n, 1.0) = n` exactly (uint8 -> float -> *1.0 -> round is the identity), so the carried state — and any optical-flow warp of it — contributes exactly zero. `--mode stateful --anchor-blend 1.0` and `--mode anchored` produce pixel-identical frames, same seeds, same tails; `dltb-distance-feedback`'s docstring already leaned on this endpoint as a correctness check (a=1.0 there = repeated independent passes = boil test). The only cost difference was reprojection: stateful defaults to `--reproject` ON, and at a=1.0 the per-frame Farneback flow + warp was computed and then blended away — pure waste. `continuous._blend_conditioning` now skips flow/warp when `anchor_blend >= 1.0`, and `run()` prints a NOTE when that skip is active (the header still says `reproject=on` because the setting, not the work, drives the run tag). So stateful a=1.0 is now pixel- AND cost-identical to anchored. Why keep the flag at all: - **klein dual-ref.** `--conditioning dual-ref` ignores `--anchor-blend` entirely (weighting is the model's job via attention over two clean references), so the blend endpoint does not exist there; `--mode anchored` is the only single-reference (boil-test) topology for klein. - Self-documenting CLI, distinct run tag (`_anchored`, `processed_anchored.mp4` — smoke.sh asserts the latter), and it costs one `or` clause in the main loop. Docs updated to say the flag is redundant except under klein dual-ref (continuous.py + klein.py docstrings/help, README, AGENTS.md gotchas).