Notes #
Loose ends and follow-ups for imgiter.
Pod SSH: sshd not running by default (2026-09-12; root cause found & solved 2026-09-12) #
Symptom: every non-interactive access path fails while the pod itself is
healthy — runpodctl ssh info ip:port gives connection refused, exec
python tunnels over that same mapping, and croc relays are independently
flaky (see 2026-09-12 transfer session: two dead relays, one Go panic in
runpodctl's croc client, then a 10-min receive timeout on a third code).
Only the console web terminal and the ssh <pod-id>-<token>@ssh.runpod.io
gateway (PTY required — plain ssh host cmd is rejected with "Your SSH client
doesn't support PTY"; scripted use needs a pty wrapper + fed commands) work.
Cause (corrected — the original "the pod image does not start sshd" was
wrong): the base image's /start.sh does set up sshd — host keys,
authorized_keys, service ssh start — but its setup_ssh block is gated
on $PUBLIC_KEY, which Runpod injects from the account's registered SSH
keys at pod start. This pod had booted with none registered, so PUBLIC_KEY
was empty and the whole block no-opped. Not an image defect; baking keys
into the image was considered and rejected (README_DEFERRED.md).
Fix (permanent): register a key once, then restart the pod. Every subsequent start brings sshd up automatically — nothing to redo manually (procedure: README_RUNPOD.md §3, "SSH access"):
runpodctl ssh add-key --key-file ~/.ssh/id_ed25519.pub
runpodctl ssh list-keys # confirm it's on the account
# then pod stop && pod start — keys added after boot are ignored until restart
Manual fallback (still valid, for a pod you cannot restart mid-run; ephemeral — redo after every pod start):
ssh-keygen -A # generates /etc/ssh/ssh_host_* keys
service ssh start # "no hostkeys available -- exiting" without the line above
ss -tlnp | grep :22
echo '<pubkey line>' >> /root/.ssh/authorized_keys # else publickey auth fails
Troubleshooting notes worth keeping:
- The gateway authenticates against the Runpod account; the in-container
sshd against
/root/.ssh/authorized_keys— two separate trust paths, each can break independently. A pod reachable via gateway but refused on the ip:port mapping means in-container sshd is down (this incident); the reverse (ip:port OK, gateway rejected) would point at the account side. ssh <pod-id>-<token>@ssh.runpod.ioneeds a PTY — scripted use requires a pty wrapper with fed commands; prefer the ip:port mapping for automation.- Host keys, authorized_keys and the running sshd live on the ephemeral
container disk and vanish on stop/restart/recreate. With the account key
registered,
/start.shregenerates all of it at every boot. - After a stop→start the external port is reassigned, and the first
runpodctl ssh infocan report a stale port for ~90 s — retry until a connection succeeds before assuming sshd is down. - With sshd reachable,
rsync -rtP --no-owner --no-group -e "ssh -i <key> -p <port>"is the preferred bulk-transfer path (measured ~14 MB/s).-r/-tare load-bearing: plain-Pon a directory printsskipping directory .and transfers nothing, and without-tthe quick-check has no preserved mtime to compare. Verified 2026-09-13 (macOS openrsync, protocol 29 ↔ pod rsync 3.2.7, protocol 31): incremental re-runs skip unchanged files,-Presumes an interrupted transfer from the partial, and a file pulled while it was still being written was re-synced correctly by a later run (final sha256 match). Add-cto verify by checksum instead.
Key choice for the template's PUBLIC_KEY (2026-09-13): the imgiter
template sets PUBLIC_KEY explicitly (currently the mm-mbp key), and the
explicit env value appears to replace Runpod's account-key injection — the
2026-09-13 pod's /root/.ssh/authorized_keys held only mm-mbp, not the
other account-registered keys (runpodctl-ssh-key, affine@lat5400). The
visible failure: runpodctl ssh info prints its suggested command with
-i ~/.runpod/ssh/runpodctl-ssh-key, but that key gets Permission denied;
direct SSH works only with -i ~/.ssh/id_ed25519 (mm-mbp). Next time set
the template's PUBLIC_KEY to the public half of
~/.runpod/ssh/runpodctl-ssh-key (~/.runpod/ssh/runpodctl-ssh-key.pub,
account-registered via runpodctl ssh add-key and cloud-synced), so the
ssh info command works as printed. The gateway
(ssh <pod-id>-<token>@ssh.runpod.io) is unaffected either way — it accepts
the key as an argument.
Jupyter Lab: auto-starts, but needs a published port (2026-09-13): the
base image's /start.sh starts Jupyter Lab on 0.0.0.0:8888 (preferred dir
/workspace, login token = $JUPYTER_PASSWORD) whenever that variable is
set — verified on the 2026-09-13 pod, so there is nothing to set up pod-side.
What is missing is a mapping: the imgiter template publishes only 22/tcp,
Runpod proxies only ports declared at pod creation, and a running pod cannot
gain one — so https://<pod-id>-8888.proxy.runpod.net is dead for this pod.
Interim access is an SSH tunnel (verified working, login page HTTP 200):
ssh -N -i ~/.ssh/id_ed25519 -p <port> -L 8888:127.0.0.1:8888 root@<ip>
# then http://localhost:8888; password = echo "$JUPYTER_PASSWORD" on the pod
Follow-up: add 8888/http to the template's ports (mark it secure in
the console) so future pods can use the proxy URL directly; template edits
apply only to newly created/recreated pods. The README §3 "Jupyter Lab
(optional, untested)" heading is stale either way — it was exercised
2026-09-13.
macOS ._* AppleDouble files in pod bundles #
Status: fixed (2026-09-12) in scripts/bundle.sh; .gitignore updated.
scripts/image-build.sh is not affected — it stages the same tree but never
creates a tar archive.
Symptom: extracting a bundle on a pod lists ._Justfile, ._.gitignore,
.___init__.py, ._imgiter-<stamp>, etc. next to the real files.
Investigation (all verified on the Mac against
bundle/imgiter-202609121612.tar.gz):
- Not committed:
git ls-files | grep -c '\._'→0. - Not the staging step: after running bundle.sh's exact staging pipeline,
find "$stage" -name '._*'→0files. - The staged files carry
com.apple.provenance;xattr -lon a staged file shows it. macOS attaches this to files created during the extraction. - The final
tar --no-xattrs -czfturns that xattr into AppleDouble members: Pythontarfilecounted 26._*members out of 52. --no-xattrsdoes not prevent this.COPYFILE_DISABLE=1→0members;--no-mac-metadata→0members.- macOS
tar -tzfhides._*members when listing, which is why the first "the archive is clean" check was wrong.
Impact: inert on Linux (ignored by uv, Python imports, and git), but they
bloat the archive and clutter extracted trees. They must not be committed.
Fix applied:
scripts/bundle.sh:export COPYFILE_DISABLE=1, plus a post-build guard that verifies the finished archive withpython3'starfile, and on failure deletes it and exits non-zero. The guard skips with a warning ifpython3is unavailable..gitignore: added._*.
Lessons:
- Never verify AppleDouble members on macOS with
tar -t/tar -tzf; use Pythontarfile. - Any future tar creation run on macOS needs
COPYFILE_DISABLE=1.
Runtime log noise (benign, optional cleanup) #
Observed during the first full sweep on the RTX 6000 Ada. None of these affect results; they are candidates for a cleanup pass on the next CLI/bundle change. Note that suppressing any of them requires a new bundle and a sweep restart.
1. There are modules in AutoencoderKL that should be kept in float32: [] #
Verdict: harmless diffusers false positive. Fires roughly twice per decoded frame for the SD/SDXL models (thousands of lines per run); does not appear for the FLUX/FLUX-2 models.
Mechanism:
diffusers/models/modeling_utils.py(~line 1523 in the installed version) has a buggy guard:fp32_modules = self._keep_in_fp32_modules or [] # never None if dtype_present_in_args and fp32_modules is not None: # always true logger.warning(f"... should be kept in float32: {fp32_modules} ...")AutoencoderKLdoes not define_keep_in_fp32_modules, so the message prints[]— nothing actually needs special handling.- The warnings come from the SD/SDXL pipeline's intentional VAE upcast around
decode:
pipeline_stable_diffusion_xl_img2img.py(~lines 1447–1475) doesself.upcast_vae()(vae.to(dtype=torch.float32)) beforevae.decode(...)andself.vae.to(dtype=torch.float16)after, because the VAE overflows in fp16. Both.to(dtype=...)calls hit the buggy guard. sd-turboandsdxl-turboboth have"force_upcast": truein their VAEconfig.json.
Outputs were verified healthy (frame mean/stddev in normal ranges, no black or
NaN frames). To silence later, add a targeted filter in dltb/models.py
(inside load_pipeline, before the pipeline loads):
import logging
class _DropFp32FalsePositive(logging.Filter):
def filter(self, record):
return "should be kept in float32" not in record.getMessage()
logging.getLogger("diffusers.models.modeling_utils").addFilter(_DropFp32FalsePositive())
2. torch.jit.script is deprecated (FutureWarning) #
From diffusers internals (torch/jit/_script.py triggered inside the
pipelines). Benign, no action planned beyond upstream updates.
3. requires torchvision (not installed); falling back to CLIPImageProcessorPil (resolved 2026-09-14) #
torchvision was not in the image, so transformers fell back to the PIL image
processors (CLIPImageProcessorPil, SiglipImageProcessorPil). Adding
dreamsim pulled torchvision into uv.lock (2026-09-14), so the default
CLIPImageProcessor / Siglip2ImageProcessor names now resolve to the
torchvision-backed classes and the warning is gone.
Verified safe for pixel output: in the pipelines dltb uses, feature_extractor
is only called from run_safety_checker (dead here — safety_checker: null
in the SD/SDXL repos) and from encode_image (dead — no IP-Adapter/image
encoder). Input frames go through diffusers' own VaeImageProcessor /
Flux2ImageProcessor, which never touch torchvision, and flux2-klein uses
no transformers image processor at all. Re-check both dead branches before
enabling a safety checker or an image encoder, and if that ever happens,
compare deliberately against pre-2026-09-14 runs instead of mixing them.
4. upcast_vae deprecation #
The SDXL pipeline calls the deprecated upcast_vae() helper internally; if the
deprecate line shows up it is internal diffusers churn, not our call site.
5. Siglip2ImageProcessorFast is deprecated #
Transformers deprecation of the Fast suffix on image processors, emitted
while loading pipelines that use Siglip/Siglip2 encoders (the FLUX.2 klein
models). Internal, benign; fixed by a transformers upgrade.
6. hf_hub_download ... local_dir_use_symlinks is deprecated and ignored #
UserWarning from huggingface_hub/utils/_validators.py. The argument is a
no-op in current huggingface_hub; something further up the pipeline stack still
passes it. Benign; disappears when that caller is updated.
7. You have disabled the safety checker ... safety_checker=None #
Printed once per SD/SDXL pipeline load: those model repos ship
safety_checker: null in model_index.json and the CLI never requests one, so
diffusers emits the license reminder. Benign, expected for sd-turbo and
sdxl-turbo; it is not something the CLI can (or should) silence.
8. Guidance scale 2.0 is ignored for step-wise distilled models. #
Not benign — it means the run is a no-op duplicate. Emitted once per pass
by Flux2KleinPipeline.check_inputs whenever guidance_scale > 1.0 with a
klein checkpoint (360 lines per probe run in the sweep log). The value is
dropped on the floor: see
FLUX.2 klein: --guidance-scale is inert
below.
flux2-klein-9b anchor-blend calibration (superseded by the dltb-klein redesign) #
Context: for sd-turbo / sdxl-turbo / flux-schnell, the stateful sweep at
--anchor-blend 0.1/0.3/0.5 produced an effect judged too strong (a
"cartoonish" degradation), so the re-sweep grid was set to
BLENDS="0.6 0.7 0.8" (BASELINE=0.7) for every model.
flux2-klein-9b did not follow that pattern:
- preview at 0.6/0.7/0.8: too weak to be useful;
- rerun at 0.1: a noticeable effect but qualitatively different — "melting" rather than the cartoonish degradation the other models show — and it drops off.
Status: experiments with this model are paused. The stateful-loop topology needs
a return to the drawing board for flux2-klein-9b before it can go into the
cross-model comparison. No conclusion yet on whether the model is suitable at
all, or whether a different conditioning/topology is needed (flux2-klein-9b
is a reference-image editor: no --strength, full 4-step regeneration).
2026-09-12: that topology redesign is now in the tree as dltb-klein +
scripts/sweep-klein.sh (prompt-as-strength ladder, guidance probes).
Later the same day the guidance-probe leg turned out to be inert for klein —
see the next section.
Preview settings for reference: stateful, --reproject,
--max-frames 30 --tail-frames 10 --tail-modes freeze.
FLUX.2 klein: --guidance-scale is inert (step-wise distilled) #
Found 2026-09-12, mid klein sweep: scripts/sweep-klein.sh reached its
guidance probes (enhance-slight at --guidance-scale 2.0 / 4.0) and the
log flooded with one Guidance scale 2.0 is ignored for step-wise distilled models. warning per pass. Verified against the deployed diffusers (0.40.0,
diffusers/pipelines/flux2/pipeline_flux2_klein.py) — the value provably
never reaches the model, via three independent points:
check_inputswarns exactly whenguidance_scale > 1.0 and self.config.is_distilled— klein checkpoints shipis_distilled: true.do_classifier_free_guidanceisself._guidance_scale > 1 and not self.config.is_distilled— alwaysFalsefor klein, and the CFG branch (noise_pred + scale * (noise_pred - neg_noise_pred)) is the only consumer ofguidance_scalein the pipeline.- Unlike FLUX.1-dev there is no guidance-embedding fallback: the transformer
is called with
guidance=Noneunconditionally.
Consequences:
- A
--guidance-scale 2.0/4.0run is bit-identical to the same-prompt default-guidance run (fixed seed) — the probes were duplicates of theprompt-enhance-slightrun and measured nothing (~360 passes each). --negative-promptis equally inert (negative embeddings are only computed under CFG).- The sweep was killed mid-probe; the meaningful legs (prompt ladder,
weathering attractor) were already on disk. (A partial
guidance2.0/duplicate ofprompt-enhance-slight/sat in the pod's output tree — moot since the pod's container disk is ephemeral.)
Follow-ups applied 2026-09-13: the guidance leg of sweep-klein.sh is
replaced by a --num-inference-steps probe (STEPS="2 8", bracketing the
default 4 that the ladder legs already run) — steps are the one remaining
direct per-pass edit-intensity knob that actually reaches klein. dltb-klein
now prints a one-time warning when --guidance-scale > 1 is passed (warn,
not refuse, so a future diffusers that implements a real guidance path does
not break the tool). The 2-frame hash A/B remains optional and only worth
doing after a diffusers upgrade. Next-experiment sketch for klein's control
problem: dual-reference conditioning, see the last section.
Klein dual-reference conditioning (--conditioning dual-ref) #
Status: IMPLEMENTED 2026-09-13. The candidate fix for klein's control
problem, replacing the pixel-blend proxy. The design below is what landed
(one deviation: continuous.run takes a make_conditioning factory that
returns (combine, tail_source) function pairs, rather than a single
condition(...) callable, because tails need their own source construction).
Runs validating it (prompt ladder + order A/B) are still pending.
Why: klein's calibration trouble (too weak at blend 0.6–0.8, "melting" at
0.1 — see the calibration section) is plausibly an artifact of pixel-blending
two frames into ONE reference image. Klein is trained as a (multi-)reference
editor, and the pipeline natively accepts a LIST of reference images:
verified in diffusers 0.40.0 Flux2KleinPipeline.__call__ (step 4 — each
image is preprocessed, downscaled to ≤ 1 MP if needed, packed, and the packed
latents are torch.cat([latents, image_latents], dim=1)-ed on the SEQUENCE
axis; batch size comes from the prompt, not the image count). So the carried
state and the fresh frame can both be conditioning inputs:
blend (today) : P_n = f(image = (1-a)*R(P_{n-1}) + a*N_n)
dual-ref (new) : P_n = f(image = [R(P_{n-1}), N_n]) # two references
Design:
- CLI in
dltb-klein:--conditioning {blend,dual-ref}(defaultblenduntil validated).--anchor-blendapplies toblendonly; dual-ref run tag:<stem>_dualref[-norepro]_tails…(no blend component). imaging.run_passneeds NO change — it forwardsimage=source, and a list of two PIL images flows straight through.--width/--heightstill set the output canvas; references are resized/packed per-image by the pipeline (our 768² frames are under the 1 MP auto-resize cap).continuous.runwas generalized (no fork): it acceptsmake_conditioning(args, estimate_flow, warp) -> (combine, tail_source), with the former blend/reproject logic as the default (_blend_conditioning);klein.pysupplies_dual_ref_conditioning.- Reprojection: probably UNNECESSARY in dual-ref (the fresh frame is an
explicit reference; the model aligns content, not pixel coordinates) — but
keep it probeable: warping the carried reference may still help temporal
stability. A/B
--reproject/--no-reproject. - Reference order is a real variable:
[P, N]vs[N, P]— likely encodes "primary vs target"; cheap 2-frame A/Bs answer it empirically. - Tails: freeze =
[P, last_source]; free =[P]alone (single-reference regeneration from state — the pure buffer-echo case); black =[P, black].
Open questions / risks:
- Token budget: each 768² reference packs to ~2.3k sequence tokens (2×2-packed VAE latents), so two references + text is a modest sequence — but measure the real VRAM/speed delta with the README_RUNPOD VRAM-probe pattern before scheduling runs.
- Does klein weight multiple references equally, or is there an implicit "first = primary" convention? (The order A/B above answers this.)
- Interaction with the steps probe: re-run the steps axis under dual-ref — intensity may interact with conditioning strength.
Suggested first runs:
uv run dltb-klein --model flux2-klein-4b --input untracked/input/video_cropped.mp4 \
--conditioning dual-ref --prompt "slightly enhance the fine details" \
--max-frames 30 --tail-frames 10 --save-every 1
# order A/B (2 frames each, compare): EXTRA_ARGS='--reproject' etc.
uv run dltb-klein --model flux2-klein-4b --input untracked/input/video_cropped.mp4 \
--conditioning dual-ref --ref-order state-first --max-frames 2
Remaining follow-ups: none in-tree — the SMOKE_KLEIN=1 smoke leg, the
reproject A/B (REPROJECT=1|0|ab, default ab under dual-ref), and
REF_ORDER-aware role-naming prompts are all in scripts/smoke.sh /
scripts/sweep-klein.sh. What is still pending is the GPU validation itself
(prompt ladder + order A/B under dual-ref, then the reproject A/B).
Update, later the same day: the validation ran — see the next section. norepro freezes motion; the order/prompt axes are dead; reproject is load-bearing.
Klein dual-ref GPU validation: norepro freezes motion; reference gain measured (2026-09-13) #
Result: dual-ref WITHOUT reprojection cannot carry motion — a regime
mismatch, not a bug. Reprojection is load-bearing. Found during the first
sweep-klein.sh dual-ref run (stopped after the enhance-slight leg, so
partial legs could be analyzed), then pinned down with 4b debug probes and a
tail-based reference-gain readout.
Freeze ledger (all --mode stateful; motion stops within a few frames
of the start — composition locks while texture keeps chattering):
- 9b
prompt-neutral…_dualref-norepro— freezes. - 9b
prompt-enhance-slight…_dualref-norepro(role-naming prompt) — freezes. - 4b frame-first, empty prompt, norepro, 30 frames — freezes.
- 4b frame-first, role-naming prompt, norepro, 30 frames — freezes.
- 9b
…_dualref(reproject ON) — motion continues. This also explains why early dual-ref results "looked like" blend mode: both carried motion via the warp.
Neither reference order nor a role-naming prompt rescues motion, so
REF_ORDER is deprioritized as an axis (kept in the CLI for completeness).
analyze_drift on the 4b frame-first pair: no fixed point — Δprev ≈ 9.6
(neutral) / 15.3 (prompt) MAD at save-every 5; Δoriginal 23.0 vs 35.7 —
frozen composition + perpetual texture churn; the prompt escalates texture
only (its mp4 is ~3× the neutral one's).
Not a pipeline-list bug. In diffusers 0.40.0 Flux2KleinPipeline each
reference is separately preprocessed, VAE-encoded and packed, then
concatenated on the sequence axis (nothing dropped; mu shifts with total
reference token count via compute_empirical_mu). Key detail:
_prepare_image_ids assigns reference i the time coordinate T = 10 + 10*i
— the reference LIST is encoded as a temporally-indexed sequence of
observations of ONE scene (restoration-style multi-frame conditioning), not
"named subjects" an editor arbitrates between. The model card documents no
multi-reference prompt convention, so role-naming phrasing is a guess — and
it only modulates texture anyway.
Reference-gain readout (4b, 5 source frames then 10-frame tails
freeze/free/black branching from ONE shared end state, empty prompt, fixed
seed; pod untracked/output/debug/refweight). MAD between tail videos, t=1 → t=10:
frz-free 4.21 → 33.71 ([P, last_source] vs [P] alone)
frz-black 1.65 → 13.01 ([P, last_source] vs [P, black])
free-black 4.92 → 31.82
per-pass chatter (Δprev): freeze ≈ 2.5, black ≈ 3.0, free ≈ 4.0
mean luma t1→t10: freeze 114→106, black 113→108, free 118→139
Reading: the second reference has real but LOW per-pass gain (~1.6 MAD vs a maximally different ref2 after one pass) that compounds over passes; ANY second reference — even black, whose content is not reproduced (no darkening: black-run luma tracks freeze-run) — stabilizes the consensus against buffer-echo drift ([P] alone drifts +21 luma in 10 passes and churns most). Motion death in the main loop follows: at 59.94 fps the inter-frame displacement is far below the frame reference's per-pass pull, so position is carried ONLY by the self-reinforcing state — and only the flow warp moves the state. Bonus finding: under dual-ref the black tail ≈ freeze tail (no decay driver) — the black scenario barely differs from freeze for klein.
Implications / follow-ups:
- dual-ref + reproject is the viable klein video regime; the artifact
ceiling is warp quality (grayscale Farneback flow, BORDER_REPLICATE smear,
no disocclusion rejection — see
imaging.make_reprojector). Next experiment: mask disocclusions via forward-backward flow consistency and patch them from the fresh frame before pairing. - The
neutralsweep leg under dual-ref is effectively "freeze from frame ~5" — not a useful preservation baseline as-is. - The anchored control (
--mode anchored) was not needed for this verdict; optional.
Klein dual-ref reproject-mask A/B (4b, 2026-09-13) #
Outcome: the disocclusion mask behaves as designed end-to-end. Run via
scripts/sweep-klein-mask.sh (4b, 300 frames + 60-frame freeze tail,
role-naming enhance-slight prompt): legs --reproject vs
--reproject --reproject-mask, identical seed/steps/geometry.
- Both legs keep motion (3rd replication: reproject is the motion carrier).
- Mask firing on the actual clip (consecutive source pairs, run-identical constants): 1.9–7.3% of pixels per pass (mean ~4%), OOB 0.2–0.6% — fires at real motion; a flow-magnitude gate would change nothing. The carried state stays ~96% warped history per pass, so dual-ref semantics survive.
- Same-seed legs DECORRELATE fully within ~10 passes (frame 1 identical → MAD 34 at frame 10 → plateau ~85–100; 75% of pixels differ >20): klein's per-pass texture redraw amplifies small conditioning deltas, so fixed-seed determinism holds only for identical conditioning. A/Bs of this loop must be statistical, not frame-identity.
- Causal artifact metric (output-vs-source MAD inside vs outside the disocclusion mask, 30 saved frames): plain excess +3.5 (smears don't dominate pixel error), fbmask excess −47 (inside-mask output nearly matches source: 31 vs 78 MAD elsewhere) — patched fresh-frame content demonstrably propagates into the output; klein preserves the patch over stale state. 30/30 frames consistent.
- Zone CHARACTER vs source (same frames, disocclusion zones): HF energy (|Laplacian|) plain 26.8 / fbmask 10.8 / source 5.5; saturation plain 146 / fbmask 105 / source 98. The unmasked streaks are high-frequency, oversaturated replicate smears (they read as "blurry" at video speed but are spectrally sharp); fbmask zones sit near source character, with a residual ~2× HF from the per-pass enhance prompt. All three metrics (pixel error, spectrum, color) favor the mask.
9b confirmation (2026-09-13, same script, MODEL=flux2-klein-9b, both
legs complete): the 4b verdict replicates, with a twist that matches
eyeball impressions ("streaks filled with 3D-looking shapes" on plain):
| zones | MAD vs src | HF | saturation |
|---|---|---|---|
| 9b plain | 58.5 | 21.8 | 215 |
| 9b fbmask | 33.1 | 11.9 | 123 |
| 4b plain / fbmask | 92.5 / 31.0 | 26.8 / 10.8 | 146 / 105 |
| source | — | 5.5 | 98 |
The stronger editor doesn't produce raw smears — it resolves invalid history into plausible coherent structure (lower MAD than 4b plain) while cranking color/contrast (saturation 215 = 2.2× source): more convincing- LOOKING artifacts, still 4× source HF. The mask pins zones near source on BOTH models (fbmask rows are nearly model-independent — correct fresh content is preserved either way). Leg decorrelation as on 4b (t10 MAD 18, plateau ~50). fbmask cuts zone error −44% and halves the saturation excursion on 9b.
Neutral-prompt control & line concluded (2026-09-13; quantified from
untracked/output/pod-20260913_debug-flux2_4). Two separable findings:
- GLOBAL compounding drift is prompt-INDEPENDENT: analyze_drift delta_original (last-10 mean) neutral 88.5/89.0 (plain/fbmask) vs enhanced 78.4/92.0 — with an empty prompt the loop departs from the source just as far and the end frames look similar. The per-pass self-reference is the driver.
- The ZONE vividness WAS prompt-driven: disocclusion-zone HF/saturation drop to near-source under the neutral prompt on BOTH legs (plain 21.8→8.6 / 215→92; fbmask 11.9→5.9 / 123→97; source 5.5/98) — the enhance prompt was rendering stale/smeared zones vividly. The mask still wins on every axis under the neutral prompt: zone MAD 20.9 vs 48.9 (−57%), and churn delta_prev 22.6 vs 52.5.
With that, the experiment line is concluded: dual-ref + reproject + mask
is the best-characterized klein regime (mask verified near-source in
disocclusion zones on both models, with and without prompt), but the
fundamental per-pass compounding of a self-referencing regeneration loop
remains unsolved. Follow-through (mask as dual-ref default, sweep axis) is
SHELVED with the line — --reproject-mask stays opt-in; back to the
drawing board (candidates: explicit history/confidence buffer outside the
model, periodic hard re-anchoring, different conditioning entirely).
(Ops note: PROMPT is not in the run tag — the neutral sweep OVERWROTE the
enhanced A/B outputs pod-side, same mask-ab/ tags; the enhanced data
survives only in the earlier rsync. Differing-knob runs need OUT_PREFIX,
same rule as REF_ORDER in sweep-klein.sh.)
GPU selection: re-check the whole Runpod catalog (action item, opened 2026-09-13) #
Runtime re-calibration (2026-09-13, RTX 6000 Ada, measured — scoped):
the klein mask A/B ran NATIVELY (no --offload) with per leg 2 videos
(processed_stateful + one freeze tail = 300 + 60 = 360 passes): 9b at
6.24 / 6.14 s/pass (fbmask/plain legs; whole sweep 74 min), 4b at
~3.2 s/pass. Scope when reusing these numbers: passes = frames summed
over ALL videos a run emits (main + every tail mode), and note that the
earlier 20–25 s/frame figure was measured WITH --offload on a 32 GB
RTX 5090 (~10× native, see below) — NOT a native-vs-native comparison;
plan against throughput-per-dollar using the matching mode.
Context: the 2026-09-13 klein-validation run is on an L40S (48 GB,
Ada, sm_89). At commissioning time every 48 GB option showed Low stock on
both clouds (runpodctl gpu list: RTX 6000 Ada, L40S, A6000, A40); A100 SXM
was the only 80 GB card at Medium. An L40S was picked over a plain L40
despite the L40 being cheaper — see the reasoning below, which is exactly what
this action item exists to verify rather than assume.
Reasoning to record (L40 vs L40S, and cheap-vs-fitting in general): raw
$/hr is the wrong metric for this workload. A cheaper-but-slower 48 GB card
is only a win if the price ratio beats the runtime ratio — compare price per
pass (throughput per dollar), not price per hour. And either way, a slower
48 GB card that fits natively beats an OOM-ing faster card:
--offload moves a whole pipeline component to the GPU per pipeline call and
measured ~20–25 s per flux2-klein-9b frame on a 32 GB RTX 5090 (~10× native).
That number is the floor for any "just rent a cheaper 32 GB card" argument.
Action: re-check every GPU rentable on Runpod (secure AND community) against
its specs and rebuild the README_RUNPOD.md §1 table. For each candidate:
- VRAM — native fit for all five models (48 GB remains the working recommendation; per-model peaks in the README table),
- compute capability / sm_ version vs the CUDA 13 wheels in
uv.lock(host driver ≥ 580; sm_89/90/120 verified supported), $/hrsecure vs community (stock changes hourly — re-runrunpodctl gpu listand note the date),- throughput: record a real number where a run exists. Right now the tree has
no measured L40S figure at all; the 2026-09-13 run should note
frames/minute for
flux2-klein-9b(plus sd-turbo/sdxl-turbo), so the next comparison works from data instead of a guess.
Rows the current README table is missing: L40 (48 GB, non-S — lower
clocks/memory bandwidth than the L40S, often cheaper), and the non-48 GB
classes already listed (32 GB consumer, 80 GB A100/H100) should stay in the
table so the price/compute trade is explicit rather than implicit. The
appendix VRAM probe is the cheap way to get peak VRAM; a fixed
--max-frames run gives the throughput denominator.
Pod image retired: stock runpod/base + scripts/setup-pod.sh (2026-09-13) #
Decision: stop building refinementsystems/imgiter. Pods run the stock
runpod/base:1.3.0-rc.164-ubuntu2404, pinned by the same digest the custom
image was built FROM; scripts/setup-pod.sh (new, ships in every bundle)
installs uv 0.12.13 into /usr/local/bin and runs uv sync --frozen.
Why: Runpod starts billing when the container image pull starts. The
~12 GB custom image was justified as skipping the multi-GB uv sync on boot,
but it instead added a billed ~12 GB Docker Hub pull (often throttled) — the
~6 GB PyPI sync it skipped is cheaper, faster, and paid only when the lock
actually changes. Secondary: every uv.lock change forced an emulated
linux/amd64 rebuild + Docker Hub push + template digest re-pin; now lock
changes ride the normal bundle workflow. For scale, both sides are dwarfed by
the ~87.5 GB of HF model downloads every fresh pod pays anyway (container
disk is wiped on stop and restart), so the whole optimization was noise.
Mechanics that replaced the baked venv:
- Bundles extract into a fixed dir (
/workspace/imgiter,--strip-components=1), so the project.venv(uv's default location, created byuv sync) anduntracked/output/survive new bundle extracts; re-runningscripts/setup-pod.shafter each extract re-points the editable install (seconds whenuv.lockis unchanged). This replaces the image'sUV_PROJECT_ENVIRONMENT=/opt/imgiter/.venvenv var, which cannot be provided container-wide without a custom image — and without it, stamp-dir extraction would re-download the stack once per bundle. scripts/inputs.sh(sourced by every driver) now fails fast when.venvis missing, souv runcan never silently sync a fresh multi-GB venv mid-sweep.DRY_RUN=1previews bypass the guard.- uv 0.12.13 is pinned to match the lockfile producer and the
uv_buildbackend constraint (>=0.12.7,<0.13.0); installed from the GitHub release tarball (nocurl | sh).
Not changed: the pod template (id 04u1mmp8nf) keeps its disk/env/ports;
its image reference needs a one-time runpodctl template update --image
re-point (README_RUNPOD.md §3). The old image tags stay on Docker Hub.
hf-cache.sh's HF_HOME guard and the SSH/Jupyter behavior are base-image
features, unaffected.
MPS (Apple Silicon) backend, local validation (2026-09-13) #
What already worked (no lockfile change needed). uv sync on this dev
Mac installs torch 2.14.0 with a real MPS backend (is_available() == True);
uv.lock already carries the macosx_14_0_arm64 wheels next to the Linux
CUDA ones. torch.Generator("mps") constructs and seeds; fp16/bf16/fp32
matmuls run; the only CUDA hardcodes were models.load_pipeline() and
imaging.make_generator().
Design. Device is a host property: models.resolve_device() picks
CUDA -> MPS -> hard error (CPU only via explicit --device cpu), resolved
once per tool and threaded into load_pipeline(spec, offload, device) and
make_generator(..., device=...). ModelSpec stays device-free. New
--device {cuda,mps,cpu} lives in args.add_output_args (so dltb-klein gets
it too). On MPS load_pipeline enables attention slicing automatically.
--offload now calls enable_model_cpu_offload(device=device) (device
strings for CUDA are unchanged); MPS offload WORKS (verified) though it is
slower than native (klein-4b 1 step: 2m29s offloaded vs ~75 s/step native).
The five bash drivers' preflights accept CUDA or MPS
(SKIP_GPU_CHECK=1 still bypasses); scripts/smoke-local.sh (new) runs one
single-frame dltb-oneshot pass per locally-viable model with a
non-uniformity check on the produced frame.
Measured on M1/16 GB, macOS 27.0, models cached (wall clock, attention slicing on, 1 denoise step where applicable):
| model | setting | result |
|---|---|---|
sd-turbo |
512², default 4x0.4 = 1 step | ~2 s/pass (10 passes in 28.6 s; 3 in 14.4 s) |
sdxl-turbo |
768², 1 step | ~30-50 s/pass (10 passes in 8m7s; 3 in 1m54s; drifts with thermal/memory pressure) |
flux2-klein-4b |
768², 4 steps (no strength) | warm single frame 5m18s (~75 s/step), ~11 GB swap touched |
Checks that passed: scripts/smoke.sh PASSes locally on sd-turbo; two
fixed-seed MPS runs produce identical frame hashes (determinism within one
MPS build); analyze_drift.py on the local frames gives normal numbers
(delta_prev ~9.4, delta_original ~20.7 over 3 frames). dltb-klein is the
same loop (delegates to continuous.run), so it inherits --device.
Correction to the port plan: flux2-klein-4b DOES fit for single frames.
It loads and generates on 16 GB unified memory, but it swaps hard (11 GB
swap touched during 4 steps) and is far too slow for video; keep it in
the single-frame smoke list, out of video runs. flux-schnell (~34 GB) and
flux2-klein-9b (gated, ~20-29 GB) remain out of scope for local runs.
Caveats.
- MPS is not bit-identical to CUDA (different kernels/reductions): never
compare pixels across devices;
analyze_drift.pyis same-device only. On MPS,upcast_vae/fp32-VAE-upcast still works (SDXL prints the known benign diffusers false positive, see the log-noise section). - If an op errors with "not implemented for the mps backend",
PYTORCH_ENABLE_MPS_FALLBACK=1runs it on CPU silently (slow; escape hatch, not a default). No fp16-VAE black-frame artifacts were seen on any of the three models. - Local model cache is the default
~/.cache/huggingface;hf-cache.sh keep/cleanstill refuse to run withoutHF_HOMEset — do not relax that guard just because the local cache now matters. - Local and pod outputs must not share a tree: run tags do not encode the
device, so use
--output-dirsubtrees (untracked/output_<model>locally vsuntracked/output/pod-*on the pod). - Fixed a pre-existing
scripts/smoke.shbug found during local validation: the dltb-iterate artifact check usedframe_${ITERATIONS}.pngbut frames are written zero-padded (frame_0003.png), so the step could never pass; it now usesprintf '%04d'. Also,python3 src/dltb/analyze_drift.pyneeds numpy/PIL, which the macOS system python does not have — useuv run python src/dltb/analyze_drift.py(docs updated). - Pod CUDA regression (bundle ->
scripts/smoke.sh+--offloadcheck) is still pending; nothing in the changed code path is CUDA-specific, but re-run it on the next pod session before trusting a cross-device comparison.
Watch: re-validate after torch/macOS upgrades (MPS op coverage and
performance move quickly), and keep uv.lock's arm64 wheels in mind when
bumping torch.
DreamSim perceptual metrics: dltb-distance (2026-09-14) #
Why: analyze_drift.py only sees pixels. A loop that has settled into a
perceptual fixed point still chatters in pixel space (its delta_prev plateaus
above the 0.5 threshold), while a loop can be pixel-stable and yet look nothing
like its source. dltb-distance (src/dltb/analyze_distance.py) embeds
frames with DreamSim and reports two distances per frame:
dreamsim_to_ref— to the run's source image (frame_0000_original.png/frame_0000_source.png, else the first frame): perceptual drift.dreamsim_to_prev— to the previous analyzed frame: perceptual fixed-point detection. ~0 means the two frames are indistinguishable.
Validated on untracked/output/computer-enhance_sd-turbo/free-running/frames (200
frames, every 5th, ensemble): analyze_drift reported CONVERGED (delta_prev
0.119) with delta_original 97.3/255; DreamSim reported dreamsim_to_ref 0.789
(a stable image that is perceptually nothing like the source) and
dreamsim_to_prev 0.0006, minimum 7.4e-05 at frame 0126. Both lenses agree on
"settled" and differ, as designed, on "how far it went".
Facts verified for dreamsim 0.2.1 (do not re-derive):
uv add dreamsimadds exactly dreamsim, ftfy, open-clip-torch 3.3.0, peft 0.20.0, scipy 1.18.1, timm 1.0.29, torchvision 0.29.0, wcwidth — no existing pin changes (torch 2.14.0, transformers 5.17.0 stay). Linux-safe: only scipy/torchvision are compiled; the rest arepy3-none-anywheels.- Side effect: torchvision is now in the lock — see log noise §3 above for why that is pixel-safe.
- All six variants (
ensemble,dino_vitb16,clip_vitb32,open_clip_vitb32,dinov2_vitb14,synclr_vitb16) load and score;--patchworks fordino_vitb16(verified) anddinov2_vitb14(the only two the package ships patch checkpoints for). - MPS == CPU to float noise (~1e-7) on the same frames; a batch of N images
is one model call (that is what
--batch-sizeuses). - Identical images score exactly 0.0, but
1 - cos()undershoots on CPU (-2.4e-07 for a self-comparison), so the tool clamps at 0. Calibration:test_512.pngvs contrast x1.15 + brightness x1.05 scores ~0.0086 (ensemble) — that is what the default--converged-below 0.01is tuned to. - Weights come from GitHub releases via
torch.hub, and that CDN intermittently answers HTTP 504 mid-transfer (open_clip_vitb32anddinov2_vitb14failed 1-3 attempts during validation, then succeeded). Hence--retrieswith 2/4/8 s backoff; a failed download leaves no partial file. - dreamsim's cache handling is crude in two ways:
download_weightsusesos.mkdir(single level, so the tool doesmkdir -pitself), and every variant downloads to the same<cache_dir>/pretrained.zipwhile the cache check only looks at extracted checkpoints — so switching--dreamsim-typere-downloads the ~1.2 GB zip. Full ensemble cache: 3.8 GB. - Cache default is
--cache-dir untracked/models(gitignored, cwd-relative).untracked/is not packed into pod bundles (tracked files +untracked/input/only), so a fresh pod re-downloads ~2.7 GB — that is why the smoke leg is gated behindSMOKE_DISTANCE=1and runs on the frames step 2 already produced (no extra model pass). - Benign load noise, both on every load: the
torch.nn.utils.weight_normFutureWarning and peft's "Already found apeft_configattribute in the model. This will lead to having multiple adapters" UserWarning. Results are unaffected; do not add suppression without re-checking numbers. - Not encoded in the CSV/JSON names:
--every,--dreamsim-typeand--patch. Pass--out/--jsonwhen comparing variants, same rule as the run tags (AGENTS.md).
--mode anchored is the --anchor-blend 1.0 endpoint (2026-09-15) #
Analysis, then a small change. For the default pixel-blend conditioning,
blend((1-a)*R(P_{n-1}) + a*N_n) --a=1.0--> f(N_n)
because Image.blend(c, n, 1.0) = n exactly (uint8 -> float -> *1.0 -> round
is the identity), so the carried state — and any optical-flow warp of it —
contributes exactly zero. --mode stateful --anchor-blend 1.0 and
--mode anchored produce pixel-identical frames, same seeds, same tails;
dltb-distance-feedback's docstring already leaned on this endpoint as a
correctness check (a=1.0 there = repeated independent passes = boil test).
The only cost difference was reprojection: stateful defaults to
--reproject ON, and at a=1.0 the per-frame Farneback flow + warp was
computed and then blended away — pure waste. continuous._blend_conditioning
now skips flow/warp when anchor_blend >= 1.0, and run() prints a NOTE
when that skip is active (the header still says reproject=on because the
setting, not the work, drives the run tag). So stateful a=1.0 is now
pixel- AND cost-identical to anchored.
Why keep the flag at all:
- klein dual-ref.
--conditioning dual-refignores--anchor-blendentirely (weighting is the model's job via attention over two clean references), so the blend endpoint does not exist there;--mode anchoredis the only single-reference (boil-test) topology for klein. - Self-documenting CLI, distinct run tag (
<stem>_anchored,processed_anchored.mp4— smoke.sh asserts the latter), and it costs oneorclause in the main loop.
Docs updated to say the flag is redundant except under klein dual-ref (continuous.py + klein.py docstrings/help, README, AGENTS.md gotchas).