Deep Learning Tripping Balls
dltb AGENTS.md
13 kB
Markdown
at main

AGENTS.md #

Instructions for AI agents working in this repository (repo: imgiter, package: dltb).

What this is #

DLTB ("Deep Learning Tripping Balls") — experiments in feeding diffusion models their own output: free-running image self-iteration, video pipeline simulation with carried state, and failure tails. Research/hobby code, not a product. Two domains exist:

  • Local (macOS): editing, bundling (just bundle), and small runs on the Apple GPU via MPS (sd-turbo, sdxl-turbo, single-frame flux2-klein-4b); --help/arg checks need no accelerator at all (heavy imports are deliberately lazy).
  • Pod (Runpod, linux/amd64): actual GPU runs via the bundle workflow.

Read NOTES.md before changing anything nontrivial — it records hard-won investigations (pod SSH, AppleDouble tar pollution, klein guidance being inert) and design sketches (e.g. klein dual-reference conditioning, not yet implemented).

Commands #

Package management is uv only (Python 3.13; uv sync resolves platform-correct torch wheels — CUDA on Linux, MPS on macOS arm64; there is no requirements.txt). uv.lock ships in every bundle and is installed pod-side by scripts/setup-pod.sh, so lock changes ride the normal bundle workflow (there is no custom image to rebuild).

uv sync                                      # set up env
uv run dltb-oneshot --help                   # cheap local sanity check (no accelerator)
uv run dltb-iterate --model sd-turbo --input input_example/test_512.png --iterations 3
uv run dltb-oneshot --model sd-turbo --input input_example/test_512.png \
    --num-inference-steps 1 --strength 1.0   # one real local pass (MPS)
just bundle                                  # stage pod-ready tarball in bundle/
just hf-status | hf-keep <model> | hf-clean  # HF cache management

Experiment drivers (bash, env-var configurable, all support DRY_RUN=1):

DRY_RUN=1 scripts/sweep.sh                   # video sweep: boil test, blends, tails
DRY_RUN=1 scripts/sweep-klein.sh             # klein prompt ladder + steps probes +
                                            # reproject A/B (default under dual-ref)
DRY_RUN=1 scripts/sweep-prompt.sh            # free-running prompt x strength sweep on one image
scripts/smoke.sh                             # tiny run of every tool; needs CUDA
                                            # or Apple MPS (SMOKE_KLEIN=1 adds
                                            # dltb-klein, both conditionings)
scripts/smoke-local.sh                       # single-frame pass per locally-viable
                                            # model (sd-turbo, sdxl-turbo, klein-4b)
uv run dltb-distance <frames_dir>            # DreamSim perceptual drift/convergence
                                            # (SMOKE_DISTANCE=1 adds it to smoke.sh)
uv run dltb-distance-feedback --model sd-turbo --input input_example/test_512.png \
    --anchor-blend 0.3 --frames 200          # static-source feedback loop + chart
uv run dltb-stability-sweep --model sd-turbo --input input_example/test_512.png \
    --min-blend 0.05 --max-blend 0.6 --step 0.05 --frames 200 \
                                             # sweep anchor-blend: stable frames + table
uv run python src/dltb/analyze_drift.py <frames_dir>   # CPU-only drift metrics
uv run dltb-stabilization <metrics.csv>     # CPU-only verdict per metric column:
                                            # stabilized? where? at what value?
uv run python tests/test_detect_stabilization.py   # its stress battery (CPU-only)

The drivers read their input files (IMG/CLIP) from untracked/input/inputs.env, falling back to the tracked input_example/inputs.env; environment variables win (see scripts/inputs.sh).

There is no linter and one standalone test script (plain Python, no pytest): uv run python tests/test_detect_stabilization.py, the CPU-only stress battery for the stabilization detector. Verification ladder:

  1. uv run <tool> --help locally (catches import/arg breakage, no accelerator needed).
  2. uv run python tests/test_detect_stabilization.py (CPU-only, catches numeric regressions in detect_stabilization.py; its real-file regression case skips itself when the untracked CSV is absent).
  3. scripts/smoke.sh on the pod, and on Apple Silicon/MPS locally with sd-turbo, after deploying a bundle; scripts/smoke-local.sh covers the other locally-viable models.
  4. analyze_drift.py on produced frames for sanity of results; dltb-distance for the perceptual counterpart of the same frames (it needs torch and downloads DreamSim weights on first use).

Architecture #

Shared library in src/dltb/ + ten console scripts (see [project.scripts] in pyproject.toml):

File Role
models.py ModelSpec/MODELS table (5 models), load_pipeline, fast-fail checks, geometry rules
imaging.py run_pass (one model pass), prepare_frame, optical-flow reprojection
output.py run-directory layout (untracked/output_<model>/<stem>_<tag>/), timelapse assembly
args.py argparse flag groups shared by the tools
oneshot.py / iterate.py single pass / free-running self-iteration
continuous.py video loop: anchored boil test vs stateful blend, reproject, freeze/free/black tails
assemble.py CPU-only local mp4 assembly from a run's saved frames (split workflow; dltb-assemble)
analyze_distance.py DreamSim perceptual drift/convergence over a run's frames (dltb-distance); companion to analyze_drift.py, needs torch + first-run weights download
distance_loop.py iterate + measure in one run (dltb-distance-loop): free-running loop with a strength-capable model, then DreamSim on the saved frames -> distance_metrics.csv + distance_plot.png; the pipeline is unloaded and released before DreamSim loads
distance_feedback.py continuous' stateful loop with a static source (dltb-distance-feedback): one image repeated for --frames frames, --anchor-blend 0 = distance-loop, 1 = repeated fresh passes; saves every frame and every model input blend, then the same three-series DreamSim CSV + plot (input-vs-original is the third series)
detect_stabilization.py CPU-only verdict on a metrics CSV (dltb-stabilization): robust tail band (median ± k·MAD), backward scan for the last out-of-band smoothed sample, plateau diagnostics in band units (drift / level shift / wandering / oscillation / too short); reads distance_metrics.csv/drift_metrics.csv, writes stabilization_report.csv next to the input; tests/test_detect_stabilization.py is its stress battery
stability_sweep.py anchor-blend stability sweep (dltb-stability-sweep): dltb-distance-feedback's loop per blend with ONE pipeline and ONE DreamSim load, detect_stabilization verdicts on to_prev/to_ref; stable blends reduced to two plateau frames + run.json under stable/, unstable moved whole to unstable/, summary stability_sweep.csv + stability_plot.png
klein.py klein-restricted wrapper; dual-ref conditioning hook; delegates to continuous.run()

Key invariants:

  • Heavy imports (torch, PIL, cv2, numpy) stay inside functions so --help and arg validation never require a torch/accelerator environment. Do not hoist them.
  • Device is a host property: models.resolve_device() (auto: CUDA → MPS → error; --device overrides) threads an explicit device into load_pipeline() and make_generator(); ModelSpec never carries one.
  • One accelerator model at a time: when a tool phases two models (dltb-distance-loop: diffusion loop, then DreamSim), del the first and call models.release_accelerator(device) before loading the second, so peak VRAM is max(A, B) rather than the sum.
  • Model differences are encoded in ModelSpec (uses_strength, pass_size, gated, dtype/variant), not in if model == ... at call sites. Adding a model = adding a MODELS entry (plus docs/tables in README.md).
  • dltb-klein shares continuous.run() wholesale but customizes conditioning via the run(args, make_conditioning=...) hook: --conditioning dual-ref passes [state, frame] as two clean reference images instead of the pixel blend (--conditioning blend, the default, exercises the default _blend_conditioning). Its argument surface also restricts models, defaults blend to 0.1, and warns about inert guidance.
  • Run-directory tags encode only mode/blend-or-conditioning/tails (dual-ref: ..._dualref[-norepro]_tails…). Runs differing in other knobs (prompt, strength, steps, --ref-order) must use --output-dir subtrees or they silently overwrite earlier results (see how sweep.sh does it; sweep-klein.sh takes OUT_PREFIX for exactly this — the REF_ORDER A/B is not in the tag).
  • User-facing errors fail fast via raise SystemExit("message").
  • Most modules and scripts carry the ISC license header; keep it on new files in src/dltb/ and scripts/.
  • Long module docstrings double as --help text (RawDescriptionHelpFormatter); update them with behavior.

Model gotchas (verified, see NOTES.md) #

  • --guidance-scale > 1 is inert for the step-distilled klein models (CFG hard-disabled, no guidance embedding); --negative-prompt equally so. dltb-klein warns once; never "fix" this by removing the warning.
  • num_inference_steps * strength >= 1 is enforced (diffusers would run 0 denoise steps). Defaults 4 × 0.4 = exactly 1 actual step for img2img models; klein runs all 4 steps (no strength).
  • --width/--height must be multiples of 16; SD/SDXL reject explicit width/height while FLUX pipelines require them (pass_size flag).
  • flux2-klein-9b is gated: needs HF_TOKEN (checked at startup).
  • Klein is a reference-image editor: no --strength; the prompt is the per-pass edit-strength knob; blend active range is far below img2img models.
  • --mode anchored ≡ --mode stateful --anchor-blend 1.0 under pixel-blend conditioning: blend(P, N, 1) = N exactly and the flow/warp is skipped at a=1.0 (its result would be blended away), so frames and cost are identical. NOT equivalent under klein dual-ref, where the blend is ignored and anchored is the only single-reference loop.

Local (macOS) gotchas #

  • Any tar creation needs COPYFILE_DISABLE=1 or macOS pollutes archives with ._* AppleDouble members (bundle.sh does this + a python3 tarfile verification guard; never verify with tar -t, it hides them).
  • Bundles ship tracked files only, with one exception: gitignored untracked/input/ (user inputs) is packed explicitly by bundle.sh. git add new scripts before just bundle, or the pod silently misses them (the bundle script warns).
  • .gitignore covers untracked/ (user inputs, run outputs, DreamSim weight cache), bundle/, ._*, .DS_Store, .pi/ — keep generated artifacts out of git.

MPS (Apple Silicon) runs #

  • Device is auto-resolved per host (models.resolve_device()): CUDA on the pod, Apple MPS locally; --device {cuda,mps,cpu} overrides (all tools, including dltb-klein, via the shared output-args group).
  • Attention slicing is enabled automatically on MPS in load_pipeline().
  • Some ops are unimplemented on MPS: PYTORCH_ENABLE_MPS_FALLBACK=1 runs them on CPU (silently slow; escape hatch only).
  • Local and pod outputs must not share a tree: run tags do not encode the device. Use --output-dir subtrees (untracked/output_<model> locally vs untracked/output/pod-* on the pod) — same rule as other non-tag knobs.
  • MPS ≠ CUDA pixel parity is not a goal; compare frames only within one device (same as any other kernel change).

Pod workflow (see README_RUNPOD.md for the full runbook) #

just bundle                     # local
runpodctl send bundle/imgiter-<stamp>.tar.gz
# on pod:
cd /workspace && mkdir -p imgiter
tar xzf imgiter-<stamp>.tar.gz --strip-components=1 -C imgiter && cd imgiter
scripts/setup-pod.sh            # pinned uv + uv sync --frozen (re-run per bundle)
scripts/smoke.sh && scripts/sweep.sh
  • Pod image = stock runpod/base pinned by digest — no custom image: Runpod bills from the start of the image pull, so baking deps buys nothing (NOTES.md, 2026-09-13). scripts/setup-pod.sh installs uv 0.12.13 + the locked deps on the pod; bundles extract into the fixed dir /workspace/imgiter so .venv and outputs survive new bundle extracts.
  • just and runpodctl are NOT in the pod image — use scripts/*.sh directly.
  • 48 GB VRAM recommended; --offload for the two big models on smaller cards.
  • Container disk is ephemeral on stop AND restart — copy untracked/output/ out before stopping. Pod sshd isn't started by default (NOTES.md has the fix).
  • HF cache: swept models are kept by default (all five ≈ 87.5 GB fit the 150 GB pod disk); EVICT_CACHE=1 (env or inputs.env, like IMG/CLIP) restores keep-one eviction at sweep model boundaries. hf-cache.sh keep/clean refuse to run without HF_HOME set — keep that safety guard.