diff --git a/doc/temp/macos.md b/doc/temp/macos.md deleted file mode 100644 index 906fbd4..0000000 --- a/doc/temp/macos.md +++ /dev/null @@ -1,82 +0,0 @@ -Good news: the port is small. PyTorch on Apple Silicon ships with a built-in GPU backend (**MPS**, built on Metal) [^1], so you don't need CUDA, and Hugging Face `diffusers` supports it directly. You mostly change the device string and adjust a few memory/perf options. - -## 1. Install (macOS, Apple Silicon) - -```bash -# default PyPI wheels for macOS arm64 include MPS support -pip install torch torchvision -pip install diffusers transformers accelerate safetensors -``` - -The latest stable PyTorch supports MPS on Apple Silicon Macs running macOS 14.0 or later [^2]. You can sanity-check with: - -```python -import torch -print(torch.backends.mps.is_available()) # -> True -``` - -## 2. Code changes - -The official SD-Turbo usage is [^3]: - -```python -from diffusers import AutoPipelineForText2Image -import torch - -pipe = AutoPipelineForText2Image.from_pretrained("stabilityai/sd-turbo", torch_dtype=torch.float16, variant="fp16") -pipe.to("cuda") - -image = pipe(prompt=prompt, num_inference_steps=1, guidance_scale=0.0).images[0] -``` - -The Apple Silicon version is the same, with `"mps"` instead of `"cuda"` and attention slicing enabled: - -```python -from diffusers import AutoPipelineForText2Image -import torch - -device = "mps" -pipe = AutoPipelineForText2Image.from_pretrained( - "stabilityai/sd-turbo", - torch_dtype=torch.float16, - variant="fp16", -) -pipe.to(device) -pipe.enable_attention_slicing() # recommended on Apple Silicon - -prompt = "A cinematic shot of a baby raccoon wearing an intricate italian priest robe." -image = pipe(prompt=prompt, num_inference_steps=1, guidance_scale=0.0).images[0] -image.save("out.png") -``` - -Notes that carry over unchanged: SD-Turbo ignores `guidance_scale` (you disable it with `0.0`) and `negative_prompt`, and a **single step** is enough [^3]. If your CUDA code loops over steps 1–4 for quality, that works identically here. - -If your existing script has CUDA-specific bits, map them like this: - -| CUDA-side construct | Apple Silicon replacement | -|---|---| -| `torch.device("cuda")` / `pipe.to("cuda")` | `torch.device("mps")` / `pipe.to("mps")` | -| `torch.cuda.amp.autocast()` | `torch.autocast("mps")` (usually unnecessary — pass `torch_dtype=float16` at load) | -| `torch.cuda.empty_cache()` | `torch.mps.empty_cache()` | -| `xformers` / FlashAttention (`enable_xformers_memory_efficient_attention`) | Remove it — not available on MPS; diffusers falls back to scaled-dot-product attention, which uses custom Metal kernels [^4] | -| `torch.compile(...)` | Remove it — MPS support is an early prototype and end-to-end use "is likely to fail" [^5] | - -## 3. Gotchas specific to MPS - -- **Attention slicing**: diffusers recommends `pipe.enable_attention_slicing()` for machines with < 64 GB RAM, or when generating at resolutions above 512×512 — it reduces memory pressure and prevents swapping [^6]. SD-Turbo is small (~2.5 GB fp16), so if you're on a 32 GB+ machine at 512×512 it's optional, but harmless to enable. -- **Batching**: generating multiple prompts in a batch can crash or fail unreliably on MPS — iterate over prompts instead [^6]. -- **Tensor size limit**: the MPS backend doesn't support arrays larger than 2³² elements [^6] — irrelevant at 512×512 but a reason some ops bail at large upscale sizes. -- **Ops not implemented**: if you hit `...not implemented for the mps backend` errors while running custom code (e.g., custom VAE decode or post-processing), set `PYTORCH_ENABLE_MPS_FALLBACK=1` in the environment so PyTorch silently runs those ops on CPU. -- **First run**: PyTorch ≥ 2.x doesn't need the old "priming" pass that PyTorch 1.13 required on MPS [^6]. -- **Speed**: expect roughly comparable-to-slower-than-CUDA latency per step (a consumer NVIDIA card is faster); the difference matters far less for SD-Turbo than for full SD since it's 1 step by default. First generation after load is slow due to model download + kernel warmup; subsequent runs are much faster. - -That's the whole port: swap the device to `"mps"`, drop any CUDA-only speedups (xformers, `torch.compile`), and add attention slicing. - -**References** - -[^1]: [MPS backend — PyTorch main documentation](https://docs.pytorch.org/docs/main/notes/mps.html) (7%) -[^2]: [Accelerated PyTorch training on Mac - Metal](https://developer.apple.com/metal/pytorch/) (15%) -[^3]: [stabilityai/sd-turbo · Hugging Face](https://huggingface.co/stabilityai/sd-turbo) (13%) -[^4]: [GitHub - jhurt/attention-mps-torch: A custom PyTorch operator for ...](https://github.com/jhurt/attention-mps-torch) (6%) -[^5]: [torch.compile on MPS progress tracker · Issue #150121...](https://github.com/pytorch/pytorch/issues/150121) (8%) -[^6]: [Metal Performance Shaders (MPS) · Hugging Face](https://huggingface.co/docs/diffusers/en/optimization/mps) (51%) diff --git a/doc/temp/plan-macos-port.md b/doc/temp/plan-macos-port.md deleted file mode 100644 index 0349280..0000000 --- a/doc/temp/plan-macos-port.md +++ /dev/null @@ -1,398 +0,0 @@ -# MPS (Apple Silicon) backend — implementation plan - -Goal: run the small img2img models (`sd-turbo`, optionally `sdxl-turbo`) -natively on the development Mac's GPU (Apple MPS) so small experiments do not -burn rented pod time. Source material: `doc/temp/macos.md` (a generic -SD-Turbo-to-MPS guide). This plan maps that guide onto the actual repo state -as of **2026-09-13**; function names, not line numbers, are the stable -references (line numbers are included for orientation only). - -Scope: device selection + real MPS execution for the existing tools. The pod -(CUDA) path must remain behaviorally unchanged — same bundle workflow, same -lockfile, same defaults. - ---- - -## 0. Verified starting state (do not re-litigate) - -Checked on the dev machine (M1, 16 GB, macOS 27.0, arm64) on 2026-09-13: - -- `uv sync` **already installs an MPS-capable torch**: torch 2.14.0, - `torch.backends.mps.is_available() == True`. `uv.lock` already carries the - `macosx_14_0_arm64` wheels alongside the Linux CUDA ones. **No dependency - or lockfile change is needed or wanted.** -- `torch.Generator("mps")` constructs and seeds successfully (torch 2.14). -- fp16, bf16 and fp32 matmuls all run on the M1 MPS backend. -- accelerate's `cpu_offload_with_hook(model, execution_device="mps")` works - at the accelerate level (validated with a small `nn.Linear`). -- Repo has exactly **two** CUDA hardcodes in `src/`: - - `models.load_pipeline()` → `pipe.to("cuda")` - - `imaging.make_generator()` → `torch.Generator("cuda")` -- Five bash drivers hard-fail on `torch.cuda.is_available()`: - `scripts/smoke.sh`, `sweep.sh`, `sweep-klein.sh`, `sweep-klein-mask.sh`, - `sweep-prompt.sh` (each has a `SKIP_GPU_CHECK=1` bypass). -- The doc's "remove xformers / `torch.compile` / autocast / empty_cache" - advice is a **no-op**: the repo uses none of them. -- `assemble.py` / `analyze_drift.py` are pure CPU and unaffected. - -## 1. Design decisions (agree before coding) - -1. **Device is a host property, not a model property.** `ModelSpec` stays - device-free; a new `models.resolve_device()` picks the backend. - Auto order: **CUDA → MPS → error**. CPU is only used when requested - explicitly (`--device cpu`), so nobody silently gets a 100× slower run. -2. **New `--device {cuda,mps,cpu}` flag, default `None` = auto**, added to - the shared `add_output_args()` group (which already owns `--offload`). - This is primarily an escape hatch for debugging; normal local use needs - no flag. -3. **Resolve the device once per tool, before `load_pipeline()`**, so an - unavailable backend fails before a multi-GB model download. Thread the - resolved string explicitly to both `load_pipeline()` and - `make_generator()`; no module-level global. -4. **Attention slicing is enabled automatically on MPS** in - `load_pipeline()` (doc's <64 GB rule; this machine has 16 GB). It is - cheap to add and revert. If validation shows a significant slowdown with - no memory benefit, replace with a `--attention-slicing` flag (follow-up, - not in this plan). -5. **`--offload` stays device-aware**: `pipe.enable_model_cpu_offload(device=device)`. - Passing `device="cuda"` explicitly must be pod-smoke-verified as - equivalent to today's no-arg call. If diffusers' MPS offload turns out - broken during validation, fall back to a fail-fast `SystemExit` on MPS - ("unified memory already fits the small models"). -6. **No output-layout changes.** Run tags do not encode the device (same - policy as other knobs, AGENTS.md). Keep local and pod runs in separate - `--output-dir` subtrees when comparing; local defaults (`output_`) - are fine while pod outputs live under `output/pod-*`. -7. **Model scope**: `sd-turbo` is the local target; `sdxl-turbo` is the - stretch goal; `flux2-klein-4b` is expected to fit for SINGLE FRAMES on - this 16 GB machine (too slow for video, but one-pass generation is the - useful local case — verified in step 6b). `flux-schnell` (~34 GB) and - `flux2-klein-9b` (gated, ~20-29 GB) are documented as not targeted on - consumer Macs, but **not** hard-blocked in code. - -## 2. Code steps - -### Step 1 — `src/dltb/models.py`: device resolver + loader - -Add next to `resolve_geometry()` (after `check_requirements()`): - -```python -def resolve_device(requested: str | None = None) -> str: - """Pick the compute backend for this host. - - None = auto: CUDA if available, else Apple MPS, else a hard error - (a silent CPU run would be ~100x slower; pass --device cpu to force it). - """ - import torch - if requested is None: - if torch.cuda.is_available(): - return "cuda" - if torch.backends.mps.is_available(): - return "mps" - raise SystemExit( - "no CUDA or MPS accelerator available; pass --device cpu to run " - "on CPU anyway (very slow)." - ) - if requested == "cuda" and not torch.cuda.is_available(): - raise SystemExit("--device cuda requested but torch reports no CUDA device") - if requested == "mps" and not torch.backends.mps.is_available(): - raise SystemExit( - "--device mps requested but torch reports no MPS device (needs " - "Apple Silicon and an MPS-enabled PyTorch build)." - ) - return requested -``` - -Change `load_pipeline(spec: ModelSpec, offload: bool)` → -`load_pipeline(spec: ModelSpec, offload: bool, device: str)`: - -```python - if offload: - pipe.enable_model_cpu_offload(device=device) - else: - pipe.to(device) - if device == "mps": - pipe.enable_attention_slicing() # 16 GB-class unified memory - pipe.set_progress_bar_config(disable=True) - print(f"device: {device}", flush=True) - return pipe -``` - -Update the module docstring (it doubles as `--help` text where used) and the -`load_pipeline` docstring to say "accelerator (CUDA or Apple MPS)" instead -of "GPU". - -### Step 2 — `src/dltb/imaging.py`: generator device - -```python -def make_generator(seed: int, fixed_seed: bool, i: int = 0, *, device: str): - """Fresh accelerator generator; i varies the seed when fixed_seed is off. - ... - """ - import torch - return torch.Generator(device).manual_seed(seed if fixed_seed else seed + i) -``` - -`device` is keyword-only and required, so no call site can silently keep the -old hardcode. Update the docstring ("Fresh CUDA generator" → "Fresh -accelerator generator") and the module header comment ("never require a -CUDA environment" → "never require a torch/accelerator environment"). - -### Step 3 — `src/dltb/args.py`: `--device` - -In `add_output_args()` (next to `--offload`): - -```python - p.add_argument("--device", choices=["cuda", "mps", "cpu"], default=None, - help="Compute backend (default: auto -- CUDA if available, " - "else Apple MPS; CPU only if requested explicitly)") -``` - -Update the `--offload` help to "CPU model offloading for smaller -accelerators (slower)". Do **not** mention klein in any of this; it shares -the group already. - -### Step 4 — tool call sites (3 files) - -`oneshot.py`, `iterate.py`, `continuous.py` — in each `run()`: - -```python -from .models import (... , resolve_device) - -device = resolve_device(args.device) # before load_pipeline() -... -pipe = load_pipeline(spec, args.offload, device) -... -make_generator(args.seed, args.fixed_seed, i, device=device) # i / idx per file -``` - -Exact current call sites: `oneshot.py:64,75` · `iterate.py:80,93` · -`continuous.py:213,242` (`idx`). `klein.py` needs no change (it delegates to -`continuous.run()` and inherits `--device` via `add_output_args()`). - -### Step 5 — script preflights (5 files) - -In `scripts/smoke.sh`, `sweep.sh`, `sweep-klein.sh`, `sweep-klein-mask.sh`, -`sweep-prompt.sh`, replace the probe: - -```bash -uv run python -c 'import sys, torch; sys.exit(0 if torch.cuda.is_available() else 1)' -``` - -with: - -```bash -uv run python -c 'import sys, torch; sys.exit(0 if (torch.cuda.is_available() or torch.backends.mps.is_available()) else 1)' -``` - -and update each failure message from "torch reports no CUDA device -- run -this on a GPU pod" to mention both backends, e.g. "torch reports no CUDA or -MPS device -- run this on a GPU pod or Apple Silicon". Keep the -`SKIP_GPU_CHECK=1` bypass and its documentation. Also update the stale -comment at `sweep.sh:128` ("the tools build a CUDA generator") and the -`smoke.sh` header ("against a real GPU" → "against a real accelerator (CUDA -or Apple MPS)"). Keep the preflight accelerator-only (no CPU). - -### Step 6 — local validation (MPS) - -```bash -# 1. lazy imports still lazy: every tool's --help must work, and show --device -uv run dltb-oneshot --help && uv run dltb-iterate --help \ - && uv run dltb-continuous --help && uv run dltb-klein --help \ - && uv run dltb-assemble --help - -# 2. resolver behavior (no model download) -uv run python -c 'from dltb.models import resolve_device; print(resolve_device())' -# expect "mps" on this Mac (or "cuda" on the pod) - -# 3. end-to-end: sd-turbo smoke (~2.5 GB download on first run, -# into the default ~/.cache/huggingface; see Risks) -scripts/smoke.sh - -# 4. determinism within MPS: two fixed-seed identical runs must match -uv run dltb-iterate --model sd-turbo --input input_example/test_512.png \ - --iterations 3 --fixed-seed --save-every 1 --output-dir /tmp/mps-det-a -uv run dltb-iterate --model sd-turbo --input input_example/test_512.png \ - --iterations 3 --fixed-seed --save-every 1 --output-dir /tmp/mps-det-b -find /tmp/mps-det-a -name 'frame_*.png' | sort | xargs shasum | awk '{print $1}' > /tmp/a.sha -find /tmp/mps-det-b -name 'frame_*.png' | sort | xargs shasum | awk '{print $1}' > /tmp/b.sha -diff /tmp/a.sha /tmp/b.sha && echo "deterministic on MPS" - -# 5. timing for NOTES.md (wall clock / passes) -time uv run dltb-iterate --model sd-turbo --input input_example/test_512.png \ - --iterations 10 --save-every 10 --output-dir /tmp/mps-timing - -# 6. drift tool sanity on local frames -python3 src/dltb/analyze_drift.py /tmp/mps-det-a/*/frames -# NOTE (found in validation): on macOS the system python3 has no numpy/PIL -- -# use `uv run python src/dltb/analyze_drift.py ...` locally. - -# 7. stretch: sdxl-turbo at its default 768 (memory pressure + attention slicing) -uv run dltb-oneshot --model sdxl-turbo --input input_example/test_512.png \ - --num-inference-steps 1 --strength 1.0 --output-dir /tmp/mps-sdxl -``` - -### Step 6b — local single-frame smoke test, every model that fits - -`scripts/smoke-local.sh` (added 2026-09-13, after the initial plan): the -local counterpart of `scripts/smoke.sh` — one `dltb-oneshot` pass per -locally-viable model, accelerator auto-selected. Default -`MODELS="sd-turbo sdxl-turbo flux2-klein-4b"` (env-overridable); -`flux-schnell` and `flux2-klein-9b` are excluded (memory/gated) but can be -added explicitly. Budget per model: img2img 1 step x strength 1.0, klein its -model-card 4 steps. Each produced frame is checked for existence AND -non-uniformity (a uniform frame is the fp16-VAE-on-MPS failure mode, not a -generation); per-model wall clock is logged (first run includes the -download). Output under `output/smoke-local//`. - -```bash -scripts/smoke-local.sh -MODELS="sd-turbo flux2-klein-4b" scripts/smoke-local.sh # subset -OFFLOAD=1 scripts/smoke-local.sh # add --offload -``` - -Expected: PASS on sd-turbo + sdxl-turbo; klein-4b settles whether the 4b -single-frame case is viable on this machine (which decides whether it stays -in the default list and the README note). - -### Step 7 — pod regression (CUDA must be untouched) - -```bash -just bundle -runpodctl send bundle/imgiter-.tar.gz -# on pod: extract, scripts/setup-pod.sh -scripts/smoke.sh - -# exercise the changed --offload code path explicitly (smoke.sh does not): -uv run dltb-oneshot --model sd-turbo --input input_example/test_512.png \ - --num-inference-steps 1 --strength 1.0 --offload --output-dir /tmp/pod-offload -# expect "device: cuda" and identical behavior to the pre-change runs -``` - -## 3. Documentation steps - -### Step 8 — module docstrings (required by AGENTS.md: docstrings double as help) - -- `src/dltb/models.py`: mention `resolve_device` and the auto order in the - module docstring. -- `src/dltb/imaging.py`: "Fresh accelerator generator". -- `src/dltb/args.py`: `add_output_args()` docstring already says "device - placement" — extend to mention backend selection. -- `src/dltb/__init__.py`: change "CPU-only local post-processing (no - torch/CUDA; ...)" to "no torch/accelerator", and add one sentence to the - intro: model tools auto-select CUDA (pod) or Apple MPS (local). - -### Step 9 — `README.md` - -- Replace `## Setup (NVIDIA GPU machine)` with `## Setup`, keeping the - existing `uv sync` text but reworded: the lock resolves platform-correct - PyTorch wheels (CUDA on Linux, MPS on macOS arm64); no requirements.txt. -- Add `### Apple Silicon (local, MPS)` under it: - - `uv sync`, then a one-pass example: - `uv run dltb-oneshot --model sd-turbo --input input_example/test_512.png --num-inference-steps 1 --strength 1.0` - - Device is auto-detected (CUDA → MPS); `--device` overrides; attention - slicing is enabled automatically on MPS. - - Local end-to-end check: `scripts/smoke.sh` (sd-turbo, ~2.5 GB download). - - Keep local and pod runs in separate `--output-dir` subtrees. - - Escape hatch for unimplemented ops: `PYTORCH_ENABLE_MPS_FALLBACK=1` - (silently slow, use only if an op errors). -- After the Models table, add a short "Local Apple Silicon" note: - `sd-turbo` is the practical target, `sdxl-turbo` is the stretch goal, and - `flux2-klein-4b` can generate single frames on a 16 GB machine (too slow - for video); `flux-schnell` and `flux2-klein-9b` are not targeted on - consumer Macs (memory/speed/gated), though nothing blocks trying them. - Point at `scripts/smoke-local.sh` as the all-fitting-models check. -- Leave `README_RUNPOD.md` alone (pod-only). - -### Step 10 — `AGENTS.md` - -- "Commands": add the local sd-turbo one-liner and `scripts/smoke-local.sh` - next to the `--help` sanity check; change the `scripts/smoke.sh` comment - from "needs GPU" to "needs CUDA or Apple MPS". -- "Local (macOS)" section: add a "MPS (Apple Silicon)" block covering: - device auto-resolution + `--device`, automatic attention slicing on MPS, - the `PYTORCH_ENABLE_MPS_FALLBACK=1` escape hatch, separation of local/pod - outputs (same rule as other knobs: `--output-dir` subtree), and that - MPS≠CUDA pixel parity is not a goal. -- Verification ladder: step 2 becomes "`scripts/smoke.sh` on the pod (and on - Apple Silicon/MPS locally with sd-turbo) after deploying a bundle". -- Key invariants: add "Device is a host property: `models.resolve_device()`; - `ModelSpec` never carries a device." - -### Step 11 — `NOTES.md` - -Add a dated section (repository note style: bold lead-ins) with: -- **What already worked**: torch 2.14 from `uv.lock` is MPS-enabled on - arm64; `Generator("mps")`; fp16/bf16/fp32; accelerate MPS offload. -- **Design**: auto order cuda → mps, explicit `--device`, attention slicing - on MPS, device threaded to the generator. -- **Measured** (fill from Step 6): sd-turbo s/pass on M1/16 GB, sdxl-turbo - s/pass, with attention slicing on. -- **Caveats**: no bit parity with CUDA (don't compare cross-device pixels; - `analyze_drift.py` only within a device); fp16 VAE fallback options - (`PYTORCH_ENABLE_MPS_FALLBACK=1`, or VAE in fp32) if artifacts appear; - local model cache is the default `~/.cache/huggingface` and - `hf-cache.sh keep/clean`'s "refuse without HF_HOME" guard must stay. -- **Watch**: re-validate on torch/macOS upgrades. - -### Step 12 — `doc/temp/macos.md` - -No edits (temporary document, will be deleted). Optionally add a one-line -pointer at its top to this plan for the duration of the work; do not -reference `doc/temp/` from non-temporary docs. - -## 4. Acceptance checklist - -- [ ] `uv run --help` works for all five tools, offline, no CUDA/MPS - initialization; the four model tools show `--device` (dltb-assemble is - CPU-only and takes none). -- [ ] `resolve_device()` returns `mps` on the Mac and `cuda` on the pod, - and fails fast (before downloads) for `--device cuda` on the Mac. -- [ ] `scripts/smoke.sh` PASSES locally on sd-turbo. -- [ ] `scripts/smoke-local.sh` produces one sane (non-uniform) frame for - every default model: sd-turbo, sdxl-turbo, flux2-klein-4b (the 4b - single-frame case is the added scope). -- [ ] Two fixed-seed MPS runs produce identical frame hashes (step 6.4). -- [ ] Pod: `scripts/smoke.sh` passes and `--offload` still works - ("device: cuda"). -- [ ] No changes to `pyproject.toml` / `uv.lock`; bundle contents unchanged - apart from the edited tracked files. -- [ ] Docs from §3 updated; NOTES.md carries measured numbers. - -## 5. Risks and fallbacks - -| Risk | Fallback | -| --- | --- | -| fp16 VAE produces NaNs/black tiles on MPS | `PYTORCH_ENABLE_MPS_FALLBACK=1` first; if insufficient, run VAE in fp32 (`pipe.vae.to(torch.float32)`) — document whichever worked | -| Unsupported op in a pipeline | `PYTORCH_ENABLE_MPS_FALLBACK=1` (silent CPU op; expect a slowdown) | -| Attention slicing slows sd-turbo noticeably | Measure (step 6.5); then gate it behind `--attention-slicing` or drop it for small models | -| diffusers `enable_model_cpu_offload(device="mps")` broken despite accelerate support | Replace with a fail-fast `SystemExit` on MPS; `--offload` remains CUDA-only | -| Passing `device="cuda"` to `enable_model_cpu_offload()` changes pod behavior | Caught by step 7; if so, keep the no-arg call for CUDA and pass `device` only otherwise | -| Local/Pod outputs silently overwrite | Use `--output-dir` subtrees (documented); do not encode device in run tags | -| 16 GB M1 swaps on sdxl-turbo@768 | Attention slicing on; treat sdxl as best-effort; sd-turbo remains the supported local path | - -## 6. Out of scope / follow-ups - -- MPS support/performance work for `flux-schnell` and `flux2-klein-9b` - (klein-4b single-frame smoke is in scope; video-length klein runs are not). -- Automatic device suffixing of output directories (documented convention - instead). -- `torch.mps.empty_cache()` in long `dltb-continuous` loops (no CUDA - equivalent exists today either). -- Batching, `torch.compile`, xformers — not used by this repo. -- CI/tests — the repo has none; the smoke script is the acceptance test. - -## Appendix — files touched - -| File | Change | -| --- | --- | -| `src/dltb/models.py` | `resolve_device()`, `load_pipeline(..., device)`, docstrings | -| `src/dltb/imaging.py` | `make_generator(..., *, device)`, docstrings | -| `src/dltb/args.py` | `--device`, `--offload` help | -| `src/dltb/oneshot.py` | resolve + pass device (2 call sites) | -| `src/dltb/iterate.py` | resolve + pass device (2 call sites) | -| `src/dltb/continuous.py` | resolve + pass device (2 call sites) | -| `scripts/{smoke,sweep,sweep-klein,sweep-klein-mask,sweep-prompt}.sh` | MPS-aware preflight + comments | -| `scripts/smoke-local.sh` | **new**: single-frame pass for every locally-viable model | -| `README.md`, `AGENTS.md`, `NOTES.md` | MPS documentation | -| `pyproject.toml`, `uv.lock`, `scripts/bundle.sh`, `scripts/setup-pod.sh` | **unchanged on purpose** |