From a4bf09888587b7a48a0140ad7d38d57c436afe22 Mon Sep 17 00:00:00 2001 From: Refinement Systems Date: Sun, 13 Sep 2026 20:35:10 +0200 Subject: [PATCH] plan for mac port --- doc/temp/plan-macos-port.md | 362 ++++++++++++++++++++++++++++++++++++ 1 file changed, 362 insertions(+) create mode 100644 doc/temp/plan-macos-port.md diff --git a/doc/temp/plan-macos-port.md b/doc/temp/plan-macos-port.md new file mode 100644 index 0000000..a8fb0ad --- /dev/null +++ b/doc/temp/plan-macos-port.md @@ -0,0 +1,362 @@ +# MPS (Apple Silicon) backend — implementation plan + +Goal: run the small img2img models (`sd-turbo`, optionally `sdxl-turbo`) +natively on the development Mac's GPU (Apple MPS) so small experiments do not +burn rented pod time. Source material: `doc/temp/macos.md` (a generic +SD-Turbo-to-MPS guide). This plan maps that guide onto the actual repo state +as of **2026-09-13**; function names, not line numbers, are the stable +references (line numbers are included for orientation only). + +Scope: device selection + real MPS execution for the existing tools. The pod +(CUDA) path must remain behaviorally unchanged — same bundle workflow, same +lockfile, same defaults. + +--- + +## 0. Verified starting state (do not re-litigate) + +Checked on the dev machine (M1, 16 GB, macOS 27.0, arm64) on 2026-09-13: + +- `uv sync` **already installs an MPS-capable torch**: torch 2.14.0, + `torch.backends.mps.is_available() == True`. `uv.lock` already carries the + `macosx_14_0_arm64` wheels alongside the Linux CUDA ones. **No dependency + or lockfile change is needed or wanted.** +- `torch.Generator("mps")` constructs and seeds successfully (torch 2.14). +- fp16, bf16 and fp32 matmuls all run on the M1 MPS backend. +- accelerate's `cpu_offload_with_hook(model, execution_device="mps")` works + at the accelerate level (validated with a small `nn.Linear`). +- Repo has exactly **two** CUDA hardcodes in `src/`: + - `models.load_pipeline()` → `pipe.to("cuda")` + - `imaging.make_generator()` → `torch.Generator("cuda")` +- Five bash drivers hard-fail on `torch.cuda.is_available()`: + `scripts/smoke.sh`, `sweep.sh`, `sweep-klein.sh`, `sweep-klein-mask.sh`, + `sweep-prompt.sh` (each has a `SKIP_GPU_CHECK=1` bypass). +- The doc's "remove xformers / `torch.compile` / autocast / empty_cache" + advice is a **no-op**: the repo uses none of them. +- `assemble.py` / `analyze_drift.py` are pure CPU and unaffected. + +## 1. Design decisions (agree before coding) + +1. **Device is a host property, not a model property.** `ModelSpec` stays + device-free; a new `models.resolve_device()` picks the backend. + Auto order: **CUDA → MPS → error**. CPU is only used when requested + explicitly (`--device cpu`), so nobody silently gets a 100× slower run. +2. **New `--device {cuda,mps,cpu}` flag, default `None` = auto**, added to + the shared `add_output_args()` group (which already owns `--offload`). + This is primarily an escape hatch for debugging; normal local use needs + no flag. +3. **Resolve the device once per tool, before `load_pipeline()`**, so an + unavailable backend fails before a multi-GB model download. Thread the + resolved string explicitly to both `load_pipeline()` and + `make_generator()`; no module-level global. +4. **Attention slicing is enabled automatically on MPS** in + `load_pipeline()` (doc's <64 GB rule; this machine has 16 GB). It is + cheap to add and revert. If validation shows a significant slowdown with + no memory benefit, replace with a `--attention-slicing` flag (follow-up, + not in this plan). +5. **`--offload` stays device-aware**: `pipe.enable_model_cpu_offload(device=device)`. + Passing `device="cuda"` explicitly must be pod-smoke-verified as + equivalent to today's no-arg call. If diffusers' MPS offload turns out + broken during validation, fall back to a fail-fast `SystemExit` on MPS + ("unified memory already fits the small models"). +6. **No output-layout changes.** Run tags do not encode the device (same + policy as other knobs, AGENTS.md). Keep local and pod runs in separate + `--output-dir` subtrees when comparing; local defaults (`output_`) + are fine while pod outputs live under `output/pod-*`. +7. **Model scope**: `sd-turbo` is the local target; `sdxl-turbo` is the + stretch goal. `flux-schnell` (~34 GB) and the klein pair are documented + as not targeted on consumer Macs, but **not** hard-blocked in code. + +## 2. Code steps + +### Step 1 — `src/dltb/models.py`: device resolver + loader + +Add next to `resolve_geometry()` (after `check_requirements()`): + +```python +def resolve_device(requested: str | None = None) -> str: + """Pick the compute backend for this host. + + None = auto: CUDA if available, else Apple MPS, else a hard error + (a silent CPU run would be ~100x slower; pass --device cpu to force it). + """ + import torch + if requested is None: + if torch.cuda.is_available(): + return "cuda" + if torch.backends.mps.is_available(): + return "mps" + raise SystemExit( + "no CUDA or MPS accelerator available; pass --device cpu to run " + "on CPU anyway (very slow)." + ) + if requested == "cuda" and not torch.cuda.is_available(): + raise SystemExit("--device cuda requested but torch reports no CUDA device") + if requested == "mps" and not torch.backends.mps.is_available(): + raise SystemExit( + "--device mps requested but torch reports no MPS device (needs " + "Apple Silicon and an MPS-enabled PyTorch build)." + ) + return requested +``` + +Change `load_pipeline(spec: ModelSpec, offload: bool)` → +`load_pipeline(spec: ModelSpec, offload: bool, device: str)`: + +```python + if offload: + pipe.enable_model_cpu_offload(device=device) + else: + pipe.to(device) + if device == "mps": + pipe.enable_attention_slicing() # 16 GB-class unified memory + pipe.set_progress_bar_config(disable=True) + print(f"device: {device}", flush=True) + return pipe +``` + +Update the module docstring (it doubles as `--help` text where used) and the +`load_pipeline` docstring to say "accelerator (CUDA or Apple MPS)" instead +of "GPU". + +### Step 2 — `src/dltb/imaging.py`: generator device + +```python +def make_generator(seed: int, fixed_seed: bool, i: int = 0, *, device: str): + """Fresh accelerator generator; i varies the seed when fixed_seed is off. + ... + """ + import torch + return torch.Generator(device).manual_seed(seed if fixed_seed else seed + i) +``` + +`device` is keyword-only and required, so no call site can silently keep the +old hardcode. Update the docstring ("Fresh CUDA generator" → "Fresh +accelerator generator") and the module header comment ("never require a +CUDA environment" → "never require a torch/accelerator environment"). + +### Step 3 — `src/dltb/args.py`: `--device` + +In `add_output_args()` (next to `--offload`): + +```python + p.add_argument("--device", choices=["cuda", "mps", "cpu"], default=None, + help="Compute backend (default: auto -- CUDA if available, " + "else Apple MPS; CPU only if requested explicitly)") +``` + +Update the `--offload` help to "CPU model offloading for smaller +accelerators (slower)". Do **not** mention klein in any of this; it shares +the group already. + +### Step 4 — tool call sites (3 files) + +`oneshot.py`, `iterate.py`, `continuous.py` — in each `run()`: + +```python +from .models import (... , resolve_device) + +device = resolve_device(args.device) # before load_pipeline() +... +pipe = load_pipeline(spec, args.offload, device) +... +make_generator(args.seed, args.fixed_seed, i, device=device) # i / idx per file +``` + +Exact current call sites: `oneshot.py:64,75` · `iterate.py:80,93` · +`continuous.py:213,242` (`idx`). `klein.py` needs no change (it delegates to +`continuous.run()` and inherits `--device` via `add_output_args()`). + +### Step 5 — script preflights (5 files) + +In `scripts/smoke.sh`, `sweep.sh`, `sweep-klein.sh`, `sweep-klein-mask.sh`, +`sweep-prompt.sh`, replace the probe: + +```bash +uv run python -c 'import sys, torch; sys.exit(0 if torch.cuda.is_available() else 1)' +``` + +with: + +```bash +uv run python -c 'import sys, torch; sys.exit(0 if (torch.cuda.is_available() or torch.backends.mps.is_available()) else 1)' +``` + +and update each failure message from "torch reports no CUDA device -- run +this on a GPU pod" to mention both backends, e.g. "torch reports no CUDA or +MPS device -- run this on a GPU pod or Apple Silicon". Keep the +`SKIP_GPU_CHECK=1` bypass and its documentation. Also update the stale +comment at `sweep.sh:128` ("the tools build a CUDA generator") and the +`smoke.sh` header ("against a real GPU" → "against a real accelerator (CUDA +or Apple MPS)"). Keep the preflight accelerator-only (no CPU). + +### Step 6 — local validation (MPS) + +```bash +# 1. lazy imports still lazy: every tool's --help must work, and show --device +uv run dltb-oneshot --help && uv run dltb-iterate --help \ + && uv run dltb-continuous --help && uv run dltb-klein --help \ + && uv run dltb-assemble --help + +# 2. resolver behavior (no model download) +uv run python -c 'from dltb.models import resolve_device; print(resolve_device())' +# expect "mps" on this Mac (or "cuda" on the pod) + +# 3. end-to-end: sd-turbo smoke (~2.5 GB download on first run, +# into the default ~/.cache/huggingface; see Risks) +scripts/smoke.sh + +# 4. determinism within MPS: two fixed-seed identical runs must match +uv run dltb-iterate --model sd-turbo --input input_example/test_512.png \ + --iterations 3 --fixed-seed --save-every 1 --output-dir /tmp/mps-det-a +uv run dltb-iterate --model sd-turbo --input input_example/test_512.png \ + --iterations 3 --fixed-seed --save-every 1 --output-dir /tmp/mps-det-b +find /tmp/mps-det-a -name 'frame_*.png' | sort | xargs shasum | awk '{print $1}' > /tmp/a.sha +find /tmp/mps-det-b -name 'frame_*.png' | sort | xargs shasum | awk '{print $1}' > /tmp/b.sha +diff /tmp/a.sha /tmp/b.sha && echo "deterministic on MPS" + +# 5. timing for NOTES.md (wall clock / passes) +time uv run dltb-iterate --model sd-turbo --input input_example/test_512.png \ + --iterations 10 --save-every 10 --output-dir /tmp/mps-timing + +# 6. drift tool sanity on local frames +python3 src/dltb/analyze_drift.py /tmp/mps-det-a/*/frames + +# 7. stretch: sdxl-turbo at its default 768 (memory pressure + attention slicing) +uv run dltb-oneshot --model sdxl-turbo --input input_example/test_512.png \ + --num-inference-steps 1 --strength 1.0 --output-dir /tmp/mps-sdxl +``` + +### Step 7 — pod regression (CUDA must be untouched) + +```bash +just bundle +runpodctl send bundle/imgiter-.tar.gz +# on pod: extract, scripts/setup-pod.sh +scripts/smoke.sh + +# exercise the changed --offload code path explicitly (smoke.sh does not): +uv run dltb-oneshot --model sd-turbo --input input_example/test_512.png \ + --num-inference-steps 1 --strength 1.0 --offload --output-dir /tmp/pod-offload +# expect "device: cuda" and identical behavior to the pre-change runs +``` + +## 3. Documentation steps + +### Step 8 — module docstrings (required by AGENTS.md: docstrings double as help) + +- `src/dltb/models.py`: mention `resolve_device` and the auto order in the + module docstring. +- `src/dltb/imaging.py`: "Fresh accelerator generator". +- `src/dltb/args.py`: `add_output_args()` docstring already says "device + placement" — extend to mention backend selection. +- `src/dltb/__init__.py`: change "CPU-only local post-processing (no + torch/CUDA; ...)" to "no torch/accelerator", and add one sentence to the + intro: model tools auto-select CUDA (pod) or Apple MPS (local). + +### Step 9 — `README.md` + +- Replace `## Setup (NVIDIA GPU machine)` with `## Setup`, keeping the + existing `uv sync` text but reworded: the lock resolves platform-correct + PyTorch wheels (CUDA on Linux, MPS on macOS arm64); no requirements.txt. +- Add `### Apple Silicon (local, MPS)` under it: + - `uv sync`, then a one-pass example: + `uv run dltb-oneshot --model sd-turbo --input input_example/test_512.png --num-inference-steps 1 --strength 1.0` + - Device is auto-detected (CUDA → MPS); `--device` overrides; attention + slicing is enabled automatically on MPS. + - Local end-to-end check: `scripts/smoke.sh` (sd-turbo, ~2.5 GB download). + - Keep local and pod runs in separate `--output-dir` subtrees. + - Escape hatch for unimplemented ops: `PYTORCH_ENABLE_MPS_FALLBACK=1` + (silently slow, use only if an op errors). +- After the Models table, add a short "Local Apple Silicon" note: + `sd-turbo` is the practical target, `sdxl-turbo` is the stretch goal; + `flux-schnell` and the klein editors are not targeted on consumer Macs + (memory/speed), though nothing blocks trying them. +- Leave `README_RUNPOD.md` alone (pod-only). + +### Step 10 — `AGENTS.md` + +- "Commands": add the local sd-turbo one-liner next to the `--help` sanity + check; change the `scripts/smoke.sh` comment from "needs GPU" to "needs + CUDA or Apple MPS". +- "Local (macOS)" section: add a "MPS (Apple Silicon)" block covering: + device auto-resolution + `--device`, automatic attention slicing on MPS, + the `PYTORCH_ENABLE_MPS_FALLBACK=1` escape hatch, separation of local/pod + outputs (same rule as other knobs: `--output-dir` subtree), and that + MPS≠CUDA pixel parity is not a goal. +- Verification ladder: step 2 becomes "`scripts/smoke.sh` on the pod (and on + Apple Silicon/MPS locally with sd-turbo) after deploying a bundle". +- Key invariants: add "Device is a host property: `models.resolve_device()`; + `ModelSpec` never carries a device." + +### Step 11 — `NOTES.md` + +Add a dated section (repository note style: bold lead-ins) with: +- **What already worked**: torch 2.14 from `uv.lock` is MPS-enabled on + arm64; `Generator("mps")`; fp16/bf16/fp32; accelerate MPS offload. +- **Design**: auto order cuda → mps, explicit `--device`, attention slicing + on MPS, device threaded to the generator. +- **Measured** (fill from Step 6): sd-turbo s/pass on M1/16 GB, sdxl-turbo + s/pass, with attention slicing on. +- **Caveats**: no bit parity with CUDA (don't compare cross-device pixels; + `analyze_drift.py` only within a device); fp16 VAE fallback options + (`PYTORCH_ENABLE_MPS_FALLBACK=1`, or VAE in fp32) if artifacts appear; + local model cache is the default `~/.cache/huggingface` and + `hf-cache.sh keep/clean`'s "refuse without HF_HOME" guard must stay. +- **Watch**: re-validate on torch/macOS upgrades. + +### Step 12 — `doc/temp/macos.md` + +No edits (temporary document, will be deleted). Optionally add a one-line +pointer at its top to this plan for the duration of the work; do not +reference `doc/temp/` from non-temporary docs. + +## 4. Acceptance checklist + +- [ ] `uv run --help` works for all five tools, offline, no CUDA/MPS + initialization, and shows `--device`. +- [ ] `resolve_device()` returns `mps` on the Mac and `cuda` on the pod, + and fails fast (before downloads) for `--device cuda` on the Mac. +- [ ] `scripts/smoke.sh` PASSES locally on sd-turbo. +- [ ] Two fixed-seed MPS runs produce identical frame hashes (step 6.4). +- [ ] Pod: `scripts/smoke.sh` passes and `--offload` still works + ("device: cuda"). +- [ ] No changes to `pyproject.toml` / `uv.lock`; bundle contents unchanged + apart from the edited tracked files. +- [ ] Docs from §3 updated; NOTES.md carries measured numbers. + +## 5. Risks and fallbacks + +| Risk | Fallback | +| --- | --- | +| fp16 VAE produces NaNs/black tiles on MPS | `PYTORCH_ENABLE_MPS_FALLBACK=1` first; if insufficient, run VAE in fp32 (`pipe.vae.to(torch.float32)`) — document whichever worked | +| Unsupported op in a pipeline | `PYTORCH_ENABLE_MPS_FALLBACK=1` (silent CPU op; expect a slowdown) | +| Attention slicing slows sd-turbo noticeably | Measure (step 6.5); then gate it behind `--attention-slicing` or drop it for small models | +| diffusers `enable_model_cpu_offload(device="mps")` broken despite accelerate support | Replace with a fail-fast `SystemExit` on MPS; `--offload` remains CUDA-only | +| Passing `device="cuda"` to `enable_model_cpu_offload()` changes pod behavior | Caught by step 7; if so, keep the no-arg call for CUDA and pass `device` only otherwise | +| Local/Pod outputs silently overwrite | Use `--output-dir` subtrees (documented); do not encode device in run tags | +| 16 GB M1 swaps on sdxl-turbo@768 | Attention slicing on; treat sdxl as best-effort; sd-turbo remains the supported local path | + +## 6. Out of scope / follow-ups + +- MPS support/performance work for flux-schnell and the klein pair. +- Automatic device suffixing of output directories (documented convention + instead). +- `torch.mps.empty_cache()` in long `dltb-continuous` loops (no CUDA + equivalent exists today either). +- Batching, `torch.compile`, xformers — not used by this repo. +- CI/tests — the repo has none; the smoke script is the acceptance test. + +## Appendix — files touched + +| File | Change | +| --- | --- | +| `src/dltb/models.py` | `resolve_device()`, `load_pipeline(..., device)`, docstrings | +| `src/dltb/imaging.py` | `make_generator(..., *, device)`, docstrings | +| `src/dltb/args.py` | `--device`, `--offload` help | +| `src/dltb/oneshot.py` | resolve + pass device (2 call sites) | +| `src/dltb/iterate.py` | resolve + pass device (2 call sites) | +| `src/dltb/continuous.py` | resolve + pass device (2 call sites) | +| `scripts/{smoke,sweep,sweep-klein,sweep-klein-mask,sweep-prompt}.sh` | MPS-aware preflight + comments | +| `README.md`, `AGENTS.md`, `NOTES.md` | MPS documentation | +| `pyproject.toml`, `uv.lock`, `scripts/bundle.sh`, `scripts/setup-pod.sh` | **unchanged on purpose** | -- 2.51.2