diff --git a/doc/temp/macos.md b/doc/temp/macos.md new file mode 100644 index 0000000..906fbd4 --- /dev/null +++ b/doc/temp/macos.md @@ -0,0 +1,82 @@ +Good news: the port is small. PyTorch on Apple Silicon ships with a built-in GPU backend (**MPS**, built on Metal) [^1], so you don't need CUDA, and Hugging Face `diffusers` supports it directly. You mostly change the device string and adjust a few memory/perf options. + +## 1. Install (macOS, Apple Silicon) + +```bash +# default PyPI wheels for macOS arm64 include MPS support +pip install torch torchvision +pip install diffusers transformers accelerate safetensors +``` + +The latest stable PyTorch supports MPS on Apple Silicon Macs running macOS 14.0 or later [^2]. You can sanity-check with: + +```python +import torch +print(torch.backends.mps.is_available()) # -> True +``` + +## 2. Code changes + +The official SD-Turbo usage is [^3]: + +```python +from diffusers import AutoPipelineForText2Image +import torch + +pipe = AutoPipelineForText2Image.from_pretrained("stabilityai/sd-turbo", torch_dtype=torch.float16, variant="fp16") +pipe.to("cuda") + +image = pipe(prompt=prompt, num_inference_steps=1, guidance_scale=0.0).images[0] +``` + +The Apple Silicon version is the same, with `"mps"` instead of `"cuda"` and attention slicing enabled: + +```python +from diffusers import AutoPipelineForText2Image +import torch + +device = "mps" +pipe = AutoPipelineForText2Image.from_pretrained( + "stabilityai/sd-turbo", + torch_dtype=torch.float16, + variant="fp16", +) +pipe.to(device) +pipe.enable_attention_slicing() # recommended on Apple Silicon + +prompt = "A cinematic shot of a baby raccoon wearing an intricate italian priest robe." +image = pipe(prompt=prompt, num_inference_steps=1, guidance_scale=0.0).images[0] +image.save("out.png") +``` + +Notes that carry over unchanged: SD-Turbo ignores `guidance_scale` (you disable it with `0.0`) and `negative_prompt`, and a **single step** is enough [^3]. If your CUDA code loops over steps 1–4 for quality, that works identically here. + +If your existing script has CUDA-specific bits, map them like this: + +| CUDA-side construct | Apple Silicon replacement | +|---|---| +| `torch.device("cuda")` / `pipe.to("cuda")` | `torch.device("mps")` / `pipe.to("mps")` | +| `torch.cuda.amp.autocast()` | `torch.autocast("mps")` (usually unnecessary — pass `torch_dtype=float16` at load) | +| `torch.cuda.empty_cache()` | `torch.mps.empty_cache()` | +| `xformers` / FlashAttention (`enable_xformers_memory_efficient_attention`) | Remove it — not available on MPS; diffusers falls back to scaled-dot-product attention, which uses custom Metal kernels [^4] | +| `torch.compile(...)` | Remove it — MPS support is an early prototype and end-to-end use "is likely to fail" [^5] | + +## 3. Gotchas specific to MPS + +- **Attention slicing**: diffusers recommends `pipe.enable_attention_slicing()` for machines with < 64 GB RAM, or when generating at resolutions above 512×512 — it reduces memory pressure and prevents swapping [^6]. SD-Turbo is small (~2.5 GB fp16), so if you're on a 32 GB+ machine at 512×512 it's optional, but harmless to enable. +- **Batching**: generating multiple prompts in a batch can crash or fail unreliably on MPS — iterate over prompts instead [^6]. +- **Tensor size limit**: the MPS backend doesn't support arrays larger than 2³² elements [^6] — irrelevant at 512×512 but a reason some ops bail at large upscale sizes. +- **Ops not implemented**: if you hit `...not implemented for the mps backend` errors while running custom code (e.g., custom VAE decode or post-processing), set `PYTORCH_ENABLE_MPS_FALLBACK=1` in the environment so PyTorch silently runs those ops on CPU. +- **First run**: PyTorch ≥ 2.x doesn't need the old "priming" pass that PyTorch 1.13 required on MPS [^6]. +- **Speed**: expect roughly comparable-to-slower-than-CUDA latency per step (a consumer NVIDIA card is faster); the difference matters far less for SD-Turbo than for full SD since it's 1 step by default. First generation after load is slow due to model download + kernel warmup; subsequent runs are much faster. + +That's the whole port: swap the device to `"mps"`, drop any CUDA-only speedups (xformers, `torch.compile`), and add attention slicing. + +**References** + +[^1]: [MPS backend — PyTorch main documentation](https://docs.pytorch.org/docs/main/notes/mps.html) (7%) +[^2]: [Accelerated PyTorch training on Mac - Metal](https://developer.apple.com/metal/pytorch/) (15%) +[^3]: [stabilityai/sd-turbo · Hugging Face](https://huggingface.co/stabilityai/sd-turbo) (13%) +[^4]: [GitHub - jhurt/attention-mps-torch: A custom PyTorch operator for ...](https://github.com/jhurt/attention-mps-torch) (6%) +[^5]: [torch.compile on MPS progress tracker · Issue #150121...](https://github.com/pytorch/pytorch/issues/150121) (8%) +[^6]: [Metal Performance Shaders (MPS) · Hugging Face](https://huggingface.co/docs/diffusers/en/optimization/mps) (51%)