diff --git a/AGENTS.md b/AGENTS.md index ce438f6..6b28338 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -9,7 +9,7 @@ their own output: free-running image self-iteration, video pipeline simulation with carried state, and failure tails. Research/hobby code, not a product. Two domains exist: -- **Local (macOS)**: editing, bundling (`just bundle`), image building. No GPU +- **Local (macOS)**: editing, bundling (`just bundle`). No GPU needed for `--help`/arg checks (heavy imports are deliberately lazy). - **Pod (Runpod, linux/amd64)**: actual GPU runs via the bundle workflow. @@ -20,15 +20,15 @@ and design sketches (e.g. klein dual-reference conditioning, not yet implemented ## Commands Package management is `uv` only (Python 3.13, CUDA-enabled torch comes via -`uv sync`; there is no requirements.txt). Never edit `uv.lock` casually; the -Dockerfile bakes it and rebuilds are only needed when it changes. +`uv sync`; there is no requirements.txt). `uv.lock` ships in every bundle and +is installed pod-side by `scripts/setup-pod.sh`, so lock changes ride the +normal bundle workflow (there is no custom image to rebuild). ```bash uv sync # set up env uv run dltb-oneshot --help # cheap local sanity check (no CUDA) uv run dltb-iterate --model sd-turbo --input input_example/test_512.png --iterations 3 just bundle # stage pod-ready tarball in bundle/ -just image-build /: # build pod image (linux/amd64) just hf-status | hf-keep | hf-clean # HF cache management ``` @@ -113,8 +113,8 @@ Key invariants: verification guard; never verify with `tar -t`, it hides them). - Bundles ship **tracked files only**, with one exception: gitignored `input/` (user inputs) is packed explicitly by `bundle.sh`. `git add` new scripts - before `just bundle`, or the pod silently misses them (both scripts warn); - image builds remain tracked-only. + before `just bundle`, or the pod silently misses them (the bundle script + warns). - `.gitignore` covers `output/`, `bundle/`, `._*`, `.DS_Store`, `.pi/` — keep generated artifacts out of git. @@ -124,13 +124,17 @@ Key invariants: just bundle # local runpodctl send bundle/imgiter-.tar.gz # on pod: -cd /workspace && tar xzf imgiter-.tar.gz && cd imgiter- -uv sync --frozen # re-points baked venv (/opt/imgiter/.venv) at new code +cd /workspace && mkdir -p imgiter +tar xzf imgiter-.tar.gz --strip-components=1 -C imgiter && cd imgiter +scripts/setup-pod.sh # pinned uv + uv sync --frozen (re-run per bundle) scripts/smoke.sh && scripts/sweep.sh ``` -- Pod image (`image/Dockerfile`) = runpod/base + uv-locked deps; rebuild only - when `uv.lock` changes. Code ships via bundles. +- Pod image = stock `runpod/base` pinned by digest — no custom image: Runpod + bills from the start of the image pull, so baking deps buys nothing + (NOTES.md, 2026-09-13). `scripts/setup-pod.sh` installs uv 0.12.13 + the + locked deps on the pod; bundles extract into the fixed dir + `/workspace/imgiter` so `.venv` and outputs survive new bundle extracts. - `just` and `runpodctl` are NOT in the pod image — use `scripts/*.sh` directly. - 48 GB VRAM recommended; `--offload` for the two big models on smaller cards. - Container disk is **ephemeral on stop AND restart** — copy `output/` out diff --git a/Justfile b/Justfile index 52e0211..8bc550b 100644 --- a/Justfile +++ b/Justfile @@ -9,11 +9,6 @@ bundle: clean: rm -rf bundle/* -# Build the pod environment image (deps + venv baked, code still ships via `bundle`). -# The tag is your Docker Hub push target, e.g.: just image-build /imgiter:1 -image-build tag: - scripts/image-build.sh {{tag}} - # Inspect the HuggingFace model cache (per-repo sizes, filesystem free) hf-status: scripts/hf-cache.sh status diff --git a/NOTES.md b/NOTES.md index eda7790..bcf7fe7 100644 --- a/NOTES.md +++ b/NOTES.md @@ -373,3 +373,43 @@ table so the price/compute trade is explicit rather than implicit. The appendix VRAM probe is the cheap way to get peak VRAM; a fixed `--max-frames` run gives the throughput denominator. + +## Pod image retired: stock runpod/base + `scripts/setup-pod.sh` (2026-09-13) + +**Decision:** stop building `refinementsystems/imgiter`. Pods run the stock +`runpod/base:1.3.0-rc.164-ubuntu2404`, pinned by the same digest the custom +image was built `FROM`; `scripts/setup-pod.sh` (new, ships in every bundle) +installs uv 0.12.13 into `/usr/local/bin` and runs `uv sync --frozen`. + +**Why:** Runpod starts billing when the container image pull starts. The +~12 GB custom image was justified as skipping the multi-GB `uv sync` on boot, +but it instead added a billed ~12 GB Docker Hub pull (often throttled) — the +~6 GB PyPI sync it skipped is cheaper, faster, and paid only when the lock +actually changes. Secondary: every `uv.lock` change forced an emulated +linux/amd64 rebuild + Docker Hub push + template digest re-pin; now lock +changes ride the normal bundle workflow. For scale, both sides are dwarfed by +the ~87.5 GB of HF model downloads every fresh pod pays anyway (container +disk is wiped on stop and restart), so the whole optimization was noise. + +**Mechanics that replaced the baked venv:** + +- Bundles extract into a **fixed dir** (`/workspace/imgiter`, + `--strip-components=1`), so the project `.venv` (uv's default location, + created by `uv sync`) and `output/` survive new bundle extracts; re-running + `scripts/setup-pod.sh` after each extract re-points the editable install + (seconds when `uv.lock` is unchanged). This replaces the image's + `UV_PROJECT_ENVIRONMENT=/opt/imgiter/.venv` env var, which cannot be + provided container-wide without a custom image — and without it, stamp-dir + extraction would re-download the stack once per bundle. +- `scripts/inputs.sh` (sourced by every driver) now fails fast when `.venv` + is missing, so `uv run` can never silently sync a fresh multi-GB venv + mid-sweep. `DRY_RUN=1` previews bypass the guard. +- uv 0.12.13 is pinned to match the lockfile producer and the `uv_build` + backend constraint (`>=0.12.7,<0.13.0`); installed from the GitHub release + tarball (no `curl | sh`). + +**Not changed:** the pod template (id `04u1mmp8nf`) keeps its disk/env/ports; +its image reference needs a one-time `runpodctl template update --image` +re-point (README_RUNPOD.md §3). The old image tags stay on Docker Hub. +`hf-cache.sh`'s `HF_HOME` guard and the SSH/Jupyter behavior are base-image +features, unaffected. diff --git a/README_DEFERRED.md b/README_DEFERRED.md index c819e75..a7c0569 100644 --- a/README_DEFERRED.md +++ b/README_DEFERRED.md @@ -1,17 +1,7 @@ -# Deferred / rejected image decisions +# Deferred / rejected pod decisions -Parked changes to the pod image (`image/Dockerfile`, published as -`refinementsystems/imgiter` on Docker Hub). Nothing here is scheduled; items -under *Rejected* are decided unless new information arrives. - -Ground rules for any future rebuild: - -- The pod template pins the image **by digest** and published tags are never - mutated, so existing pods and other users cannot break. A changed image - ships as a **new tag** (`0.2.0` for dependency/behavior changes, `0.1.1` for - metadata-only), plus a `runpodctl template update` of the image reference. -- Prefer bundling cosmetic changes into the next rebuild that is forced anyway - (i.e. whenever `uv.lock` changes). +Parked or rejected ideas around the pod setup. Nothing here is scheduled; +items under *Rejected* are decided unless new information arrives. ## Parked @@ -20,22 +10,16 @@ Ground rules for any future rebuild: NOTES.md, "Runtime log noise" item 3: without torchvision, transformers falls back to `CLIPImageProcessorPil` — slightly slower preprocessing, same results. Adding it to `pyproject.toml`/`uv.lock` switches the backend back, so runs -from before/after must not be mixed (reproducibility); that makes it an image -`0.2.0`. Do it only if preprocessing ever becomes a bottleneck or a -correctness question. - -### OCI labels + `EXPOSE 22` (cosmetic) +from before/after must not be mixed (reproducibility); ship it with the next +`uv.lock` change that happens anyway. Do it only if preprocessing ever becomes +a bottleneck or a correctness question. -`org.opencontainers.image.{source,description,licenses}` (ISC) for the public -Docker Hub page; `EXPOSE 22` documents the SSH port the template maps. -Trivial — ride along with the next rebuild. +### Re-pin the pod image off the rc tag -### Re-pin the base image off the rc tag - -`runpod/base:1.3.0-rc.164-ubuntu2404` is pinned by digest (immutable), so -this is not urgent; when a stable `1.3.0`+ tag exists, re-pin during the next -rebuild and re-check the uv version constraint (`>=0.12.7,<0.13.0`) at the -same time. +The pod image is stock `runpod/base:1.3.0-rc.164-ubuntu2404`, pinned by digest +(immutable), so this is not urgent; when a stable `1.3.0`+ tag exists, re-pin +the template / `pod create` image reference and re-check the uv version pinned +in `scripts/setup-pod.sh` (constraint `>=0.12.7,<0.13.0`) at the same time. ### Exercise the Jupyter option @@ -45,24 +29,15 @@ or drop the section. ## Rejected -### Bake SSH host keys / authorized_keys into the image - -A public image would give every user's pods identical host keys -(impersonation risk), and baked `authorized_keys` pin one person's key into a -public artifact. Unnecessary anyway: the base image's `/start.sh` generates -host keys at container start and builds `authorized_keys` from `$PUBLIC_KEY` -(account keys injected by Runpod at pod start). The 2026-09-12 "no sshd" -incident (NOTES.md) was a missing registered key at first boot, not an image -defect — see README_RUNPOD.md §3, "SSH access". - -### Bake `runpodctl` into the image - -rsync over direct SSH (automatic once account keys are registered, ~14 MB/s -measured) is the preferred transfer path; a baked CLI would duplicate it and -drift out of date. - -### Bake model weights or the HF token into the image - -~87.5 GB of weights, one model is gated (needs a per-pod `HF_TOKEN` anyway), -and the keep-all cache policy on the 150 GB disk already works -(README_RUNPOD.md §5). +### Build a custom pod image (retired 2026-09-13) + +The first commissioning baked the locked venv into a ~12 GB custom image +(published as `refinementsystems/imgiter` on Docker Hub; the old tags remain +there, unchanged) to skip the multi-GB `uv sync` on pod boot. Retired because +Runpod starts billing when the image pull starts: the ~12 GB pull (from Docker +Hub, often throttled) cost more billed GPU time than the ~6 GB PyPI sync it +skipped — and every `uv.lock` change forced an emulated linux/amd64 rebuild, a +push, and a template digest re-pin. The stock base image + `scripts/setup-pod.sh` +(README_RUNPOD.md §4) does the same job with no build train. For scale: the +~87.5 GB of HF model downloads every fresh pod pays dwarfs both sides of the +trade anyway. diff --git a/README_RUNPOD.md b/README_RUNPOD.md index d834e5a..cd44481 100644 --- a/README_RUNPOD.md +++ b/README_RUNPOD.md @@ -1,8 +1,11 @@ # imgiter on Runpod — runbook -Practical notes from commissioning and testing the `imgiter` pod template -(`refinementsystems/imgiter`, tag `0.1.0`). Everything here was measured on a -real pod; prices and availability drift, so re-check the console. +Practical notes from commissioning and running `imgiter` pods. Everything +here was measured on real pods; prices and availability drift, so re-check +the console. Pods run the stock `runpod/base` image pinned by digest (§3); +the dependency-baked custom image used for the first commissioning was +retired 2026-09-13 — Runpod bills from the start of the image pull, so it +cost more than the `uv sync` it skipped (NOTES.md). ## TL;DR @@ -21,9 +24,10 @@ git add scripts/hf-cache.sh scripts/sweep.sh Justfile README_RUNPOD.md # untrac just bundle # pod -cd /workspace && tar xzf imgiter-.tar.gz && cd imgiter- -uv sync --frozen -scripts/smoke.sh # post-deploy check (all three tools) +cd /workspace && mkdir -p imgiter +tar xzf imgiter-.tar.gz --strip-components=1 -C imgiter && cd imgiter +scripts/setup-pod.sh # pinned uv + uv sync --frozen (once per pod; re-run per bundle) +scripts/smoke.sh # post-deploy check (all tools) scripts/sweep.sh ``` @@ -81,18 +85,22 @@ the number of frames, not steps. ``` /workspace <- container disk (ephemeral), 150 GB - imgiter-/ <- extracted bundle (code only; no .venv) + imgiter/ <- extracted bundle (fixed dir, reused across bundles) + .venv/ <- locked deps (scripts/setup-pod.sh; ~7 GB) output/ <- results, sweep logs .cache/huggingface/ <- HF_HOME (base image default) hub/models----/ <- model weights xet/ <- xet chunk cache, hard-capped at 10 GB + .cache/uv/ <- UV_CACHE_DIR (base image default): wheel cache ``` -Dependencies live in the image at `/opt/imgiter/.venv` -(`UV_PROJECT_ENVIRONMENT`), not in the extracted tree. +Dependencies live in `/workspace/imgiter/.venv`, created by +`scripts/setup-pod.sh` (it installs uv 0.12.13 and provisions CPython 3.13 — +`runpod/base` ships neither). The fixed extraction dir keeps `.venv` and +`output/` alive across bundle updates. -The image itself does **not** count against the container disk: `df /workspace` -showed ~85 MB used with the 12 GB image present. +The image itself does **not** count against the container disk (`df +/workspace` showed ~85 MB used next to a 12 GB image). ### Measured cache sizes (real downloads, not repo totals) @@ -155,12 +163,14 @@ The `imgiter` template is **private and stays that way on purpose**: it wires account-internal secrets by name (`{{ RUNPOD_SECRET_HF_TOKEN }}`), and other users have no reason to use the same secret names (no leak risk either way — secrets resolve per account — but a public template would simply not work for -others). The **image** is public (`refinementsystems/imgiter` on Docker Hub), -so anyone can run the stack without the template: +others). The image is the **stock** `runpod/base`, pinned by digest — the +same base the retired custom image was built `FROM`, so the pod plumbing +(`/start.sh` sshd, `/workspace` cache conventions) is unchanged. Anyone can +run the stack without the template: ```bash runpodctl pod create \ - --image refinementsystems/imgiter:0.1.0 \ + --image runpod/base:1.3.0-rc.164-ubuntu2404@sha256:95357957d7660542b37226fb31863c4510502006f2018efe23e2869137daf000 \ --gpu-id "NVIDIA RTX 6000 Ada" \ --container-disk-in-gb 150 \ --ports "22/tcp" \ @@ -173,8 +183,8 @@ of this section documents the author's template as the reference configuration. Template `imgiter` (`04u1mmp8nf`): container disk 150 GB, no volume, env -`HF_TOKEN={{ RUNPOD_SECRET_HF_TOKEN }}`, port `22/tcp`, image pinned by digest -`sha256:68d934…` (tag `0.1.0`). +`HF_TOKEN={{ RUNPOD_SECRET_HF_TOKEN }}`, port `22/tcp`, image = the +`runpod/base` digest above. ```bash runpodctl pod create \ @@ -201,6 +211,7 @@ the template if you need to change it): ```bash runpodctl template update 04u1mmp8nf \ + --image runpod/base:1.3.0-rc.164-ubuntu2404@sha256:95357957d7660542b37226fb31863c4510502006f2018efe23e2869137daf000 \ --container-disk-in-gb 150 \ --env '{"HF_TOKEN":"{{ RUNPOD_SECRET_HF_TOKEN }}"}' ``` @@ -268,17 +279,22 @@ has the URL. ```bash cd /workspace -tar xzf imgiter-.tar.gz -cd imgiter- -uv sync --frozen # re-points the editable install from /opt/imgiter to this tree +mkdir -p imgiter # fixed dir: .venv and output/ survive bundle updates +tar xzf imgiter-.tar.gz --strip-components=1 -C imgiter +cd imgiter +scripts/setup-pod.sh # uv 0.12.13 + uv sync --frozen; re-run after every bundle scripts/sweep.sh # full sweep; add OFFLOAD=1 on <48 GB GPUs for the big models scripts/sweep-klein.sh # klein prompt ladder + steps probes + reproject A/B (single model) scripts/sweep-prompt.sh # free-running prompt (x strength) sweep on one image ``` -- Dependencies are baked into the image at `/opt/imgiter/.venv` - (`UV_PROJECT_ENVIRONMENT`), so `uv sync` only rebuilds the project itself. +- Dependencies install into `imgiter/.venv` on first `scripts/setup-pod.sh` + (~6 GB of wheels + CPython 3.13, a couple of minutes from PyPI). Re-running + it after a new bundle extract re-points the editable install — seconds when + `uv.lock` is unchanged. Extracting over the fixed dir can leave files that + were deleted from the repo lingering in the tree; harmless, or wipe the + tree (`.venv` included) to fully reset. - **`just` is not installed in the image** — call `scripts/*.sh` directly, or use `uv run dltb-oneshot|dltb-iterate|dltb-continuous --help` for individual runs. - `runpodctl` is not in the image either; install it on the pod @@ -389,7 +405,8 @@ disk. the template. 4. **Volume-size changes are impossible via `runpodctl template update`**; only the console or a template recreate. -5. **`just` and `runpodctl` are absent from the image.** Use the shell scripts. +5. **`just`, `runpodctl` and `uv` are absent from the image.** Use the shell + scripts; `scripts/setup-pod.sh` installs uv (pinned 0.12.13) once per pod. 6. `--num-inference-steps 4 --strength 0.4` (the sweep defaults) results in **1 actual denoise step** for `sd-turbo` / `sdxl-turbo` / `flux-schnell` (`int(4 × 0.4)`); the `flux2-klein-*` models run all 4 steps @@ -400,14 +417,19 @@ disk. so it is not corruption; treat per-model peak as ≤ ~35 GB either way. 8. Cost example: RTX 5090 at $0.99/hr, whole disk test (image pull + ~95 GB of downloads + a few runs) was well under an hour. -9. Downloads from Hugging Face have no egress cost; the only cost is GPU time - while downloading. +9. Pod billing starts when the **image pull** starts — a key reason the custom + baked-deps image was retired for stock `runpod/base` + `setup-pod.sh` + (2026-09-13, NOTES.md): a billed ~12 GB Docker Hub pull to skip a ~6 GB + PyPI sync is a net loss, and fresh pods re-download wheels and models + anyway (container disk is wiped on stop and restart). +10. Downloads from Hugging Face have no egress cost; the only cost is GPU time + while downloading. ## Appendix: quick VRAM probe Runs one pass per model in a fresh process, reporting native fit, peak VRAM and -whether `--offload` is needed. Run it after `uv sync`, from the extracted repo -root (`input_example/test_768.png` ships in every bundle): +whether `--offload` is needed. Run it after `scripts/setup-pod.sh`, from the +extracted repo root (`input_example/test_768.png` ships in every bundle): ```bash for m in sd-turbo sdxl-turbo flux2-klein-4b flux-schnell flux2-klein-9b; do diff --git a/image/Dockerfile b/image/Dockerfile deleted file mode 100644 index 380b827..0000000 --- a/image/Dockerfile +++ /dev/null @@ -1,58 +0,0 @@ -# imgiter pod image: Runpod base + the uv-locked python environment baked in. -# -# Why bake it: a fresh pod skips the multi-GB `uv sync` (torch/diffusers stack) -# that used to run on pod boot. Code is NOT frozen into the image — the normal -# workflow (`just bundle` + `runpodctl send`) still delivers it. Extracting a -# bundle on the pod and re-running `uv sync --frozen` upgrades the code in -# seconds: UV_PROJECT_ENVIRONMENT below pins every `uv sync`/`uv run` to the -# baked venv, so only the project itself (editable install) gets re-pointed, -# never the dependencies. Rebuild the image only when uv.lock changes. -# -# Base: runpod/base (no CUDA toolkit, no torch) rather than runpod/pytorch. -# - The toolkit is unnecessary: nothing compiles CUDA here, and the locked -# torch 2.14 linux wheels bundle their whole CUDA 13 userspace via the -# nvidia-*/cuda-* PyPI wheels inside the venv (that is also how the earlier -# bundle+`uv sync` runs worked — they never used the pod's system CUDA). -# Only the host driver must match, and that is injected per-pod by the -# NVIDIA container toolkit regardless of base image. -# - runpod/pytorch would additionally ship a system torch that the venv's -# torch 2.14 shadows — ~4 GB of dead, confusing weight. -# - runpod/base keeps the pod plumbing we DO need: /start.sh (SSH etc., kept -# by not overriding CMD) and the /workspace cache conventions (HF_HOME, -# UV_CACHE_DIR, HF_XET_HIGH_PERFORMANCE — already set in the base env). -# Pinned by digest: the tag is an rc and Runpod may mutate or drop it. -FROM docker.io/runpod/base:1.3.0-rc.164-ubuntu2404@sha256:95357957d7660542b37226fb31863c4510502006f2018efe23e2869137daf000 - -# uv, pinned to the version the lockfile was produced with (must also satisfy -# the uv_build backend constraint in pyproject.toml: >=0.12.7,<0.13.0). -# Static binaries — no python needed in the base at this point. -COPY --from=ghcr.io/astral-sh/uv:0.12.13 /uv /uvx /usr/local/bin/ - -# All `uv sync` / `uv run` invocations — build-time and pod-side — resolve to -# this one venv, instead of a project-local .venv (which, in a bundle extracted -# under /workspace, would start empty and re-download the stack). -ENV UV_PROJECT_ENVIRONMENT=/opt/imgiter/.venv - -WORKDIR /opt/imgiter - -# Lock-first layering: copy only the dependency inputs so this layer caches -# across code-only changes. README.md is required by uv_build (readme field), -# .python-version pins the interpreter (uv provisions CPython 3.13 itself — -# the base image ships no python 3.13). -COPY pyproject.toml uv.lock README.md .python-version ./ -COPY src ./src - -# --frozen : uv.lock is the source of truth; never re-resolve in-image. -# --no-dev : no dev deps exist today; keeps the flag honest if they appear. -# --compile-bytecode : faster interpreter startup in the image. -# UV_CACHE_DIR is redirected off the base's /workspace default for this RUN: -# a BuildKit cache mount keeps the ~6 GB of downloaded wheels out of the final -# image while still warming rebuilds on the build host. -RUN --mount=type=cache,target=/tmp/uv-cache \ - UV_CACHE_DIR=/tmp/uv-cache uv sync --frozen --compile-bytecode --no-dev - -# The imgiter CLI (and venv tools) on PATH. After a bundle overlay + `uv sync`, -# this same entry point runs the freshly sent code. -ENV PATH="/opt/imgiter/.venv/bin:${PATH}" - -# no CMD override: inherit /start.sh (SSH + pod bootstrap) from the base image. diff --git a/scripts/hf-cache.sh b/scripts/hf-cache.sh index b4520a2..17e4960 100755 --- a/scripts/hf-cache.sh +++ b/scripts/hf-cache.sh @@ -19,8 +19,8 @@ # # Model keys are the same values as the tools' --model flag; the key -> repo # id mapping is read from dltb.models.MODELS so it cannot drift from the -# tools. The command is run through `uv run --frozen`, i.e. the baked venv on -# a pod. +# tools. The command is run through `uv run --frozen`, i.e. the project +# `.venv` on a pod (created by `scripts/setup-pod.sh`). # # Usage: # scripts/hf-cache.sh status # sizes per repo, filesystem free diff --git a/scripts/image-build.sh b/scripts/image-build.sh deleted file mode 100755 index 3f72477..0000000 --- a/scripts/image-build.sh +++ /dev/null @@ -1,59 +0,0 @@ -#!/usr/bin/env bash - -# Permission to use, copy, modify, and/or distribute this software for -# any purpose with or without fee is hereby granted. -# -# THE SOFTWARE IS PROVIDED “AS IS” AND THE AUTHOR DISCLAIMS ALL -# WARRANTIES WITH REGARD TO THIS SOFTWARE INCLUDING ALL IMPLIED WARRANTIES -# OF MERCHANTABILITY AND FITNESS. IN NO EVENT SHALL THE AUTHOR BE LIABLE -# FOR ANY SPECIAL, DIRECT, INDIRECT, OR CONSEQUENTIAL DAMAGES OR ANY -# DAMAGES WHATSOEVER RESULTING FROM LOSS OF USE, DATA OR PROFITS, WHETHER IN -# AN ACTION OF CONTRACT, NEGLIGENCE OR OTHER TORTIOUS ACTION, ARISING OUT -# OF OR IN CONNECTION WITH THE USE OR PERFORMANCE OF THIS SOFTWARE. - -set -euo pipefail - -# Build the imgiter pod image (see image/Dockerfile). Usage: -# scripts/image-build.sh //: -# The tag is chosen by the caller; push it yourself afterwards: -# docker push - -tag="${1:?usage: image-build.sh //:}" - -temp_dir="$(mktemp -d)" -trap 'rm -rf "${temp_dir}"' EXIT -stage="${temp_dir}/ctx" -mkdir -p "${stage}" - -# Tracked files only, copied from the working tree so images match the repo -# (same policy as scripts/bundle.sh). --no-xattrs keeps macOS xattrs out of -# the build context; the Dockerfile itself is copied explicitly so an -# uncommitted edit still builds. -git ls-files -z | tar --no-xattrs --null -T - -cf - | tar -xf - -C "${stage}" -cp image/Dockerfile "${stage}/Dockerfile" - -# Guard against building something broken. -for f in pyproject.toml uv.lock \ - src/dltb/models.py src/dltb/imaging.py src/dltb/output.py src/dltb/args.py \ - src/dltb/oneshot.py src/dltb/iterate.py src/dltb/continuous.py src/dltb/klein.py \ - Dockerfile; do - [[ -f "${stage}/$f" ]] || { echo "image-build: missing $f" >&2; exit 1; } -done - -# Untracked files are NOT baked in — surface them so nothing silently gets -# left behind. -untracked="$(git ls-files --others --exclude-standard)" -if [[ -n "${untracked}" ]]; then - echo "image-build: warning — untracked files are NOT included:" - printf '%s\n' "${untracked}" | sed 's/^/ /' -fi - -if ! git diff --quiet HEAD 2>/dev/null; then - echo "image-build: note — building with uncommitted changes to tracked files." -fi - -# linux/amd64 explicitly: builds correctly (under emulation) on Apple Silicon. -docker buildx build --platform linux/amd64 -t "${tag}" "${stage}" - -echo -echo "Next: docker push ${tag}" diff --git a/scripts/inputs.sh b/scripts/inputs.sh index 7e06bb5..133a002 100755 --- a/scripts/inputs.sh +++ b/scripts/inputs.sh @@ -31,6 +31,15 @@ # Callers still apply their own ${VAR:-...} fallbacks after sourcing, so a # missing conf file only costs the built-in defaults. +# Guard: every driver invokes the tools via `uv run`, which would silently +# create and sync a fresh project venv (~6 GB on a pod) if none exists yet. +# Fail fast instead; DRY_RUN=1 previews bypass this. +if [[ "${DRY_RUN:-0}" != "1" && ! -d .venv ]]; then + echo "inputs: no .venv in ${PWD} -- run scripts/setup-pod.sh first" >&2 + echo " (locally: uv sync)" >&2 + exit 1 +fi + if [[ -f input/inputs.env ]]; then source input/inputs.env elif [[ -f input_example/inputs.env ]]; then diff --git a/scripts/setup-pod.sh b/scripts/setup-pod.sh new file mode 100755 index 0000000..e8fa41e --- /dev/null +++ b/scripts/setup-pod.sh @@ -0,0 +1,85 @@ +#!/usr/bin/env bash + +# Permission to use, copy, modify, and/or distribute this software for +# any purpose with or without fee is hereby granted. +# +# THE SOFTWARE IS PROVIDED “AS IS” AND THE AUTHOR DISCLAIMS ALL +# WARRANTIES WITH REGARD TO THIS SOFTWARE INCLUDING ALL IMPLIED WARRANTIES +# OF MERCHANTABILITY AND FITNESS. IN NO EVENT SHALL THE AUTHOR BE LIABLE +# FOR ANY SPECIAL, DIRECT, INDIRECT, OR CONSEQUENTIAL DAMAGES OR ANY +# DAMAGES WHATSOEVER RESULTING FROM LOSS OF USE, DATA OR PROFITS, WHETHER +# IN AN ACTION OF CONTRACT, NEGLIGENCE OR OTHER TORTIOUS ACTION, ARISING +# OUT OF OR IN CONNECTION WITH THE USE OR PERFORMANCE OF THIS SOFTWARE. + +# +# setup-pod.sh -- pod bootstrap: pinned uv + the locked dependency venv. +# +# The pod runs the stock runpod/base image, pinned by digest (README_RUNPOD.md +# §3) -- no custom image is built: Runpod starts billing when the image pull +# starts, so a ~12 GB baked-deps image cost more billed pull time than the +# ~6 GB `uv sync` it was meant to skip (NOTES.md, 2026-09-13). runpod/base +# ships neither uv nor Python 3.13, so this script installs the one and +# provisions the other: +# +# 1. uv 0.12.13 -- the version uv.lock was produced with, and inside the +# uv_build backend constraint (>=0.12.7,<0.13.0) -- from the GitHub +# release tarball into /usr/local/bin (static binaries, no curl|sh). +# Skipped when that exact version is already installed. +# 2. `uv sync --frozen` -- creates .venv in this tree, downloading CPython +# 3.13 (per .python-version) and the locked torch stack (~6 GB of +# wheels, a couple of minutes from PyPI; the wheels bundle the whole +# CUDA userspace, so only the host driver version matters). +# +# Run from the extracted bundle root right after extraction, and re-run after +# every new bundle extract over the same tree (idempotent; with an unchanged +# uv.lock the re-run only re-points the editable install, in seconds): +# +# cd /workspace && mkdir -p imgiter +# tar xzf imgiter-.tar.gz --strip-components=1 -C imgiter +# cd imgiter && scripts/setup-pod.sh +# +# Linux only. The wheel cache lands under UV_CACHE_DIR (runpod/base default: +# /workspace/.cache/uv, i.e. the ephemeral container disk) -- expect a fresh +# sync on every new pod, same as for the HF model cache. + +set -euo pipefail + +UV_VERSION="0.12.13" + +if [[ "$(uname -s)" != "Linux" ]]; then + echo "setup-pod: for the Runpod pod (Linux), not $(uname -s)" >&2 + exit 1 +fi + +cd "$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" + +# 1. uv, pinned to the lockfile's version. +if command -v uv >/dev/null 2>&1 && [[ "$(uv --version)" == "uv ${UV_VERSION} "* ]]; then + echo "setup-pod: uv ${UV_VERSION} already at $(command -v uv)" +else + [[ ${EUID} -eq 0 ]] || { echo "setup-pod: need root to install uv into /usr/local/bin" >&2; exit 1; } + command -v curl >/dev/null || { echo "setup-pod: curl not found" >&2; exit 1; } + + case "$(uname -m)" in + x86_64) uv_arch="x86_64-unknown-linux-gnu" ;; + aarch64) uv_arch="aarch64-unknown-linux-gnu" ;; + *) echo "setup-pod: unsupported architecture $(uname -m)" >&2; exit 1 ;; + esac + + tmp="$(mktemp -d)" + trap 'rm -rf "${tmp}"' EXIT + echo "setup-pod: installing uv ${UV_VERSION} (${uv_arch}) -> /usr/local/bin" + curl -fL "https://github.com/astral-sh/uv/releases/download/${UV_VERSION}/uv-${uv_arch}.tar.gz" \ + | tar xz -C "${tmp}" + install -m 0755 "${tmp}/uv-${uv_arch}/uv" /usr/local/bin/uv + if [[ -f "${tmp}/uv-${uv_arch}/uvx" ]]; then + install -m 0755 "${tmp}/uv-${uv_arch}/uvx" /usr/local/bin/uvx + fi +fi + +# 2. The locked environment into ./.venv (default project venv location, so +# plain `uv run` from any later shell finds it without extra env vars). +echo "setup-pod: uv sync --frozen (first run on a fresh pod: ~6 GB of wheels)" +uv sync --frozen --no-dev + +echo "setup-pod: ready -- venv at ${PWD}/.venv; next: scripts/smoke.sh"