diff --git a/13-inch-thin-cannon/docs/iso-build.md b/13-inch-thin-cannon/docs/iso-build.md index e7cb185..391f7bb 100644 --- a/13-inch-thin-cannon/docs/iso-build.md +++ b/13-inch-thin-cannon/docs/iso-build.md @@ -210,15 +210,14 @@ Once SSH is up it verifies: Secure Boot enabled (sbctl), `cryptroot` is LUKS, `/` is btrfs with `/nix`, `/persistent` and `/root-blank`, `/etc/nixos` persisted, kernel cmdline, hostname, no failed services. -**Status (2026-08-13):** the closure-finder blocker is **fixed and boot now -works** — the reinstalled system comes up to a LUKS prompt, unlocks, boots to -the login prompt, and SSH answers (verified: impermanence mounts live, -Secure Boot enabled, `/nix`/`/persistent`/`/root-blank` all present). Two -follow-ups remain before the boot phase is green: (1) an **intermittent -initrd watchdog reboot** on some cold boots (see known issues; root cause in -`kubernetes/docs/debugging.md`) and (2) a few failed services on the installed -system (persisted `/etc/passwd`-family links, home-manager, wireguard). -Previous blocker summary (kept for history): the boot failure was diagnosed via +**Status (2026-08-13, green):** the full pipeline **passes end-to-end** +(`just up`: build → live → install → boot) and repeated cold boots are +consistent (3/3 green). The installed system boots, unlocks LUKS, reaches the +login prompt, and SSH answers with no failed services except the expected +`wireguard-wg0` WARN (age secret undecryptable with the VM's fresh host key — +VM-only, not a config bug). The last blockers were the swapfile systemd +ordering cycle (see known issues) and the earlier failed services, all since +fixed. Previous blocker summary (kept for history): the boot failure was diagnosed via VGA screendump + tesseract OCR (see debugging.md): `initrd-find-nixos-closure` failed because the disk had **no `/nix` subvolume** — the Nix store lived inside `/root`, which the impermanence rollback wipes and replaces with `root-blank` @@ -248,6 +247,28 @@ before the closure finder runs). "subvol=/nix"]`. The `init=` kernel param is confirmed baked into the UKI by lanzaboote (`rust/tool/systemd/src/install.rs:624`), so the closure-finder's other requirement is already satisfied. +- **RESOLVED (2026-08-13): nondeterministic "first boot" SSH death was a + systemd ordering cycle, not a race.** `create-swapfile.service` was ordered + `Before=swap.target` while depending on the `/swap` mount. Because every + service has implicit `After=sysinit.target` default deps, and + `sysinit.target` has `After=swap.target`, that closed a cycle + (`create-swapfile → swap.target → sysinit.target → local-fs.target → + /swap.mount → create-swapfile`). systemd breaks ordering cycles by deleting a + job, and it did so **nondeterministically** — on some boots it dropped + `run-wrappers.mount`, so NixOS's setuid wrappers + (`/run/wrappers/bin/unix_chkpwd`) were never created. sshd then failed every + login at the PAM account stage (`pam_unix(sshd:account): helper binary + execve failed: No such file or directory` → `fatal: Access denied for user + root by PAM account configuration`), making SSH unusable even though the + login prompt appeared. This is why the same install booted fine sometimes + and silently broke on others. **Fix:** drop the systemd swap unit + (`swapDevices`) and `swap.target` ordering entirely; a self-contained + `swapfile.service` (multi-user.target, `RequiresMountsFor=/swap`) creates + the file (guarded by `swaplabel`, not `-e` — a 0-byte placeholder passes an + existence check but fails swapon with "insecure permissions") and runs + `swapon` itself. No ordering edge crosses `swap.target` → no cycle, no job + deletion. Verified: `systemctl show` shows no `Before=swap.target`, and + fresh install + 3× cold boots are all green. - **INTERMITTENT: first cold boot sometimes dies in the initrd with a watchdog reboot (QEMU iTCO artifact, not a config bug).** Serial log shows `kvm_intel: VMX not supported by CPU 7` (~10.8 s), then `watchdog: watchdog0: @@ -272,22 +293,32 @@ before the closure finder runs). The harness must tolerate the extra reboot (and possibly a second LUKS prompt) — `iso-test.sh` handles this with responsive unlock (it watches the serial log for the prompt instead of sending a fixed count). -- **Failed services on the installed system (to fix, not yet root-caused):** - `persist-persistent-etc-passwd.service`, `persist-persistent-etc-group.service`, - `persist-persistent-etc-shadow.service` (persistenced can't bind-mount/link - `/etc/passwd`-family into `/persistent/etc`), `home-manager-file_magic.service`, - `wireguard-wg0.service`, and `swap-swapfile.swap` (`/swap/swapfile` missing). - Likely related to the **`persist-files` activation snippet failing during - `nixos-install`** (the impermanence directories weren't populated at install - time). Also: `/etc/nixos` is **not persisted on the installed system** - (`test -f /etc/nixos/configuration.nix` fails) — it WAS verified working in an - earlier session, so this is a regression or an intermittent persist race. - The `boot` phase now dumps per-unit `journalctl` on failure (see "Start next" - below) to root-cause these. -- **Harness: Secure Boot check has a parsing bug.** `sbctl status` prints - `Secure Boot: ✓ Enabled` (tab + checkmark), but the harness greps for - `"Secure Boot: Enabled"` → always FAIL even though Secure Boot is enabled. - Fix: grep for `Secure Boot:.*Enabled`. +- **Failed services on the installed system — all RESOLVED (2026-08-13):** + - `persist-persistent-etc-{passwd,group,shadow,gshadow}`: impermanence's + `mount-file.bash` (pin `7b1d382f…`) requires an *empty* target file + (`[[ -s $mountPoint ]]` skips), but NixOS writes non-empty `/nix/store` + symlinks into `/etc/{passwd,…}` → bind impossible. Users are declarative + (`mutableUsers = true` in `modules/user.nix`), so these files must **not** + be persisted — removed from `environment.persistence`. (SSH public-key + auth needs `/etc/passwd` only as a symlink to the store, which survives + rollback.) + - `home-manager-file_magic.service`: `.local/state/nix/profiles` and + `.local/state/home-manager` did not exist on first boot → HM activation + failed. Now persisted via `users.file_magic.directories`, and the service + is ordered `after = ["nix-daemon.socket" "nix-daemon.service"]` (first + activation installs a generation through the Nix daemon). + - `swap-swapfile.swap`: see the ordering-cycle bullet above (swapfile + handling moved to a self-contained service). + - `wireguard-wg0.service`: age secret undecryptable with the VM's fresh SSH + host key — expected in the harness and treated as a WARN, not a FAIL. + - `/etc/nixos` not persisted on the installed system: the repo was never + copied into `/persistent/etc/nixos`. Fixed by the installer's + `persist-config` script (mounts the LUKS `/persistent` subvol itself — + disko-install unmounts the target before exiting — and `cp -a`s the baked + `/iso/repo`). Harness now checks `flake.nix`, not `configuration.nix`. +- **Harness: Secure Boot check parsing bug — RESOLVED.** `sbctl status` prints + `Secure Boot: ✓ Enabled` (tab + checkmark); the grep is now + `Secure Boot:.*Enabled`. - **Diagnosis technique that worked:** QEMU monitor `screendump` + tesseract OCR and serial SysRq-trigger to prove kernel liveness — details in `kubernetes/docs/debugging.md` ("Boot-blocker diagnosis without a working @@ -320,48 +351,34 @@ before the closure finder runs). ## Current status & missing work -**Save point (2026-08-13, late):** `just` is installed (`home.nix`), the -justfile now has `build`/`debug`/`live`/`install`/`boot`/`up`/`ovmf`/`help` -targets and a fixed `ovmf` recipe, and `iso-test.sh` accepts `QEMU_EXTRA=...` -for debug flags plus dumps per-unit journals on boot-verification failure. The -watchdog-reboot verdict is in (guest software reboot; see `debugging.md`). -Uncommitted at this save point: `home.nix` (just), `justfile`, `iso-test.sh`, -and the two doc updates (this file + `debugging.md`). - -- **Green:** `live` phase, `install` phase (`disko-install succeeded`, - exit 0, `RAM=10240`). -- **Boot phase: kernel boots, SSH comes up, but verification FAILS.** The - closure-finder fix works. Remaining failures (all reproducible via - `just boot`): - 1. Failed services: `persist-persistent-etc-{passwd,group,shadow}`, - `home-manager-file_magic`, `wireguard-wg0`, `swap-swapfile.swap`. - 2. `/etc/nixos` not persisted on the installed system (was OK earlier — - regression or persist race). - 3. Harness Secure-Boot check bug (see known issues). - 4. Intermittent initrd watchdog self-reboot (harness survives via retry; - real hardware unaffected). +**Green (2026-08-13):** `just up` passes end-to-end (build → live → install → +boot), including **first boot after a fresh install** (the previous flakiest +step). Repeated `just boot` runs are consistent (3/3 green, only the expected +wireguard WARN). Committed in this session: `modules/impermanence.nix` +(`/var/log/journal` persistence for boot diagnostics), installer +`persist-config`, harness robustness, and the four `configuration.nix` fixes +above. + +Remaining work (not blockers): + +- **Watchdog reboot on some cold boots** — the only intermittent noise left. + QEMU/guest-software artifact (see known issues + `kubernetes/docs/debugging.md`); + real hardware has no TCO. Harness auto-retries; a clean cold-boot pass still + shows 0 failed services. Fallback if it ever fails: `boot.blacklistedKernelModules + = ["iTCO_wdt"]` on the host config. - **Deferred hardening:** `hardened.apparmor.enable` and `hardened.kernel.hardened` are one-liners on the host config (enabled on the ISO, off on the machine "for now"). ### Start next (ordered) -1. **Gather the failure detail.** Run `just boot`; the harness now prints - `journalctl -n 6` per failed unit plus `/etc/nixos`, `/persistent/etc` and - `/swap` state. (The last run was aborted before this landed.) -2. **Fix the harness Secure Boot check** — `sbctl status` emits - `Secure Boot: ✓ Enabled`; change the grep to `Secure Boot:.*Enabled`. -3. **Root-cause + fix `/etc/nixos` not persisted and the failed services** — - expected to trace back to the `persist-files` activation failure during - `nixos-install` (impermanence directories empty on the installed disk). - Verify the same `persist-files` snippet also handles `/etc/nixos`. -4. **Confirm harness tolerance of the watchdog self-reboot** — run `just boot` - 3-5× cold; it should pass every time via the second boot. If it ever fails, - fall back to `boot.blacklistedKernelModules = [ "iTCO_wdt" ]` on the host - config (documented in `debugging.md`). -5. **Reach green then `just up` end-to-end** (build → live → install → boot). - The pending changes (`home.nix`, `justfile`, `iso-test.sh`, doc updates) - are a natural first commit once the boot phase is green. +1. **Re-run `just up`** after any future ISO/host change; the full pipeline + is the commit gate. +2. If a cold boot ever fails again, read the *persisted* journal from the + failed boot (`journalctl --list-boots`, then `-b ` — `/var/log/journal` + now survives reboots, which is what finally cracked the first-boot mystery). +3. Consider the `iTCO_wdt` blacklist if the watchdog reboot noise ever + matters for CI. ## Useful links diff --git a/kubernetes/docs/debugging.md b/kubernetes/docs/debugging.md index 1f9edf7..f4207dd 100644 --- a/kubernetes/docs/debugging.md +++ b/kubernetes/docs/debugging.md @@ -938,6 +938,14 @@ glyph blocks, letting you "see" the VGA layout in a terminal without opening the config, so only use if the harness-tolerance route is insufficient. 3. Real-hardware note: 13-inch-thin-cannon has no TCO; this bug is purely a QEMU harness artifact and will not occur on physical hardware. +- **Status update (2026-08-13, end of session):** harness tolerance confirmed + sufficient — fresh install + 3× repeated cold `just boot` runs all green + (only the expected wireguard WARN). No kernel blacklist needed. +- **Related but distinct:** a *separate* intermittent "SSH dead on some boots" + failure was root-caused to a **systemd ordering cycle** (swapfile service + `Before=swap.target` → systemd deleted `run-wrappers.mount` → missing + `unix_chkpwd` → sshd PAM denial). Full write-up in + `13-inch-thin-cannon/docs/iso-build.md` (Known issues & gotchas). ### Technique: distinguishing the reboot type (tested 2026-08-13) - `-global ICH9-LPC.noreboot=on`: stops TCO/hardware-chipset resets. Applied