diff --git a/kubernetes/docs/debugging.md b/kubernetes/docs/debugging.md index 24e29bc..cd4659c 100644 --- a/kubernetes/docs/debugging.md +++ b/kubernetes/docs/debugging.md @@ -831,9 +831,12 @@ After NixOS rebuild, the OVMF path changes. Find it: virsh dumpxml k8s-master-0 | grep -i ovmf # Use the path from the element -# Or evaluate directly: -nix eval --raw nixpkgs#OVMF.fd 2>/dev/null +# Or build the flake's pinned OVMF (realizes the path, see 2026-08-01 entry): +nix build --no-link --print-out-paths .#ovmf +# -> /nix/store/-OVMF--fd (contains FV/OVMF_CODE.fd + FV/OVMF_VARS.fd) ``` +Do NOT use `nix eval --raw nixpkgs#OVMF.fd` — it prints the store path but does not +realize it, so `FV/OVMF_CODE.fd` does not exist unless OVMF was previously built. ## virsh UEFI VM Management @@ -844,3 +847,76 @@ virsh -c qemu:///system undefine k8s-master-0 --nvram # List VMs and their state: virsh -c qemu:///system list --all ``` + +--- + +## 2026-08-01: `tofu destroy` leaves k8s-* disks/networks behind + +### Symptom +- After `just up` then `just down`, libvirt still had orphaned resources: an active + `k8s-1` NAT network with autostart, while the tofu state was empty and no domains/volumes + were tracked. +- `tofu destroy` itself correctly deletes tracked volumes (verified: 13 resources, + including `worker_data` raw disks, fully removed). The leftovers were resources that + had drifted out of the tofu state (interrupted/failed applies, or switching + `CLUSTER_INDEX`/`WORKER_COUNT` between runs). + +### Root cause +1. **State drift.** Volumes/domains are named without a cluster index + (`k8s-master-0.qcow2` in the shared `default` pool) while the network IS index-scoped + (`k8s-${index}`). If an apply is interrupted or fails after libvirt objects are created + but before tofu records them, or if you switch `CLUSTER_INDEX`/`WORKER_COUNT`, objects + exist in libvirt but are invisible to `tofu destroy` — so they are never deleted. +2. **`nix eval --raw nixpkgs#OVMF.fd` returns an unrealized store path.** `tofu apply` + then fails at `libvirt_domain` creation ("Failed to Start Domain") after the disks and + network were already created, leaving a half-applied state and forcing manual cleanup. + This is what caused the initial orphaned disks: a `just up` whose apply died partway. + +### Fix +- Root cause was the OVMF path: it is now realized via the flake + `nix build --no-link --print-out-paths .#ovmf` (new `ovmf` package output in + `flake.nix`) instead of `nix eval --raw`. Applies no longer fail at domain start, so + tofu state stays consistent and `tofu destroy` removes everything. +- `just down` stays a pure `tofu destroy` (no virsh side-commands): tofu state is the + single source of cleanup. A virsh prune was tried as a workaround but removed — + the goal is that a failure like this should be fixed at its root, not papered over + with separate manual cleanup. + +### Technique notes (not currently used) +- `virsh vol-list ` does **not** support `--name` (unlike `list`/`net-list`). + Parse the name column: `virsh vol-list --pool default | awk 'NR>2 && $1 != "" {print $1}'`. +- `virsh undefine --nvram` also removes the OVMF NVRAM file. +- When `tofu destroy` reports nothing destroyed but libvirt still has `k8s-*` objects, + the state was drifted by a failed apply; re-apply once (or recover state) so destroy + can track them again. + +--- + +## 2026-08-01: ansible fails with missing `/var/lib/kubelet/secrets/cluster-admin.pem` on fresh cluster + +### Symptom +- `just up` succeeded through deploy + wait-ssh, then the Cilium Helm task failed: + `kubernetes cluster unreachable: invalid configuration: unable to read client-cert + /var/lib/kubelet/secrets/cluster-admin.pem: no such file or directory`. +- The certs existed minutes later — this is a boot race, not a config error. + +### Root cause +The NixOS `kubernetes` module with `easyCerts = true` generates the cluster PKI at boot: +the master's `cfssl` service + `kube-certmgr-bootstrap` create `/var/lib/kubelet/secrets/*.pem`. +The ansible playbook only waited for SSH (`wait_for_connection`), so on a fast-booting +VM (SSH ready in ~10s) ansible ran helm before cfssl finished generating certs. + +### Fix +`ansible/site.yml` gained a "Wait for Kubernetes API server" play before Cilium install: +`kubectl get --raw /readyz` with `KUBECONFIG=/etc/kubernetes/cluster-admin.kubeconfig`, +retried 60 × 5s until rc==0 (per AGENTS.md: `until` + `retries` + `delay`). This waits +for both the certs AND the apiserver to be healthy. Use `ansible.builtin.command` + +task-level `environment:` (not `shell`) to satisfy ansible-lint's +`command-instead-of-shell` rule. + +### Technique: SSH around stale known_hosts +After redeploys, the VM host key changes and ssh aborts on the fingerprint mismatch +(see the `Offending RSA key` warning). Use `-o UserKnownHostsFile=/dev/null` for manual +diagnostic SSH; the justfile already sets `-o StrictHostKeyChecking=no`. + +--- diff --git a/kubernetes/docs/security-hardening.md b/kubernetes/docs/security-hardening.md index ee4e109..1cf7b83 100644 --- a/kubernetes/docs/security-hardening.md +++ b/kubernetes/docs/security-hardening.md @@ -63,6 +63,18 @@ All tasks verified on live cluster (`ansible-playbook site.yml: failed=0`). See - Kubelet extraOpts: `--protect-kernel-defaults=true --read-only-port=0 --streaming-connection-idle-timeout=5m` - All workers `Ready` and running kubelet +## Verified Test Results (2026-08-01) + +Full lifecycle verified end-to-end: `just up` (genkey → build → deploy → wait-ssh → apiserver-wait → ansible → verify) with `failed=0` on all nodes, then `just down` destroyed all 13 tofu resources and left libvirt clean (no domains, no `k8s-*` volumes/networks). + +**Changes since 2026-07-16:** +- OVMF path realized via flake: `OVMF_CODE_PATH = nix build --no-link --print-out-paths .#ovmf` (new `ovmf` package in `flake.nix`). Fixes `nix eval --raw` returning an unrealized store path that broke `tofu apply` and left disks outside tofu state. +- `just down` remains a pure `tofu destroy` — tofu state is the single source of cleanup. With the OVMF path realized, applies no longer fail halfway and leave orphaned volumes. +- `genkey` recipe generates the SSH keypair (`ssh-key` + tracked `ssh-public-key`) when missing; `up` depends on it. Private key was lost (gitignored, never committed); keypair regenerated 2026-08-01 and baked into a fresh image build. +- `site.yml` now waits for apiserver readiness (`kubectl get --raw /readyz`, retries 60 × 5s) before installing Cilium. Fixes a boot race where ansible started as soon as SSH was up but cfssl/easyCerts had not finished generating the cluster PKI (`/var/lib/kubelet/secrets/cluster-admin.pem` missing). + +See `docs/debugging.md` "2026-08-01" entries for root causes. + ### Pod Networking Verification From a test pod (`kubectl run test --image=nginx:alpine --rm -it --restart=Never -- /bin/sh`):