From 353284693defa39c6fc85079037067b822903778 Mon Sep 17 00:00:00 2001 From: File Magic Date: Sat, 1 Aug 2026 17:25:03 -0400 Subject: [PATCH] resume-state.md [kubernetes/docs]: update cluster state snapshot for 2026-08-01 --- kubernetes/docs/resume-state.md | 35 +++++++++++++++++++-------------- 1 file changed, 20 insertions(+), 15 deletions(-) diff --git a/kubernetes/docs/resume-state.md b/kubernetes/docs/resume-state.md index eb9af25..61f4476 100644 --- a/kubernetes/docs/resume-state.md +++ b/kubernetes/docs/resume-state.md @@ -1,8 +1,8 @@ -# Cluster State — Save Point (2026-07-16) +# Cluster State — Save Point (2026-08-01) ## What's Running -**4 nodes, all `Ready`:** +**4 nodes, all `Ready` (just `just up` 2026-08-01, verify `failed=0`):** - k8s-master-0 (192.168.122.10) - k8s-worker-0 (192.168.122.11) - k8s-worker-1 (192.168.122.12) @@ -18,7 +18,7 @@ **Rook-Ceph:** - Operator: Running, available -- CephCluster: HEALTH_WARN (OSDs still converging — normal for fresh cluster) +- CephCluster: HEALTH_OK (converged since 2026-07-16 HEALTH_WARN) - 1 mon pod running - 3 OSD pods running (one per worker) - Ceph CSI node plugin DaemonSet: Ready on all 3 workers @@ -50,23 +50,25 @@ fc2967e verify.yml [kubernetes/ansible/tasks/]: add checks for Hubble, Rook-Ceph ## Known Issues -1. **CephCluster health is HEALTH_WARN** — OSDs converging. Verify playbook accepts both HEALTH_OK and HEALTH_WARN. -2. **ceph-csi-controller-manager in CrashLoopBackOff** — exit code 0 (normal for one-shot process). See `security-hardening.md` for details. -3. **Hubble relay + UI disabled** — non-essential. Can re-enable later if needed. -4. **K8s API token still hardcoded in flake.nix** — Task 4 (secrets management) not yet implemented. +1. **ceph-csi-controller-manager in CrashLoopBackOff** — exit code 0 (normal for one-shot process). See `security-hardening.md` for details. +2. **Hubble relay + UI disabled** — non-essential. Can re-enable later if needed. +3. **K8s API token still hardcoded in flake.nix** — Task 4 (secrets management) not yet implemented. ## SSH Access +The keypair is now generated by `just genkey` (gitignored `kubernetes/ssh-key`, tracked +`kubernetes/ssh-public-key`). Regenerated 2026-08-01 after the old key was lost: + ```bash ssh -F /dev/null -o StrictHostKeyChecking=no -o IdentityAgent=none \ - -i /home/file_magic/Projects/github/dotfiles/kubernetes/ssh-key \ + -i kubernetes/ssh-key \ root@192.168.122.10 ``` ## tofu / Nix Commands ```bash -TOFU="/nix/store/2q5865pn7ygz8n1bjdzgsxxagbhpvisw-opentofu-1.12.3/bin/tofu" +TOFU="/nix/store/ypjys1z1rj0df3aa3rm7xi2p87kcjn8y-opentofu-1.12.4/bin/tofu" # Check VMs virsh -c qemu:///system list --all @@ -89,10 +91,11 @@ The justfile provides a complete lifecycle for managing the cluster. All targets | Target | Description | |--------|-------------| -| `just up` | **Full lifecycle**: build images → deploy VMs → wait for SSH → run Ansible | +| `just up` | **Full lifecycle**: genkey → build images → deploy VMs → wait for SSH → run Ansible | | `just build` | Build NixOS VM images only | +| `just genkey` | Generate SSH keypair if missing (writes `ssh-key` + `ssh-public-key`) | | `just deploy` | Destroy old VMs and create new ones from images | -| `just down` | Destroy all VMs | +| `just down` | Destroy all VMs, disks, and networks (`tofu destroy` — single source of cleanup) | | `just wait-ssh` | Wait for SSH on master (checks every 5s, 60 retries) | | `just ssh` | SSH into master node | | `just ansible` | Run full Ansible playbook (installs Cilium, Rook-Ceph, verifies) | @@ -150,11 +153,13 @@ just fmt # Auto-fix formatting issues ### Target Details -- **`just up`** — The primary target. Chains: `build` → `deploy` → `wait-ssh` → `ansible` +- **`just up`** — The primary target. Chains: `genkey` → `build` → `deploy` → `wait-ssh` → `ansible` - **`just build`** — Runs `nix build .#images` to create QCOW2 images +- **`just genkey`** — Generates the SSH keypair if missing (no-op if `ssh-key` exists) - **`just deploy`** — Runs `tofu destroy` then `tofu apply` with the built images - **`just ansible`** — Runs the full Ansible playbook which: 1. Waits for SSH on all nodes - 2. Installs Cilium CNI with eBPF masquerade - 3. Installs Rook-Ceph storage (operator, CephCluster, CSI, StorageClass) - 4. Verifies cluster health (all pods Running, all services Ready) + 2. Waits for the Kubernetes apiserver readiness (covers cfssl PKI generation race) + 3. Installs Cilium CNI with eBPF masquerade + 4. Installs Rook-Ceph storage (operator, CephCluster, CSI, StorageClass) + 5. Verifies cluster health (all pods Running, all services Ready) -- 2.51.2