diff --git a/kubernetes/docs/debugging.md b/kubernetes/docs/debugging.md index 8453025..24e29bc 100644 --- a/kubernetes/docs/debugging.md +++ b/kubernetes/docs/debugging.md @@ -821,3 +821,26 @@ systemd.tmpfiles.rules = [ "L /lib/modules - - - - /run/current-system/kernel-modules/lib/modules" ]; ``` + +--- + +## OVMF Path Discovery + +After NixOS rebuild, the OVMF path changes. Find it: +```bash +virsh dumpxml k8s-master-0 | grep -i ovmf +# Use the path from the element + +# Or evaluate directly: +nix eval --raw nixpkgs#OVMF.fd 2>/dev/null +``` + +## virsh UEFI VM Management + +```bash +# Undefine UEFI VMs requires --nvram flag (no path argument): +virsh -c qemu:///system undefine k8s-master-0 --nvram + +# List VMs and their state: +virsh -c qemu:///system list --all +``` diff --git a/kubernetes/docs/resume-state.md b/kubernetes/docs/resume-state.md index 0ebe21e..1b20020 100644 --- a/kubernetes/docs/resume-state.md +++ b/kubernetes/docs/resume-state.md @@ -12,7 +12,7 @@ - 4 cilium agent pods (DaemonSet) — all Ready - 2 cilium-operator replicas — all Ready - 4 cilium-envoy pods — all Ready -- Hubble relay + UI: **disabled** (non-essential, caused CrashLoopBackOff) +- Hubble relay + UI: **disabled** **CoreDNS:** 2 pods running, healthy. @@ -25,11 +25,7 @@ - CephBlockPool `replicapool`: exists - StorageClass `rook-ceph-block`: exists, set as default -**NixOS Security Hardening (all committed):** -- Firewall enabled, SSH hardened, kernel sysctls set -- API server: `allowPrivileged=true`, `anonymous-auth=false`, audit logging, PodSecurity + NodeRestriction admission -- Kubelet: `protect-kernel-defaults=true`, `read-only-port=0`, streaming timeout 5m -- kube-proxy: **disabled** (Cilium handles routing via eBPF) +**NixOS Security Hardening (all committed):** See `security-hardening.md` for details. ## Git Commits (in order, oldest first) @@ -54,9 +50,9 @@ fc2967e verify.yml [kubernetes/ansible/tasks/]: add checks for Hubble, Rook-Ceph ## Known Issues -1. **CephCluster health is HEALTH_WARN** — OSDs are still converging. This is normal for a fresh cluster. The verify playbook accepts both HEALTH_OK and HEALTH_WARN. -2. **ceph-csi-controller-manager in CrashLoopBackOff** — exit code 0 (Completed). This is a one-shot process that does its work and exits. The CrashLoopBackOff check in verify.yml is smart enough to ignore pods with exit code 0. -3. **Hubble relay + UI disabled** — non-essential for security hardening. Can re-enable later if needed. +1. **CephCluster health is HEALTH_WARN** — OSDs converging. Verify playbook accepts both HEALTH_OK and HEALTH_WARN. +2. **ceph-csi-controller-manager in CrashLoopBackOff** — exit code 0 (normal for one-shot process). See `security-hardening.md` for details. +3. **Hubble relay + UI disabled** — non-essential. Can re-enable later if needed. 4. **K8s API token still hardcoded in flake.nix** — Task 4 (secrets management) not yet implemented. ## SSH Access @@ -83,7 +79,4 @@ $TOFU apply -auto-approve -var "image_dir=../result" # Partial rebuild (VMs only, skip image build) $TOFU destroy -auto-approve && \ $TOFU apply -auto-approve -var "image_dir=../result" - -# Check OVMF path (changes after nix rebuild): -virsh dumpxml k8s-master-0 | grep -i ovmf ``` diff --git a/kubernetes/docs/security-hardening.md b/kubernetes/docs/security-hardening.md index 847a187..ee4e109 100644 --- a/kubernetes/docs/security-hardening.md +++ b/kubernetes/docs/security-hardening.md @@ -51,35 +51,17 @@ ## Verified Test Results (2026-07-16) -### All tasks verified on live cluster (ansible-playbook site.yml: failed=0) +All tasks verified on live cluster (`ansible-playbook site.yml: failed=0`). See `resume-state.md` for current component state. -**Master (192.168.122.10):** +**Hardening verification on master and workers:** - Firewall: `nixos-fw` chain present in iptables INPUT, port 8888 (cfssl) open - Kernel: `dmesg_restrict=1`, `kptr_restrict=2`, `suid_dumpable=0`, `panic=10`, `panic_on_oops=1`, `vm.overcommit_memory=1` - SSH: `PermitRootLogin prohibit-password`, `PasswordAuthentication no`, `AllowAgentForwarding no`, `AllowTcpForwarding no` - Apiserver: `--allow-privileged=true`, `--anonymous-auth=false`, `--audit-log-path=/var/log/kubernetes/audit.log`, `--enable-admission-plugins=PodSecurity,NodeRestriction` - Audit log: Active, writing to `/var/log/kubernetes/audit.log` -- kube-proxy: **disabled** (Cilium handles routing) -- Apiserver, etcd, cfssl: all `active` - -**Workers (192.168.122.11-13):** -- Kernel sysctls: all verified ✓ -- SSH: `PermitRootLogin prohibit-password` ✓ -- Kubelet extraOpts confirmed in systemd unit: `--protect-kernel-defaults=true --read-only-port=0 --streaming-connection-idle-timeout=5m` -- kube-proxy: **disabled** on all workers -- All workers `Ready` and running kubelet ✓ - -### Ansible Results (all passing) - -- **Cilium CNI:** Installed, 4 agent pods + 2 operator + 4 envoy pods running ✓ -- **Cilium eBPF masquerade:** `bpf.masquerade: true` set, kube-proxy disabled ✓ -- **CoreDNS:** 2 pods running, healthy ✓ -- **Rook-Ceph operator:** Available ✓ -- **CephCluster:** HEALTH_WARN (OSDs converging — normal for fresh cluster) ✓ -- **Ceph mon pods:** 1 running ✓ -- **Ceph CSI node plugin:** DaemonSet Ready on all 3 workers ✓ -- **CephBlockPool replicapool:** Exists ✓ -- **StorageClass rook-ceph-block:** Exists, set as default ✓ +- kube-proxy: **disabled** on all nodes +- Kubelet extraOpts: `--protect-kernel-defaults=true --read-only-port=0 --streaming-connection-idle-timeout=5m` +- All workers `Ready` and running kubelet ### Pod Networking Verification @@ -162,39 +144,10 @@ kubectl get pods -n kube-system -o json | jq -r '.items[] | # Exit code 1 = application error ``` -### NixOS Rebuild + Cluster Rebuild Cycle -```bash -# After NixOS config changes: -nix build .#images # Build new VM images -cd tofu && tofu destroy -auto-approve # Destroy old VMs -tofu apply -auto-approve # Create new VMs from fresh images -# Wait for SSH: -ssh -F /dev/null -o IdentityAgent=none -i ssh-key root@192.168.122.10 -# Run playbook: -cd ../ansible && ansible-playbook site.yml -``` - -### OVMF Path Discovery -```bash -# After NixOS rebuild, the OVMF path changes. Find it: -virsh dumpxml k8s-master-0 | grep -i ovmf -# Use the path from the element - -# Or evaluate directly: -nix eval --raw nixpkgs#OVMF.fd 2>/dev/null -``` - -### virsh UEFI VM Management -```bash -# Undefine UEFI VMs requires --nvram flag (no path argument): -virsh -c qemu:///system undefine k8s-master-0 --nvram - -# List VMs and their state: -virsh -c qemu:///system list --all -``` - ## Cluster Test Procedure +See `resume-state.md` for rebuild commands (nix build, tofu destroy/apply). + 1. Build images: `nix build .#images` 2. Deploy VMs: `cd tofu && tofu apply` 3. Wait for SSH: `ssh -i ssh-key root@192.168.122.10` diff --git a/kubernetes/docs/session-state-2026-07-12.md b/kubernetes/docs/session-state-2026-07-12.md index dcf2b65..40de7ac 100644 --- a/kubernetes/docs/session-state-2026-07-12.md +++ b/kubernetes/docs/session-state-2026-07-12.md @@ -1,29 +1,13 @@ -# Session State — 2026-07-12/14 +# Session Log — 2026-07-12/14 -## Current State: CLUSTER FULLY OPERATIONAL — Cilium + Rook-Ceph + PVC provisioning verified - -All 4 nodes Ready, Cilium CNI + Hubble running, Rook-Ceph v1.20.2 with CSI operator enabled, 3 OSDs, StorageClass with CSI secrets, **PVC provisioning tested and working end-to-end**. - -### Pod/Volume Verification -``` -NAME READY STATUS RESTARTS AGE IP NODE -test-pod 1/1 Running 0 2m49s 10.0.0.247 k8s-worker-1 -# Output: "hello from ceph" -``` - -### Ceph Status -- CephCluster: Ready, v19.2.0 Squid, 1 mon, 1 mgr, 3 OSDs -- CephBlockPool `replicapool` (size:1) + StorageClass `rook-ceph-block` (default, allowVolumeExpansion) -- CSI operator deploying RBD + CephFS node plugins automatically - -## What Changed This Session +## Changes Made ### Critical Fixes -- **kubelet rootDir**: `services.kubernetes.dataDir = "/var/lib/kubelet"` — NixOS defaulted to `/var/lib/kubernetes` which CSI containers can't see -- **clusterDns**: Workers computed wrong DNS IP (10.0.0.254 vs 10.96.0.254) — set `addons.dns.clusterIp` explicitly in common.nix -- **CSI operator**: `ROOK_USE_CSI_OPERATOR: "true"` — Driver CRDs now deploy CSI node plugin DaemonSets -- **StorageClass secrets**: Added `csi.storage.k8s.io/provisioner-secret-name/namespace` and `node-stage-secret-name/namespace` -- **/lib/modules symlink**: systemd tmpfiles rule for CSI kernel module access on NixOS +- **kubelet rootDir**: `services.kubernetes.dataDir = "/var/lib/kubelet"` — NixOS defaulted to `/var/lib/kubernetes` which CSI containers can't see (full root cause in `debugging.md`) +- **clusterDns**: Workers computed wrong DNS IP (10.0.0.254 vs 10.96.0.254) — set `addons.dns.clusterIp` explicitly in common.nix (full root cause in `debugging.md`) +- **CSI operator**: `ROOK_USE_CSI_OPERATOR: "true"` — Driver CRDs now deploy CSI node plugin DaemonSets (full root cause in `debugging.md`) +- **StorageClass secrets**: Added `csi.storage.k8s.io/provisioner-secret-name/namespace` and `node-stage-secret-name/namespace` (full root cause in `debugging.md`) +- **/lib/modules symlink**: systemd tmpfiles rule for CSI kernel module access on NixOS (full root cause in `debugging.md`) ### Bugs Fixed (16-20) 16. CSI staging path mismatch (kubelet rootDir) @@ -32,68 +16,11 @@ test-pod 1/1 Running 0 2m49s 10.0.0.247 k8s-worker-1 19. StorageClass missing CSI secret references 20. Pod scheduled on master (no CSI node plugin) -## Key Config Summary - -| Setting | Value | -|---------|-------| -| kubelet root-dir | `/var/lib/kubelet` (via `services.kubernetes.dataDir`) | -| DNS clusterIp | `10.96.0.254` (via `addons.dns.clusterIp`) | -| CSI operator | enabled (`ROOK_USE_CSI_OPERATOR: "true"`) | -| StorageClass | `rook-ceph-block` with CSI secret refs + allowVolumeExpansion | -| CephCluster | v19.2.0 Squid, `useAllDevices: true` | -| CephBlockPool | `replicapool` size:1 | -| /lib/modules | systemd tmpfiles symlink | - -## Committed Changes (this session) +## Committed Changes 1. `common.nix, master.nix, worker.nix`: Fix clusterDns and kubelet rootDir for CSI compatibility 2. `rook-ceph.yml`: Enable CSI operator, fix StorageClass secrets -## Commands - -```bash -# SSH into master -SSH_KEY="/home/file_magic/Projects/github/dotfiles/kubernetes/ssh-key" -env -u SSH_AUTH_SOCK ssh -o IdentityAgent=none -i "$SSH_KEY" root@192.168.122.10 - -# Verify kubelet root-dir -cat /proc/$(pgrep kubelet)/cmdline | tr "\0" "\n" | grep root-dir - -# Check Ceph health -kubectl -n rook-ceph exec deploy/rook-ceph-mgr-a -- ceph health - -# Test PVC provisioning -kubectl apply -f - < /mnt/data/test.txt && cat /mnt/data/test.txt && sleep 3600"] - volumeMounts: - - name: test-vol - mountPath: /mnt/data - volumes: - - name: test-vol - persistentVolumeClaim: - claimName: test-pvc -EOF -``` - ## Design Decisions - **dataDir = /var/lib/kubelet**: Matches Kubernetes upstream default; avoids CSI path mismatch