From 375add5b9eee016e4b0c065a7dbb6b2eeba0e435 Mon Sep 17 00:00:00 2001 From: File Magic Date: Thu, 20 Aug 2026 20:48:16 -0400 Subject: [PATCH] [kubernetes/docs] nixidy: add docs for nixidy --- kubernetes/docs/nixidy.md | 208 ++++++++++++++++++++++++++++++++++++++ 1 file changed, 208 insertions(+) diff --git a/kubernetes/docs/nixidy.md b/kubernetes/docs/nixidy.md index 22255e3..da25948 100644 --- a/kubernetes/docs/nixidy.md +++ b/kubernetes/docs/nixidy.md @@ -89,6 +89,214 @@ checks have been tested independently. - The current IPv6 DNS egress issue must be resolved independently of Nixidy; it is in the Cilium service path for host-network CoreDNS endpoints. +## Full Migration Design + +### Current Responsibility Split + +Ansible currently performs five different jobs: + +1. SSH and API readiness +2. Offline image distribution +3. Cilium installation +4. Rook/Ceph installation +5. Cluster verification + +Nixidy directly replaces application installation (items 3 and 4). It does not replace +VM provisioning, node-local NixOS configuration, offline image import, or health checks. + +The target boundary is: + +- **NixOS:** node services, containerd, CoreDNS, kubelet, PKI, CNI extraction, and the + pod sandbox image +- **OpenTofu:** VMs, disks, libvirt networking, and node metadata +- **Nixidy:** Cilium, Rook, Ceph, CSI, namespaces, CRDs, workloads, and services +- **Operations tool:** SSH/API waits, image transfer/import, and live verification + +### Migration Phases + +#### Phase 0: Stabilize the Existing Cluster + +Before changing application ownership: + +- Fix the Cilium host-to-local DNS failure affecting workers that host CoreDNS. +- Make the IPv6 egress verification reliable. +- Make desired node-count checks dynamic instead of hardcoded to four nodes. +- Run a complete `just up` with the existing Ansible workflow and require `verify.yml` + to pass. + +Nixidy will reproduce the current networking configuration, so it should not be used to +mask this existing failure. + +#### Phase 1: Add Nixidy in Parallel + +Add a pinned Nixidy flake input and a dedicated cluster environment. Keep the current +NixOS packages and offline-image output in the same flake initially. + +Use `lib.helm.downloadHelmChart` for Cilium with a fixed chart hash. Do not download the +chart with Ansible at deployment time. + +Do not change the current deployment path in this phase. + +#### Phase 2: Render and Compare Cilium + +Move the values from `ansible/playbooks/tasks/cilium.yml` into a Nixidy Helm release. +Compare the Nixidy output with the current Helm output and `just helm-images` output. + +The expected Cilium images are: + +- `quay.io/cilium/cilium:v1.20.0` +- `quay.io/cilium/cilium-envoy:v1.37.5-1782911245-7cffc778c923f68a77954a53b1a98d6b5353f004` +- `quay.io/cilium/operator-generic:v1.20.0` + +Only switch Cilium deployment to Nixidy after the rendered resources and image set match. + +#### Phase 3: Apply Cilium With Nixidy + +The temporary lifecycle becomes: + +```text +deploy -> wait SSH -> wait API -> import images -> nixidy apply Cilium -> verify Cilium +``` + +Keep Rook installation and the existing verification playbook unchanged while testing +this phase. + +#### Phase 4: Render Rook/Ceph + +The current Rook workflow uses raw manifests rather than a Helm chart. Represent these +with Nixidy raw resources or Kustomize support: + +- Rook namespace +- Rook CRDs and common resources +- CSI operator CRDs and resources +- Rook operator Deployment +- NixOS-specific operator ConfigMap +- IPv6 CSI controller patch +- CephCluster +- CephBlockPool +- StorageClass + +Preserve the existing NixOS-specific settings for `/nix`, kernel modules, CSI mounts, +host networking, probes, IPv6 Ceph networking, and the disabled dashboard. + +#### Phase 5: Use Staged Rook Applications + +Do not assume one Nixidy apply can converge Rook. Rook has readiness dependencies beyond +CRD and namespace ordering: + +```text +Rook CRDs + -> CSI operator CRDs + -> Rook operator + -> operator-created resources + -> CephCluster + -> CephBlockPool + -> StorageClass +``` + +Use separate Nixidy applications or environments, for example: + +```text +nixidy apply .#bootstrap +wait for Rook/CSI CRDs and operator +nixidy apply .#storage +``` + +#### Phase 6: Remove Application Ansible + +After Cilium and Rook converge through Nixidy, remove: + +- `ansible/playbooks/tasks/cilium.yml` +- `ansible/playbooks/tasks/rook-ceph.yml` +- Application installation imports from `site.yml` +- Ansible chart and manifest downloads +- Remote Helm execution +- Remote Python/YAML resource extraction + +Keep SSH readiness, API readiness, image import, and verification temporarily. + +#### Phase 7: Replace Image Import + +Replace `import-images.yml` with a Nix-built operations script using `ssh`, `rsync` or +`scp`, `ctr`, `jq`, and `kubectl`. + +For each node it must: + +1. Ensure containerd is active. +2. Copy the `.tar.zst` archives. +3. Import each archive into the `k8s.io` containerd namespace. +4. Remove temporary files. + +The script should check expected image references before importing so repeated runs are +idempotent. + +#### Phase 8: Replace Verification + +Create a Nix-built verifier independent of Nixidy. It must preserve the checks currently +in `verify.yml`: + +- Expected nodes are Ready and IPv6-only +- No non-zero CrashLoopBackOff pods +- Cilium and CoreDNS are available +- Cilium is IPv6-only +- Rook operator and CephCluster are healthy +- Ceph monitors and CSI DaemonSets are ready +- CephBlockPool and StorageClass exist +- Pod IPv6 HTTPS egress works +- Host IPv6 egress works from every node + +Use `kubectl`, `jq`, and SSH. The verifier must test live state, not only generated +Nixidy output. + +### Kubeconfig Handling + +The master currently owns the admin certificate files referenced by its kubeconfig. A +raw copy of that kubeconfig cannot reliably be used on the control machine. + +Before running Nixidy or control-machine verification, create a flattened kubeconfig on +the master and fetch it temporarily: + +```bash +kubectl config view --raw --flatten +``` + +Use that temporary kubeconfig for `nixidy apply` and `kubectl` checks. Remove it after +the operation or store it only in a mode-600 temporary path. + +### Final Lifecycle + +The desired no-Ansible lifecycle is: + +```text +genkey +build +deploy +wait-ssh +wait-api +offline-images +nixidy-bootstrap +nixidy-storage +verify +``` + +Ansible can be removed only after this lifecycle passes repeatedly from a clean +`just down` state. + +### Removal Criteria + +Remove Ansible and its development dependencies only when: + +- Cilium and Rook are applied exclusively through Nixidy. +- All image archives are imported without Ansible. +- The replacement verifier covers every existing verification assertion. +- The full lifecycle passes after a clean destroy. +- A second clean lifecycle passes without relying on pre-existing containerd images. +- No resources owned by NixOS are accidentally pruned by Nixidy. + +Nixidy's `kubectl apply --prune` behavior makes ownership boundaries important. CoreDNS, +NixOS-seeded resources, and other resources not declared by Nixidy must not receive +Nixidy ownership labels or be included in its prune scope. + ## References - [Nixidy](https://nixidy.dev/) - Nix-based Kubernetes manifest generation and apply -- 2.51.2