diff --git a/.gitignore b/.gitignore index 6d38d52..c99a201 100644 --- a/.gitignore +++ b/.gitignore @@ -1,14 +1,13 @@ # terraform -infra/**/.terraform/ -infra/**/.terraform.lock.hcl -infra/**/terraform.tfstate -infra/**/terraform.tfstate.backup -infra/**/*.tfvars -!infra/terraform.tfvars.example +**/infra/**/.terraform/ +**/infra/**/.terraform.lock.hcl +**/infra/**/terraform.tfstate +**/infra/**/terraform.tfstate.backup +**/infra/**/*.tfvars +!**/infra/terraform.tfvars.example # kubeconfig (fetched from server) -kubeconfig.yaml -zlay-kubeconfig.yaml +**/kubeconfig.yaml # secrets *.secret diff --git a/README.md b/README.md index 866fe7d..2751240 100644 --- a/README.md +++ b/README.md @@ -59,11 +59,22 @@ consumes the simplified JSON firehose via [jetstream](https://github.com/bluesky ``` . -├── docs/ # architecture, deployment guide, backfill +├── indigo/ # Go relay (indigo) — justfile, deploy configs, terraform +├── zlay/ # zig relay (zlay) — justfile, deploy configs, terraform +├── shared/deploy/ # helm values shared by both deployments ├── scripts/ # uv scripts — firehose, jetstream, backfill -├── justfile # all commands: deploy, status, logs, backfill, etc. -├── infra/ # terraform — hetzner server + k3s -└── deploy/ # helm values + k8s manifests +├── docs/ # architecture, deployment guide, backfill +└── justfile # root — `just indigo ` / `just zlay ` +``` + +each relay is a `just` module with symmetric recipes: + +```bash +just indigo deploy # deploy Go relay +just indigo status # check Go relay pods +just zlay deploy # deploy zig relay +just zlay status # check zig relay pods +just --list # see all available recipes ``` ## why diff --git a/docs/architecture.md b/docs/architecture.md index b7f1d40..a02b950 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -45,7 +45,7 @@ the relay and collectiondir ServiceMonitors are standalone manifests (`kubectl a relays try to reconnect to PDS hosts when connections drop, but eventually give up after repeated failures (exponential backoff). PDS hosts re-announce themselves to bluesky's relay when they come back online, but not to third-party relays like ours. this causes a natural decay in connected host count over time. -fix: a k8s CronJob (`deploy/reconnect-cronjob.yaml`) runs every 4 hours, fetching the [community PDS list](https://github.com/mary-ext/atproto-scraping) and sending `requestCrawl` for each host. this can also be run manually via `just reconnect`. +fix: a k8s CronJob (`indigo/deploy/reconnect-cronjob.yaml`) runs every 4 hours, fetching the [community PDS list](https://github.com/mary-ext/atproto-scraping) and sending `requestCrawl` for each host. this can also be run manually via `just indigo reconnect`. ## steady-state specs (indigo relay) @@ -80,16 +80,16 @@ a second relay implementation in [Zig](https://ziglang.org/), deployed on a sepa ### deployment -separate Hetzner cpx41 in Hillsboro OR (`hil`), independent k3s cluster. all `zlay-*` justfile recipes use `ZLAY_KUBECONFIG`. terraform in `infra/zlay/`. +separate Hetzner cpx41 in Hillsboro OR (`hil`), independent k3s cluster. terraform in `zlay/infra/`. ```bash -just zlay-init # terraform init -just zlay-infra # create server -just zlay-kubeconfig # pull kubeconfig -just zlay-deploy # full deploy (cert-manager, postgres, relay, monitoring) -just zlay-publish # build and push image -just zlay-status # check pods + health -just zlay-logs # tail logs +just zlay init # terraform init +just zlay infra # create server +just zlay kubeconfig # pull kubeconfig +just zlay deploy # full deploy (cert-manager, postgres, relay, monitoring) +just zlay publish-remote # build and push image +just zlay status # check pods + health +just zlay logs # tail logs ``` ### collection index backfill diff --git a/docs/backfill.md b/docs/backfill.md index 871bab4..22cbe1c 100644 --- a/docs/backfill.md +++ b/docs/backfill.md @@ -20,14 +20,14 @@ the calls are sequential per host, rate-limited at 100 req/s (configurable via ` ## running the backfill -the `just backfill` recipe handles port-forwarding and host list extraction: +the `just indigo backfill` recipe handles port-forwarding and host list extraction: ```bash # backfill all connected hosts (extracts list from relay automatically) -just backfill +just indigo backfill # backfill from a specific host list with custom batch size -just backfill --hosts /tmp/bsky-shards.txt --batch-size 10 +just indigo backfill --hosts /tmp/bsky-shards.txt --batch-size 10 ``` or run the script directly (requires a port-forward to the collectiondir): diff --git a/docs/deploying.md b/docs/deploying.md index fcb89e5..25bf46b 100644 --- a/docs/deploying.md +++ b/docs/deploying.md @@ -25,40 +25,40 @@ then: ```bash source .env -just init # terraform init -just infra # creates a CPX41 in Ashburn (~$30/mo) with k3s via cloud-init -just kubeconfig # waits for k3s, pulls kubeconfig (~2 min) -just deploy # installs cert-manager, postgresql, relay, jetstream, monitoring +just indigo init # terraform init +just indigo infra # creates a CPX41 in Ashburn (~$30/mo) with k3s via cloud-init +just indigo kubeconfig # waits for k3s, pulls kubeconfig (~2 min) +just indigo deploy # installs cert-manager, postgresql, relay, jetstream, monitoring ``` -point a DNS A record at the server IP (`just server-ip`) before running deploy, so the Let's Encrypt HTTP-01 challenge succeeds. +point a DNS A record at the server IP (`just indigo server-ip`) before running deploy, so the Let's Encrypt HTTP-01 challenge succeeds. after deploy, seed the relay with the network's PDS hosts: ```bash -just bootstrap # pulls hosts from upstream + restarts relay so slurper picks them up +just indigo bootstrap # pulls hosts from upstream + restarts relay so slurper picks them up ``` ## available commands ```bash -just status # nodes, pods, health check -just logs # tail relay logs -just health # curl the public health endpoint -just reconnect # re-announce all known PDS hosts to the relay -just backfill # backfill collectiondir with full network data -just firehose # consume the firehose (passes args through) -just jetstream # consume the jetstream (passes args through) -just ssh # ssh into the server -just destroy # tear down everything +just indigo status # nodes, pods, health check +just indigo logs # tail relay logs +just indigo health # curl the public health endpoint +just indigo reconnect # re-announce all known PDS hosts to the relay +just indigo backfill # backfill collectiondir with full network data +just indigo firehose # consume the firehose (passes args through) +just indigo jetstream # consume the jetstream (passes args through) +just indigo ssh # ssh into the server +just indigo destroy # tear down everything ``` ## maintenance -a k8s CronJob (`deploy/reconnect-cronjob.yaml`) runs every 4 hours to re-announce PDS hosts to the relay — see [architecture](architecture.md#pds-connection-maintenance) for why this is needed. `just reconnect` runs the same logic manually. +a k8s CronJob (`indigo/deploy/reconnect-cronjob.yaml`) runs every 4 hours to re-announce PDS hosts to the relay — see [architecture](architecture.md#pds-connection-maintenance) for why this is needed. `just indigo reconnect` runs the same logic manually. ## targeted deployments -`just deploy` deploys everything. for targeted updates: +`just indigo deploy` deploys everything. for targeted updates: -- `just deploy-monitoring` — only the monitoring stack (prometheus, grafana, dashboards, ServiceMonitors). useful for dashboard changes or prometheus config tweaks without touching the relay. +- `just indigo deploy-monitoring` — only the monitoring stack (prometheus, grafana, dashboards, ServiceMonitors). useful for dashboard changes or prometheus config tweaks without touching the relay. diff --git a/deploy/collectiondir-servicemonitor.yaml b/indigo/deploy/collectiondir-servicemonitor.yaml similarity index 100% rename from deploy/collectiondir-servicemonitor.yaml rename to indigo/deploy/collectiondir-servicemonitor.yaml diff --git a/deploy/collectiondir-values.yaml b/indigo/deploy/collectiondir-values.yaml similarity index 100% rename from deploy/collectiondir-values.yaml rename to indigo/deploy/collectiondir-values.yaml diff --git a/deploy/ingress.yaml b/indigo/deploy/ingress.yaml similarity index 100% rename from deploy/ingress.yaml rename to indigo/deploy/ingress.yaml diff --git a/deploy/jetstream-ingress.yaml b/indigo/deploy/jetstream-ingress.yaml similarity index 100% rename from deploy/jetstream-ingress.yaml rename to indigo/deploy/jetstream-ingress.yaml diff --git a/deploy/jetstream-servicemonitor.yaml b/indigo/deploy/jetstream-servicemonitor.yaml similarity index 100% rename from deploy/jetstream-servicemonitor.yaml rename to indigo/deploy/jetstream-servicemonitor.yaml diff --git a/deploy/jetstream-values.yaml b/indigo/deploy/jetstream-values.yaml similarity index 100% rename from deploy/jetstream-values.yaml rename to indigo/deploy/jetstream-values.yaml diff --git a/deploy/monitoring-values.yaml b/indigo/deploy/monitoring-values.yaml similarity index 100% rename from deploy/monitoring-values.yaml rename to indigo/deploy/monitoring-values.yaml diff --git a/deploy/reconnect-cronjob.yaml b/indigo/deploy/reconnect-cronjob.yaml similarity index 100% rename from deploy/reconnect-cronjob.yaml rename to indigo/deploy/reconnect-cronjob.yaml diff --git a/deploy/relay-dashboard.json b/indigo/deploy/relay-dashboard.json similarity index 100% rename from deploy/relay-dashboard.json rename to indigo/deploy/relay-dashboard.json diff --git a/deploy/relay-servicemonitor.yaml b/indigo/deploy/relay-servicemonitor.yaml similarity index 100% rename from deploy/relay-servicemonitor.yaml rename to indigo/deploy/relay-servicemonitor.yaml diff --git a/deploy/relay-values.yaml b/indigo/deploy/relay-values.yaml similarity index 100% rename from deploy/relay-values.yaml rename to indigo/deploy/relay-values.yaml diff --git a/infra/main.tf b/indigo/infra/main.tf similarity index 100% rename from infra/main.tf rename to indigo/infra/main.tf diff --git a/infra/outputs.tf b/indigo/infra/outputs.tf similarity index 100% rename from infra/outputs.tf rename to indigo/infra/outputs.tf diff --git a/infra/variables.tf b/indigo/infra/variables.tf similarity index 100% rename from infra/variables.tf rename to indigo/infra/variables.tf diff --git a/infra/versions.tf b/indigo/infra/versions.tf similarity index 100% rename from infra/versions.tf rename to indigo/infra/versions.tf diff --git a/indigo/justfile b/indigo/justfile new file mode 100644 index 0000000..99f943e --- /dev/null +++ b/indigo/justfile @@ -0,0 +1,301 @@ +# indigo (Go) relay deployment +# required env vars: HCLOUD_TOKEN, RELAY_DOMAIN, RELAY_ADMIN_PASSWORD, POSTGRES_PASSWORD, LETSENCRYPT_EMAIL +# optional env vars: GRAFANA_DOMAIN (default: relay-metrics.waow.tech), GRAFANA_ADMIN_PASSWORD, JETSTREAM_DOMAIN (default: jetstream.waow.tech) + +export KUBECONFIG := source_directory() / "kubeconfig.yaml" + +# --- infrastructure --- + +# initialize terraform +init: + terraform -chdir=infra init + +# create the hetzner server with k3s +infra: + terraform -chdir=infra apply -var="hcloud_token=$HCLOUD_TOKEN" + +# destroy all infrastructure +destroy: + terraform -chdir=infra destroy -var="hcloud_token=$HCLOUD_TOKEN" + +# get the server IP from terraform +server-ip: + @terraform -chdir=infra output -raw server_ip + +# ssh into the server +ssh: + ssh root@$(just server-ip) + +# --- cluster access --- + +# fetch kubeconfig from the server (run after cloud-init finishes, ~2 min) +kubeconfig: + #!/usr/bin/env bash + set -euo pipefail + IP=$(just server-ip) + echo "fetching kubeconfig from $IP..." + + # wait for k3s to be ready + until ssh -o ConnectTimeout=5 -o StrictHostKeyChecking=accept-new root@$IP test -f /run/k3s-ready 2>/dev/null; do + echo " waiting for k3s..." + sleep 5 + done + + scp root@$IP:/etc/rancher/k3s/k3s.yaml kubeconfig.yaml + # replace localhost with public IP + if [[ "$(uname)" == "Darwin" ]]; then + sed -i '' "s|127.0.0.1|$IP|g" kubeconfig.yaml + else + sed -i "s|127.0.0.1|$IP|g" kubeconfig.yaml + fi + chmod 600 kubeconfig.yaml + echo "kubeconfig written to kubeconfig.yaml" + kubectl get nodes + +# --- deployment --- + +# deploy everything to the cluster +deploy: + #!/usr/bin/env bash + set -euo pipefail + + helm repo add bjw-s https://bjw-s-labs.github.io/helm-charts + helm repo add bitnami https://charts.bitnami.com/bitnami + helm repo add jetstack https://charts.jetstack.io + helm repo add prometheus-community https://prometheus-community.github.io/helm-charts + helm repo update + + : "${RELAY_DOMAIN:?set RELAY_DOMAIN}" + : "${RELAY_ADMIN_PASSWORD:?set RELAY_ADMIN_PASSWORD}" + : "${POSTGRES_PASSWORD:?set POSTGRES_PASSWORD}" + : "${LETSENCRYPT_EMAIL:?set LETSENCRYPT_EMAIL}" + + echo "==> creating namespace" + kubectl create namespace relay --dry-run=client -o yaml | kubectl apply -f - + + echo "==> installing cert-manager" + helm upgrade --install cert-manager jetstack/cert-manager \ + --namespace cert-manager --create-namespace \ + --set crds.enabled=true \ + --wait + + echo "==> applying cluster issuer" + sed "s|you@example.com|$LETSENCRYPT_EMAIL|g" ../shared/deploy/cluster-issuer.yaml \ + | kubectl apply -f - + + echo "==> installing postgresql" + helm upgrade --install relay-db bitnami/postgresql \ + --namespace relay \ + --values ../shared/deploy/postgres-values.yaml \ + --set auth.password="$POSTGRES_PASSWORD" \ + --wait + + echo "==> creating relay secret" + kubectl create secret generic relay-secret \ + --namespace relay \ + --from-literal=DATABASE_URL="postgres://relay:${POSTGRES_PASSWORD}@relay-db-postgresql.relay.svc.cluster.local:5432/relay" \ + --from-literal=RELAY_ADMIN_PASSWORD="$RELAY_ADMIN_PASSWORD" \ + --dry-run=client -o yaml | kubectl apply -f - + + echo "==> installing relay" + helm upgrade --install relay bjw-s/app-template \ + --namespace relay \ + --values deploy/relay-values.yaml \ + --wait --timeout 5m + + echo "==> applying ingress" + sed "s|RELAY_DOMAIN_PLACEHOLDER|$RELAY_DOMAIN|g" deploy/ingress.yaml \ + | kubectl apply -f - + + GRAFANA_DOMAIN="${GRAFANA_DOMAIN:-relay-metrics.waow.tech}" + + echo "==> installing monitoring stack" + kubectl create namespace monitoring --dry-run=client -o yaml | kubectl apply -f - + kubectl create configmap relay-dashboard \ + --namespace monitoring \ + --from-file=relay-dashboard.json=deploy/relay-dashboard.json \ + --dry-run=client -o yaml | kubectl apply -f - + helm upgrade --install kube-prometheus-stack prometheus-community/kube-prometheus-stack \ + --namespace monitoring \ + --values deploy/monitoring-values.yaml \ + --set grafana.adminPassword="${GRAFANA_ADMIN_PASSWORD:-prom-operator}" \ + --wait --timeout 5m + kubectl apply -f deploy/relay-servicemonitor.yaml + + echo "==> applying grafana ingress" + sed "s|GRAFANA_DOMAIN_PLACEHOLDER|$GRAFANA_DOMAIN|g" ../shared/deploy/grafana-ingress.yaml \ + | kubectl apply -f - + + echo "==> creating collectiondir secret" + kubectl create secret generic collectiondir-secret \ + --namespace relay \ + --from-literal=COLLECTIONS_ADMIN_TOKEN="${COLLECTIONDIR_ADMIN_TOKEN:-}" \ + --dry-run=client -o yaml | kubectl apply -f - + + echo "==> installing collectiondir" + helm upgrade --install collectiondir bjw-s/app-template \ + --namespace relay \ + --values deploy/collectiondir-values.yaml \ + --wait --timeout 5m + kubectl apply -f deploy/collectiondir-servicemonitor.yaml + + echo "==> installing reconnect cronjob" + kubectl apply -f deploy/reconnect-cronjob.yaml + + echo "==> installing jetstream" + JETSTREAM_DOMAIN="${JETSTREAM_DOMAIN:-jetstream.waow.tech}" + helm upgrade --install jetstream bjw-s/app-template \ + --namespace relay \ + --values deploy/jetstream-values.yaml \ + --wait --timeout 5m + + echo "==> applying jetstream ingress" + sed "s|JETSTREAM_DOMAIN_PLACEHOLDER|$JETSTREAM_DOMAIN|g" deploy/jetstream-ingress.yaml \ + | kubectl apply -f - + kubectl apply -f deploy/jetstream-servicemonitor.yaml + + echo "" + echo "done. point DNS:" + echo " $RELAY_DOMAIN -> $(just server-ip)" + echo " $GRAFANA_DOMAIN -> $(just server-ip)" + echo " $JETSTREAM_DOMAIN -> $(just server-ip)" + echo "then check:" + echo " curl https://$RELAY_DOMAIN/xrpc/_health" + echo " curl https://$GRAFANA_DOMAIN" + echo " curl https://$JETSTREAM_DOMAIN" + +# deploy only the monitoring stack (grafana + prometheus) +deploy-monitoring: + #!/usr/bin/env bash + set -euo pipefail + + helm repo add prometheus-community https://prometheus-community.github.io/helm-charts + helm repo update + + GRAFANA_DOMAIN="${GRAFANA_DOMAIN:-relay-metrics.waow.tech}" + + echo "==> installing monitoring stack" + kubectl create namespace monitoring --dry-run=client -o yaml | kubectl apply -f - + kubectl create configmap relay-dashboard \ + --namespace monitoring \ + --from-file=relay-dashboard.json=deploy/relay-dashboard.json \ + --dry-run=client -o yaml | kubectl apply -f - + helm upgrade --install kube-prometheus-stack prometheus-community/kube-prometheus-stack \ + --namespace monitoring \ + --values deploy/monitoring-values.yaml \ + --set grafana.adminPassword="${GRAFANA_ADMIN_PASSWORD:-prom-operator}" \ + --wait --timeout 5m + kubectl apply -f deploy/relay-servicemonitor.yaml + + echo "==> applying grafana ingress" + sed "s|GRAFANA_DOMAIN_PLACEHOLDER|$GRAFANA_DOMAIN|g" ../shared/deploy/grafana-ingress.yaml \ + | kubectl apply -f - + + echo "done." + +# seed the relay with hosts from the network (includes restart so slurper picks them up) +bootstrap: + kubectl exec -n relay deploy/relay -- /relay pull-hosts --relay-host https://relay1.us-west.bsky.network + kubectl rollout restart deploy/relay -n relay + kubectl rollout status deploy/relay -n relay --timeout=2m + +# sync PDS host list from upstream (run periodically to discover new hosts) +sync-hosts: + kubectl exec -n relay deploy/relay -- /relay pull-hosts --relay-host https://relay1.us-west.bsky.network + +# --- status --- + +# check the state of everything +status: + @echo "==> nodes" + @kubectl get nodes + @echo "" + @echo "==> pods" + @kubectl get pods -n relay + @echo "" + @echo "==> relay health (in-cluster)" + @kubectl exec -n relay deploy/relay -- curl -sf localhost:2470/xrpc/_health 2>/dev/null || echo "(relay not ready yet)" + +# tail relay logs +logs: + kubectl logs -n relay deploy/relay -f + +# check relay health via public endpoint +health: + #!/usr/bin/env bash + : "${RELAY_DOMAIN:?set RELAY_DOMAIN}" + curl -sf "https://$RELAY_DOMAIN/xrpc/_health" | jq . + +# get the grafana admin password from the cluster +grafana-password: + @kubectl get secret -n monitoring kube-prometheus-stack-grafana -o jsonpath="{.data.admin-password}" | base64 -d && echo + +# --- images --- + +# build and push collectiondir image from indigo source +collectiondir-publish: + #!/usr/bin/env bash + set -euo pipefail + TMPDIR=$(mktemp -d) + trap "rm -rf $TMPDIR" EXIT + git clone --depth 1 https://github.com/bluesky-social/indigo "$TMPDIR" + docker build --platform linux/amd64 \ + -f "$TMPDIR/cmd/collectiondir/Dockerfile" \ + -t atcr.io/zzstoatzz.io/collectiondir:latest "$TMPDIR" + ATCR_AUTO_AUTH=1 docker push atcr.io/zzstoatzz.io/collectiondir:latest + +# --- scripts --- + +# reconnect relay to all known PDS hosts (run periodically, e.g. every 4 hours) +reconnect *args: + #!/usr/bin/env bash + set -euo pipefail + : "${RELAY_ADMIN_PASSWORD:?set RELAY_ADMIN_PASSWORD}" + ../scripts/reconnect --password "$RELAY_ADMIN_PASSWORD" {{ args }} + +# consume the firehose (default: 10s of bsky posts) +firehose *args: + ../scripts/firehose {{ args }} + +# consume the jetstream (default: 10s of all events) +jetstream *args: + ../scripts/jetstream {{ args }} + +# backfill collectiondir with full network PDS hosts +# pass --hosts to use a specific host list, otherwise extracts from relay +backfill *args: + #!/usr/bin/env bash + set -euo pipefail + : "${COLLECTIONDIR_ADMIN_TOKEN:?set COLLECTIONDIR_ADMIN_TOKEN}" + + PIDS=() + cleanup() { kill "${PIDS[@]}" 2>/dev/null; } + trap cleanup EXIT + + # port-forward to collectiondir + kubectl port-forward -n relay svc/collectiondir 2510:2510 >/dev/null 2>&1 & + PIDS+=($!) + + EXTRA_ARGS=({{ args }}) + + # if --hosts not provided, extract from relay + if ! printf '%s\n' "${EXTRA_ARGS[@]}" | grep -q '^--hosts$'; then + : "${RELAY_ADMIN_PASSWORD:?set RELAY_ADMIN_PASSWORD}" + kubectl port-forward -n relay svc/relay 12470:2470 >/dev/null 2>&1 & + PIDS+=($!) + sleep 3 + + echo "==> extracting connected PDS host list from relay" + curl -sf -u "admin:$RELAY_ADMIN_PASSWORD" http://localhost:12470/admin/pds/list \ + | jq -r '.[] | select(.HasActiveConnection) | .Host' > /tmp/relay-hosts.txt + TOTAL=$(curl -sf -u "admin:$RELAY_ADMIN_PASSWORD" http://localhost:12470/admin/pds/list | jq 'length') + echo " $(wc -l < /tmp/relay-hosts.txt | tr -d ' ') connected hosts (of $TOTAL total)" + echo + EXTRA_ARGS+=(--hosts /tmp/relay-hosts.txt) + else + sleep 3 + fi + + ../scripts/backfill \ + --token "$COLLECTIONDIR_ADMIN_TOKEN" \ + "${EXTRA_ARGS[@]}" diff --git a/justfile b/justfile index d347247..8a983d6 100644 --- a/justfile +++ b/justfile @@ -1,66 +1,15 @@ # ATProto relay deployment -# required env vars: HCLOUD_TOKEN, RELAY_DOMAIN, RELAY_ADMIN_PASSWORD, POSTGRES_PASSWORD, LETSENCRYPT_EMAIL -# optional env vars: GRAFANA_DOMAIN (default: relay-metrics.waow.tech), GRAFANA_ADMIN_PASSWORD, JETSTREAM_DOMAIN (default: jetstream.waow.tech) -# zlay env vars: ZLAY_DOMAIN, ZLAY_ADMIN_PASSWORD, ZLAY_POSTGRES_PASSWORD, LETSENCRYPT_EMAIL +# usage: just indigo | just zlay -set dotenv-load +set dotenv-load := true -export KUBECONFIG := justfile_directory() / "kubeconfig.yaml" +mod indigo +mod zlay # show available recipes default: @just --list -# --- infrastructure --- - -# initialize terraform -init: - terraform -chdir=infra init - -# create the hetzner server with k3s -infra: - terraform -chdir=infra apply -var="hcloud_token=$HCLOUD_TOKEN" - -# destroy all infrastructure -destroy: - terraform -chdir=infra destroy -var="hcloud_token=$HCLOUD_TOKEN" - -# get the server IP from terraform -server-ip: - @terraform -chdir=infra output -raw server_ip - -# ssh into the server -ssh: - ssh root@$(just server-ip) - -# --- cluster access --- - -# fetch kubeconfig from the server (run after cloud-init finishes, ~2 min) -kubeconfig: - #!/usr/bin/env bash - set -euo pipefail - IP=$(just server-ip) - echo "fetching kubeconfig from $IP..." - - # wait for k3s to be ready - until ssh -o ConnectTimeout=5 -o StrictHostKeyChecking=accept-new root@$IP test -f /run/k3s-ready 2>/dev/null; do - echo " waiting for k3s..." - sleep 5 - done - - scp root@$IP:/etc/rancher/k3s/k3s.yaml kubeconfig.yaml - # replace localhost with public IP - if [[ "$(uname)" == "Darwin" ]]; then - sed -i '' "s|127.0.0.1|$IP|g" kubeconfig.yaml - else - sed -i "s|127.0.0.1|$IP|g" kubeconfig.yaml - fi - chmod 600 kubeconfig.yaml - echo "kubeconfig written to kubeconfig.yaml" - kubectl get nodes - -# --- deployment --- - # add required helm repos helm-repos: helm repo add bjw-s https://bjw-s-labs.github.io/helm-charts @@ -68,426 +17,3 @@ helm-repos: helm repo add jetstack https://charts.jetstack.io helm repo add prometheus-community https://prometheus-community.github.io/helm-charts helm repo update - -# deploy everything to the cluster -deploy: helm-repos - #!/usr/bin/env bash - set -euo pipefail - - : "${RELAY_DOMAIN:?set RELAY_DOMAIN}" - : "${RELAY_ADMIN_PASSWORD:?set RELAY_ADMIN_PASSWORD}" - : "${POSTGRES_PASSWORD:?set POSTGRES_PASSWORD}" - : "${LETSENCRYPT_EMAIL:?set LETSENCRYPT_EMAIL}" - - echo "==> creating namespace" - kubectl create namespace relay --dry-run=client -o yaml | kubectl apply -f - - - echo "==> installing cert-manager" - helm upgrade --install cert-manager jetstack/cert-manager \ - --namespace cert-manager --create-namespace \ - --set crds.enabled=true \ - --wait - - echo "==> applying cluster issuer" - sed "s|you@example.com|$LETSENCRYPT_EMAIL|g" deploy/cluster-issuer.yaml \ - | kubectl apply -f - - - echo "==> installing postgresql" - helm upgrade --install relay-db bitnami/postgresql \ - --namespace relay \ - --values deploy/postgres-values.yaml \ - --set auth.password="$POSTGRES_PASSWORD" \ - --wait - - echo "==> creating relay secret" - kubectl create secret generic relay-secret \ - --namespace relay \ - --from-literal=DATABASE_URL="postgres://relay:${POSTGRES_PASSWORD}@relay-db-postgresql.relay.svc.cluster.local:5432/relay" \ - --from-literal=RELAY_ADMIN_PASSWORD="$RELAY_ADMIN_PASSWORD" \ - --dry-run=client -o yaml | kubectl apply -f - - - echo "==> installing relay" - helm upgrade --install relay bjw-s/app-template \ - --namespace relay \ - --values deploy/relay-values.yaml \ - --wait --timeout 5m - - echo "==> applying ingress" - sed "s|RELAY_DOMAIN_PLACEHOLDER|$RELAY_DOMAIN|g" deploy/ingress.yaml \ - | kubectl apply -f - - - GRAFANA_DOMAIN="${GRAFANA_DOMAIN:-relay-metrics.waow.tech}" - - echo "==> installing monitoring stack" - kubectl create namespace monitoring --dry-run=client -o yaml | kubectl apply -f - - kubectl create configmap relay-dashboard \ - --namespace monitoring \ - --from-file=relay-dashboard.json=deploy/relay-dashboard.json \ - --dry-run=client -o yaml | kubectl apply -f - - helm upgrade --install kube-prometheus-stack prometheus-community/kube-prometheus-stack \ - --namespace monitoring \ - --values deploy/monitoring-values.yaml \ - --set grafana.adminPassword="${GRAFANA_ADMIN_PASSWORD:-prom-operator}" \ - --wait --timeout 5m - kubectl apply -f deploy/relay-servicemonitor.yaml - - echo "==> applying grafana ingress" - sed "s|GRAFANA_DOMAIN_PLACEHOLDER|$GRAFANA_DOMAIN|g" deploy/grafana-ingress.yaml \ - | kubectl apply -f - - - echo "==> creating collectiondir secret" - kubectl create secret generic collectiondir-secret \ - --namespace relay \ - --from-literal=COLLECTIONS_ADMIN_TOKEN="${COLLECTIONDIR_ADMIN_TOKEN:-}" \ - --dry-run=client -o yaml | kubectl apply -f - - - echo "==> installing collectiondir" - helm upgrade --install collectiondir bjw-s/app-template \ - --namespace relay \ - --values deploy/collectiondir-values.yaml \ - --wait --timeout 5m - kubectl apply -f deploy/collectiondir-servicemonitor.yaml - - echo "==> installing reconnect cronjob" - kubectl apply -f deploy/reconnect-cronjob.yaml - - echo "==> installing jetstream" - JETSTREAM_DOMAIN="${JETSTREAM_DOMAIN:-jetstream.waow.tech}" - helm upgrade --install jetstream bjw-s/app-template \ - --namespace relay \ - --values deploy/jetstream-values.yaml \ - --wait --timeout 5m - - echo "==> applying jetstream ingress" - sed "s|JETSTREAM_DOMAIN_PLACEHOLDER|$JETSTREAM_DOMAIN|g" deploy/jetstream-ingress.yaml \ - | kubectl apply -f - - kubectl apply -f deploy/jetstream-servicemonitor.yaml - - echo "" - echo "done. point DNS:" - echo " $RELAY_DOMAIN -> $(just server-ip)" - echo " $GRAFANA_DOMAIN -> $(just server-ip)" - echo " $JETSTREAM_DOMAIN -> $(just server-ip)" - echo "then check:" - echo " curl https://$RELAY_DOMAIN/xrpc/_health" - echo " curl https://$GRAFANA_DOMAIN" - echo " curl https://$JETSTREAM_DOMAIN" - -# deploy only the monitoring stack (grafana + prometheus) -deploy-monitoring: helm-repos - #!/usr/bin/env bash - set -euo pipefail - - GRAFANA_DOMAIN="${GRAFANA_DOMAIN:-relay-metrics.waow.tech}" - - echo "==> installing monitoring stack" - kubectl create namespace monitoring --dry-run=client -o yaml | kubectl apply -f - - kubectl create configmap relay-dashboard \ - --namespace monitoring \ - --from-file=relay-dashboard.json=deploy/relay-dashboard.json \ - --dry-run=client -o yaml | kubectl apply -f - - helm upgrade --install kube-prometheus-stack prometheus-community/kube-prometheus-stack \ - --namespace monitoring \ - --values deploy/monitoring-values.yaml \ - --set grafana.adminPassword="${GRAFANA_ADMIN_PASSWORD:-prom-operator}" \ - --wait --timeout 5m - kubectl apply -f deploy/relay-servicemonitor.yaml - - echo "==> applying grafana ingress" - sed "s|GRAFANA_DOMAIN_PLACEHOLDER|$GRAFANA_DOMAIN|g" deploy/grafana-ingress.yaml \ - | kubectl apply -f - - - echo "done." - -# seed the relay with hosts from the network (includes restart so slurper picks them up) -bootstrap: - kubectl exec -n relay deploy/relay -- /relay pull-hosts --relay-host https://relay1.us-west.bsky.network - kubectl rollout restart deploy/relay -n relay - kubectl rollout status deploy/relay -n relay --timeout=2m - -# sync PDS host list from upstream (run periodically to discover new hosts) -sync-hosts: - kubectl exec -n relay deploy/relay -- /relay pull-hosts --relay-host https://relay1.us-west.bsky.network - -# --- status --- - -# check the state of everything -status: - @echo "==> nodes" - @kubectl get nodes - @echo "" - @echo "==> pods" - @kubectl get pods -n relay - @echo "" - @echo "==> relay health (in-cluster)" - @kubectl exec -n relay deploy/relay -- curl -sf localhost:2470/xrpc/_health 2>/dev/null || echo "(relay not ready yet)" - -# tail relay logs -logs: - kubectl logs -n relay deploy/relay -f - -# check relay health via public endpoint -health: - #!/usr/bin/env bash - : "${RELAY_DOMAIN:?set RELAY_DOMAIN}" - curl -sf "https://$RELAY_DOMAIN/xrpc/_health" | jq . - -# get the grafana admin password from the cluster -grafana-password: - @kubectl get secret -n monitoring kube-prometheus-stack-grafana -o jsonpath="{.data.admin-password}" | base64 -d && echo - -# --- images --- - -# build and push collectiondir image from indigo source -collectiondir-publish: - #!/usr/bin/env bash - set -euo pipefail - TMPDIR=$(mktemp -d) - trap "rm -rf $TMPDIR" EXIT - git clone --depth 1 https://github.com/bluesky-social/indigo "$TMPDIR" - docker build --platform linux/amd64 \ - -f "$TMPDIR/cmd/collectiondir/Dockerfile" \ - -t atcr.io/zzstoatzz.io/collectiondir:latest "$TMPDIR" - ATCR_AUTO_AUTH=1 docker push atcr.io/zzstoatzz.io/collectiondir:latest - -# --- scripts --- - -# reconnect relay to all known PDS hosts (run periodically, e.g. every 4 hours) -reconnect *args: - #!/usr/bin/env bash - set -euo pipefail - : "${RELAY_ADMIN_PASSWORD:?set RELAY_ADMIN_PASSWORD}" - ./scripts/reconnect --password "$RELAY_ADMIN_PASSWORD" {{ args }} - -# consume the firehose (default: 10s of bsky posts) -firehose *args: - ./scripts/firehose {{ args }} - -# consume the jetstream (default: 10s of all events) -jetstream *args: - ./scripts/jetstream {{ args }} - -# backfill collectiondir with full network PDS hosts -# pass --hosts to use a specific host list, otherwise extracts from relay -backfill *args: - #!/usr/bin/env bash - set -euo pipefail - : "${COLLECTIONDIR_ADMIN_TOKEN:?set COLLECTIONDIR_ADMIN_TOKEN}" - - PIDS=() - cleanup() { kill "${PIDS[@]}" 2>/dev/null; } - trap cleanup EXIT - - # port-forward to collectiondir - kubectl port-forward -n relay svc/collectiondir 2510:2510 >/dev/null 2>&1 & - PIDS+=($!) - - EXTRA_ARGS=({{ args }}) - - # if --hosts not provided, extract from relay - if ! printf '%s\n' "${EXTRA_ARGS[@]}" | grep -q '^--hosts$'; then - : "${RELAY_ADMIN_PASSWORD:?set RELAY_ADMIN_PASSWORD}" - kubectl port-forward -n relay svc/relay 12470:2470 >/dev/null 2>&1 & - PIDS+=($!) - sleep 3 - - echo "==> extracting connected PDS host list from relay" - curl -sf -u "admin:$RELAY_ADMIN_PASSWORD" http://localhost:12470/admin/pds/list \ - | jq -r '.[] | select(.HasActiveConnection) | .Host' > /tmp/relay-hosts.txt - TOTAL=$(curl -sf -u "admin:$RELAY_ADMIN_PASSWORD" http://localhost:12470/admin/pds/list | jq 'length') - echo " $(wc -l < /tmp/relay-hosts.txt | tr -d ' ') connected hosts (of $TOTAL total)" - echo - EXTRA_ARGS+=(--hosts /tmp/relay-hosts.txt) - else - sleep 3 - fi - - ./scripts/backfill \ - --token "$COLLECTIONDIR_ADMIN_TOKEN" \ - "${EXTRA_ARGS[@]}" - -# === zlay (zig relay) === - -export ZLAY_KUBECONFIG := justfile_directory() / "zlay-kubeconfig.yaml" - -# initialize zlay terraform -zlay-init: - terraform -chdir=infra/zlay init - -# create the zlay hetzner server with k3s -zlay-infra: - terraform -chdir=infra/zlay apply -var="hcloud_token=$HCLOUD_TOKEN" - -# destroy zlay infrastructure -zlay-destroy: - terraform -chdir=infra/zlay destroy -var="hcloud_token=$HCLOUD_TOKEN" - -# get the zlay server IP -zlay-server-ip: - @terraform -chdir=infra/zlay output -raw server_ip - -# ssh into the zlay server -zlay-ssh: - ssh root@$(just zlay-server-ip) - -# fetch zlay kubeconfig -zlay-kubeconfig: - #!/usr/bin/env bash - set -euo pipefail - IP=$(just zlay-server-ip) - echo "fetching kubeconfig from $IP..." - until ssh -o ConnectTimeout=5 -o StrictHostKeyChecking=accept-new root@$IP test -f /run/k3s-ready 2>/dev/null; do - echo " waiting for k3s..." - sleep 5 - done - scp root@$IP:/etc/rancher/k3s/k3s.yaml zlay-kubeconfig.yaml - if [[ "$(uname)" == "Darwin" ]]; then - sed -i '' "s|127.0.0.1|$IP|g" zlay-kubeconfig.yaml - else - sed -i "s|127.0.0.1|$IP|g" zlay-kubeconfig.yaml - fi - chmod 600 zlay-kubeconfig.yaml - echo "kubeconfig written to zlay-kubeconfig.yaml" - KUBECONFIG=zlay-kubeconfig.yaml kubectl get nodes - -# build and push zlay image via docker (slow on mac — prefer zlay-publish-remote) -zlay-publish-docker: - #!/usr/bin/env bash - set -euo pipefail - TMPDIR=$(mktemp -d) - trap "rm -rf $TMPDIR" EXIT - git clone --depth 1 https://tangled.org/zzstoatzz.io/zlay "$TMPDIR" - cd "$TMPDIR" - TAG=$(git rev-parse --short HEAD) - IMAGE="atcr.io/zzstoatzz.io/zlay:${TAG}" - docker build --platform linux/amd64 -t "${IMAGE}" . - ATCR_AUTO_AUTH=1 docker push "${IMAGE}" - echo "==> pushed ${IMAGE}" - -# build zlay on the server and import into k3s containerd (fast — native x86_64 build) -# usage: just zlay-publish-remote (debug build) -# just zlay-publish-remote ReleaseSafe (optimized — needs 8 MiB stacks) -zlay-publish-remote optimize="": - #!/usr/bin/env bash - set -euo pipefail - ssh root@$(just zlay-server-ip) <<'DEPLOY' - set -euo pipefail - cd /opt/zlay - git pull --ff-only - - TAG=$(git rev-parse --short HEAD) - IMAGE="atcr.io/zzstoatzz.io/zlay:{{ if optimize != "" { optimize + "-" } else { "debug-" } }}${TAG}" - - echo "==> building binary (${TAG}{{ if optimize != "" { ", " + optimize } else { ", debug" } }})" - zig build {{ if optimize != "" { "-Doptimize=" + optimize + " " } else { "" } }}-Dtarget=x86_64-linux-gnu - - echo "==> building container image (${IMAGE})" - buildah bud -t "${IMAGE}" -f Dockerfile.runtime . - - echo "==> importing into k3s containerd" - buildah push "${IMAGE}" docker-archive:/tmp/zlay.tar:"${IMAGE}" - ctr -n k8s.io images import /tmp/zlay.tar - rm -f /tmp/zlay.tar - - echo "==> updating deployment image" - kubectl set image deployment/zlay -n zlay main="${IMAGE}" - kubectl rollout status deployment/zlay -n zlay --timeout=120s - - echo "==> deployed ${IMAGE}" - DEPLOY - -# deploy zlay to its k3s cluster -zlay-deploy: helm-repos - #!/usr/bin/env bash - set -euo pipefail - export KUBECONFIG="$ZLAY_KUBECONFIG" - - : "${ZLAY_DOMAIN:?set ZLAY_DOMAIN}" - : "${ZLAY_POSTGRES_PASSWORD:?set ZLAY_POSTGRES_PASSWORD}" - : "${LETSENCRYPT_EMAIL:?set LETSENCRYPT_EMAIL}" - ZLAY_ADMIN_PASSWORD="${ZLAY_ADMIN_PASSWORD:-}" - - echo "==> creating namespace" - kubectl create namespace zlay --dry-run=client -o yaml | kubectl apply -f - - - echo "==> installing cert-manager" - helm upgrade --install cert-manager jetstack/cert-manager \ - --namespace cert-manager --create-namespace \ - --set crds.enabled=true \ - --wait - - echo "==> applying cluster issuer" - sed "s|you@example.com|$LETSENCRYPT_EMAIL|g" deploy/cluster-issuer.yaml \ - | kubectl apply -f - - - echo "==> installing postgresql" - helm upgrade --install zlay-db bitnami/postgresql \ - --namespace zlay \ - --values deploy/postgres-values.yaml \ - --set auth.password="$ZLAY_POSTGRES_PASSWORD" \ - --wait - - echo "==> creating zlay secret" - kubectl create secret generic zlay-secret \ - --namespace zlay \ - --from-literal=DATABASE_URL="postgres://relay:${ZLAY_POSTGRES_PASSWORD}@zlay-db-postgresql.zlay.svc.cluster.local:5432/relay" \ - --from-literal=RELAY_ADMIN_PASSWORD="$ZLAY_ADMIN_PASSWORD" \ - --dry-run=client -o yaml | kubectl apply -f - - - echo "==> installing zlay" - helm upgrade --install zlay bjw-s/app-template \ - --namespace zlay \ - --values deploy/zlay-values.yaml \ - --wait --timeout 5m - - echo "==> applying ingress" - sed "s|ZLAY_DOMAIN_PLACEHOLDER|$ZLAY_DOMAIN|g" deploy/zlay-ingress.yaml \ - | kubectl apply -f - - - echo "==> installing monitoring" - ZLAY_METRICS_DOMAIN="${ZLAY_METRICS_DOMAIN:-zlay-metrics.waow.tech}" - kubectl create namespace monitoring --dry-run=client -o yaml | kubectl apply -f - - kubectl create configmap zlay-dashboard \ - --namespace monitoring \ - --from-file=zlay-dashboard.json=deploy/zlay-dashboard.json \ - --dry-run=client -o yaml | kubectl apply -f - - helm upgrade --install kube-prometheus-stack prometheus-community/kube-prometheus-stack \ - --namespace monitoring \ - --values deploy/zlay-monitoring-values.yaml \ - --set grafana.adminPassword="${GRAFANA_ADMIN_PASSWORD:-prom-operator}" \ - --wait --timeout 5m - kubectl apply -f deploy/zlay-servicemonitor.yaml - - echo "==> applying grafana ingress" - sed "s|GRAFANA_DOMAIN_PLACEHOLDER|$ZLAY_METRICS_DOMAIN|g" deploy/grafana-ingress.yaml \ - | kubectl apply -f - - - echo "" - echo "done. point DNS:" - echo " $ZLAY_DOMAIN -> $(just zlay-server-ip)" - echo " $ZLAY_METRICS_DOMAIN -> $(just zlay-server-ip)" - echo "then check:" - echo " curl https://$ZLAY_DOMAIN/_health" - -# check zlay status -zlay-status: - #!/usr/bin/env bash - export KUBECONFIG="$ZLAY_KUBECONFIG" - echo "==> nodes" - kubectl get nodes - echo "" - echo "==> pods" - kubectl get pods -n zlay - echo "" - echo "==> health" - curl -sf "https://$ZLAY_DOMAIN/_health" | jq . || echo "(zlay not ready)" - -# tail zlay logs -zlay-logs: - KUBECONFIG="$ZLAY_KUBECONFIG" kubectl logs -n zlay deploy/zlay -f - -# check zlay health via public endpoint -zlay-health: - #!/usr/bin/env bash - : "${ZLAY_DOMAIN:?set ZLAY_DOMAIN}" - curl -sf "https://$ZLAY_DOMAIN/_health" | jq . diff --git a/deploy/cluster-issuer.yaml b/shared/deploy/cluster-issuer.yaml similarity index 100% rename from deploy/cluster-issuer.yaml rename to shared/deploy/cluster-issuer.yaml diff --git a/deploy/grafana-ingress.yaml b/shared/deploy/grafana-ingress.yaml similarity index 100% rename from deploy/grafana-ingress.yaml rename to shared/deploy/grafana-ingress.yaml diff --git a/deploy/postgres-values.yaml b/shared/deploy/postgres-values.yaml similarity index 100% rename from deploy/postgres-values.yaml rename to shared/deploy/postgres-values.yaml diff --git a/deploy/zlay-dashboard.json b/zlay/deploy/zlay-dashboard.json similarity index 100% rename from deploy/zlay-dashboard.json rename to zlay/deploy/zlay-dashboard.json diff --git a/deploy/zlay-ingress.yaml b/zlay/deploy/zlay-ingress.yaml similarity index 100% rename from deploy/zlay-ingress.yaml rename to zlay/deploy/zlay-ingress.yaml diff --git a/deploy/zlay-monitoring-values.yaml b/zlay/deploy/zlay-monitoring-values.yaml similarity index 100% rename from deploy/zlay-monitoring-values.yaml rename to zlay/deploy/zlay-monitoring-values.yaml diff --git a/deploy/zlay-reconnect-cronjob.yaml b/zlay/deploy/zlay-reconnect-cronjob.yaml similarity index 100% rename from deploy/zlay-reconnect-cronjob.yaml rename to zlay/deploy/zlay-reconnect-cronjob.yaml diff --git a/deploy/zlay-servicemonitor.yaml b/zlay/deploy/zlay-servicemonitor.yaml similarity index 100% rename from deploy/zlay-servicemonitor.yaml rename to zlay/deploy/zlay-servicemonitor.yaml diff --git a/deploy/zlay-values.yaml b/zlay/deploy/zlay-values.yaml similarity index 100% rename from deploy/zlay-values.yaml rename to zlay/deploy/zlay-values.yaml diff --git a/infra/zlay/main.tf b/zlay/infra/main.tf similarity index 100% rename from infra/zlay/main.tf rename to zlay/infra/main.tf diff --git a/infra/zlay/outputs.tf b/zlay/infra/outputs.tf similarity index 100% rename from infra/zlay/outputs.tf rename to zlay/infra/outputs.tf diff --git a/infra/zlay/variables.tf b/zlay/infra/variables.tf similarity index 100% rename from infra/zlay/variables.tf rename to zlay/infra/variables.tf diff --git a/infra/zlay/versions.tf b/zlay/infra/versions.tf similarity index 100% rename from infra/zlay/versions.tf rename to zlay/infra/versions.tf diff --git a/zlay/justfile b/zlay/justfile new file mode 100644 index 0000000..5ec0821 --- /dev/null +++ b/zlay/justfile @@ -0,0 +1,205 @@ +# zlay (zig) relay deployment +# required env vars: HCLOUD_TOKEN, ZLAY_DOMAIN, ZLAY_POSTGRES_PASSWORD, LETSENCRYPT_EMAIL +# optional env vars: ZLAY_ADMIN_PASSWORD, ZLAY_METRICS_DOMAIN (default: zlay-metrics.waow.tech), GRAFANA_ADMIN_PASSWORD + +export KUBECONFIG := source_directory() / "kubeconfig.yaml" + +# --- infrastructure --- + +# initialize terraform +init: + terraform -chdir=infra init + +# create the hetzner server with k3s +infra: + terraform -chdir=infra apply -var="hcloud_token=$HCLOUD_TOKEN" + +# destroy all infrastructure +destroy: + terraform -chdir=infra destroy -var="hcloud_token=$HCLOUD_TOKEN" + +# get the server IP from terraform +server-ip: + @terraform -chdir=infra output -raw server_ip + +# ssh into the server +ssh: + ssh root@$(just server-ip) + +# --- cluster access --- + +# fetch kubeconfig from the server (run after cloud-init finishes, ~2 min) +kubeconfig: + #!/usr/bin/env bash + set -euo pipefail + IP=$(just server-ip) + echo "fetching kubeconfig from $IP..." + until ssh -o ConnectTimeout=5 -o StrictHostKeyChecking=accept-new root@$IP test -f /run/k3s-ready 2>/dev/null; do + echo " waiting for k3s..." + sleep 5 + done + scp root@$IP:/etc/rancher/k3s/k3s.yaml kubeconfig.yaml + if [[ "$(uname)" == "Darwin" ]]; then + sed -i '' "s|127.0.0.1|$IP|g" kubeconfig.yaml + else + sed -i "s|127.0.0.1|$IP|g" kubeconfig.yaml + fi + chmod 600 kubeconfig.yaml + echo "kubeconfig written to kubeconfig.yaml" + kubectl get nodes + +# --- deployment --- + +# deploy zlay to its k3s cluster +deploy: + #!/usr/bin/env bash + set -euo pipefail + + helm repo add bjw-s https://bjw-s-labs.github.io/helm-charts + helm repo add bitnami https://charts.bitnami.com/bitnami + helm repo add jetstack https://charts.jetstack.io + helm repo add prometheus-community https://prometheus-community.github.io/helm-charts + helm repo update + + : "${ZLAY_DOMAIN:?set ZLAY_DOMAIN}" + : "${ZLAY_POSTGRES_PASSWORD:?set ZLAY_POSTGRES_PASSWORD}" + : "${LETSENCRYPT_EMAIL:?set LETSENCRYPT_EMAIL}" + ZLAY_ADMIN_PASSWORD="${ZLAY_ADMIN_PASSWORD:-}" + + echo "==> creating namespace" + kubectl create namespace zlay --dry-run=client -o yaml | kubectl apply -f - + + echo "==> installing cert-manager" + helm upgrade --install cert-manager jetstack/cert-manager \ + --namespace cert-manager --create-namespace \ + --set crds.enabled=true \ + --wait + + echo "==> applying cluster issuer" + sed "s|you@example.com|$LETSENCRYPT_EMAIL|g" ../shared/deploy/cluster-issuer.yaml \ + | kubectl apply -f - + + echo "==> installing postgresql" + helm upgrade --install zlay-db bitnami/postgresql \ + --namespace zlay \ + --values ../shared/deploy/postgres-values.yaml \ + --set auth.password="$ZLAY_POSTGRES_PASSWORD" \ + --wait + + echo "==> creating zlay secret" + kubectl create secret generic zlay-secret \ + --namespace zlay \ + --from-literal=DATABASE_URL="postgres://relay:${ZLAY_POSTGRES_PASSWORD}@zlay-db-postgresql.zlay.svc.cluster.local:5432/relay" \ + --from-literal=RELAY_ADMIN_PASSWORD="$ZLAY_ADMIN_PASSWORD" \ + --dry-run=client -o yaml | kubectl apply -f - + + echo "==> installing zlay" + helm upgrade --install zlay bjw-s/app-template \ + --namespace zlay \ + --values deploy/zlay-values.yaml \ + --wait --timeout 5m + + echo "==> applying ingress" + sed "s|ZLAY_DOMAIN_PLACEHOLDER|$ZLAY_DOMAIN|g" deploy/zlay-ingress.yaml \ + | kubectl apply -f - + + echo "==> installing monitoring" + ZLAY_METRICS_DOMAIN="${ZLAY_METRICS_DOMAIN:-zlay-metrics.waow.tech}" + kubectl create namespace monitoring --dry-run=client -o yaml | kubectl apply -f - + kubectl create configmap zlay-dashboard \ + --namespace monitoring \ + --from-file=zlay-dashboard.json=deploy/zlay-dashboard.json \ + --dry-run=client -o yaml | kubectl apply -f - + helm upgrade --install kube-prometheus-stack prometheus-community/kube-prometheus-stack \ + --namespace monitoring \ + --values deploy/zlay-monitoring-values.yaml \ + --set grafana.adminPassword="${GRAFANA_ADMIN_PASSWORD:-prom-operator}" \ + --wait --timeout 5m + kubectl apply -f deploy/zlay-servicemonitor.yaml + + echo "==> applying grafana ingress" + sed "s|GRAFANA_DOMAIN_PLACEHOLDER|$ZLAY_METRICS_DOMAIN|g" ../shared/deploy/grafana-ingress.yaml \ + | kubectl apply -f - + + echo "" + echo "done. point DNS:" + echo " $ZLAY_DOMAIN -> $(just server-ip)" + echo " $ZLAY_METRICS_DOMAIN -> $(just server-ip)" + echo "then check:" + echo " curl https://$ZLAY_DOMAIN/_health" + +# --- images --- + +# build and push zlay image via docker (slow on mac — prefer publish-remote) +publish-docker: + #!/usr/bin/env bash + set -euo pipefail + TMPDIR=$(mktemp -d) + trap "rm -rf $TMPDIR" EXIT + git clone --depth 1 https://tangled.org/zzstoatzz.io/zlay "$TMPDIR" + cd "$TMPDIR" + TAG=$(git rev-parse --short HEAD) + IMAGE="atcr.io/zzstoatzz.io/zlay:${TAG}" + docker build --platform linux/amd64 -t "${IMAGE}" . + ATCR_AUTO_AUTH=1 docker push "${IMAGE}" + echo "==> pushed ${IMAGE}" + +# build zlay on the server and import into k3s containerd (fast — native x86_64 build) +# usage: just zlay publish-remote (debug build) +# just zlay publish-remote ReleaseSafe (optimized — needs 8 MiB stacks) +publish-remote optimize="": + #!/usr/bin/env bash + set -euo pipefail + ssh root@$(just server-ip) <<'DEPLOY' + set -euo pipefail + cd /opt/zlay + git pull --ff-only + + TAG=$(git rev-parse --short HEAD) + IMAGE="atcr.io/zzstoatzz.io/zlay:{{ if optimize != "" { optimize + "-" } else { "debug-" } }}${TAG}" + + echo "==> building binary (${TAG}{{ if optimize != "" { ", " + optimize } else { ", debug" } }})" + zig build {{ if optimize != "" { "-Doptimize=" + optimize + " " } else { "" } }}-Dtarget=x86_64-linux-gnu + + echo "==> building container image (${IMAGE})" + buildah bud -t "${IMAGE}" -f Dockerfile.runtime . + + echo "==> importing into k3s containerd" + buildah push "${IMAGE}" docker-archive:/tmp/zlay.tar:"${IMAGE}" + ctr -n k8s.io images import /tmp/zlay.tar + rm -f /tmp/zlay.tar + + echo "==> updating deployment image" + kubectl set image deployment/zlay -n zlay main="${IMAGE}" + kubectl rollout status deployment/zlay -n zlay --timeout=120s + + echo "==> deployed ${IMAGE}" + DEPLOY + +# --- status --- + +# check zlay status +status: + #!/usr/bin/env bash + echo "==> nodes" + kubectl get nodes + echo "" + echo "==> pods" + kubectl get pods -n zlay + echo "" + echo "==> health" + curl -sf "https://$ZLAY_DOMAIN/_health" | jq . || echo "(zlay not ready)" + +# tail zlay logs +logs: + kubectl logs -n zlay deploy/zlay -f + +# check zlay health via public endpoint +health: + #!/usr/bin/env bash + : "${ZLAY_DOMAIN:?set ZLAY_DOMAIN}" + curl -sf "https://$ZLAY_DOMAIN/_health" | jq . + +# get the grafana admin password from the cluster +grafana-password: + @kubectl get secret -n monitoring kube-prometheus-stack-grafana -o jsonpath="{.data.admin-password}" | base64 -d && echo