From abdac9327d85e7cd885bbeb67b7198ce3b6b72a4 Mon Sep 17 00:00:00 2001 From: Bretton Date: Wed, 30 Sep 2026 22:59:06 -0700 Subject: [PATCH] chore(ops): nightly production backups for AppView Postgres and PDS The production host had no scheduled backups: the newest AppView dump was 27 days old and on the same disk, and cron is not installed, so nothing documented as a cron line ever ran. Changes: - scripts/pg-backup.sh: pg_dump -Fc streamed out of coves-prod-postgres, root-only from the first byte, verified with a full pg_restore read before it gets its final name, 14-day retention; a failed run keeps one fixed-name partial so retries do not pile up. - scripts/pds-backup.sh: consistent copy of the live PDS data directory via SQLite's online backup API (account.sqlite first, sequencer second), every copy integrity-checked, and the run fails if any account lacks its store or signing key. Blobs live in S3 and are not included. - scripts/systemd/: oneshot services and timers (02:47 and 03:07 UTC) with retry on failure; installed copies live in /usr/local/sbin so deploys' git pull never meets a modified checkout. - docs/PRODUCTION_BACKUPS.md: what runs, what is not covered (.env secrets, off-host copies, the S3 bucket), install steps, and fail-fast restore procedures for Postgres and the PDS. - docs/CREDENTIAL_ENCRYPTION.md: describe the -Fc dumps and 14-day retention. - .gitignore and .dockerignore: exclude backups/; the root-only dumps otherwise break `docker build` with permission denied. - Remove scripts/backup.sh (unverified gzip SQL, stale .env.prod). Co-Authored-By: Claude Opus 5.5 (1M context) --- .dockerignore | 5 +- .gitignore | 3 + docs/CREDENTIAL_ENCRYPTION.md | 10 +- docs/PRODUCTION_BACKUPS.md | 184 +++++++++++++++++++++++ scripts/backup.sh | 57 ------- scripts/pds-backup.sh | 157 +++++++++++++++++++ scripts/pg-backup.sh | 69 +++++++++ scripts/setup-production.sh | 2 +- scripts/systemd/coves-pds-backup.service | 14 ++ scripts/systemd/coves-pds-backup.timer | 9 ++ scripts/systemd/coves-pg-backup.service | 15 ++ scripts/systemd/coves-pg-backup.timer | 9 ++ 12 files changed, 472 insertions(+), 62 deletions(-) create mode 100644 docs/PRODUCTION_BACKUPS.md delete mode 100755 scripts/backup.sh create mode 100755 scripts/pds-backup.sh create mode 100755 scripts/pg-backup.sh create mode 100644 scripts/systemd/coves-pds-backup.service create mode 100644 scripts/systemd/coves-pds-backup.timer create mode 100644 scripts/systemd/coves-pg-backup.service create mode 100644 scripts/systemd/coves-pg-backup.timer diff --git a/.dockerignore b/.dockerignore index 986564a..35ec696 100644 --- a/.dockerignore +++ b/.dockerignore @@ -8,8 +8,11 @@ .git .gitignore -# Per-run and local state, never an input to a build. +# Per-run and local state, never an input to a build. backups/ holds the +# nightly dumps: root-only on the production host (a context that includes it +# fails with permission denied) and full of secrets. .ci-out/ +backups/ cache/ .claude/ .beads/ diff --git a/.gitignore b/.gitignore index dc4a881..35f82a4 100644 --- a/.gitignore +++ b/.gitignore @@ -35,6 +35,9 @@ Thumbs.db # Application data /data/ +# Production backup dumps (scripts/pg-backup.sh, scripts/pds-backup.sh): keep +# them out of `git status` and out of reach of `git clean`. +/backups/ /local_dev_data/ /test_db_data/ diff --git a/docs/CREDENTIAL_ENCRYPTION.md b/docs/CREDENTIAL_ENCRYPTION.md index 647c198..2ab1784 100644 --- a/docs/CREDENTIAL_ENCRYPTION.md +++ b/docs/CREDENTIAL_ENCRYPTION.md @@ -121,15 +121,19 @@ aggregators. the counts, and migration 046 drops the key table. Startup fails loudly instead of dropping the key if anything is left unconverted. 3. Take a fresh backup and confirm it no longer contains the table. Backups are - plain SQL, gzipped: + `pg_dump -Fc` archives (`/opt/coves/backups/coves-.dump`, root-only); + `pg_restore -f -` turns one back into SQL: ```bash - zcat .sql.gz | grep -c encryption_keys # expect 0 + sudo cat | sudo docker exec -i coves-prod-postgres pg_restore -f - \ + | grep -c encryption_keys # expect 0 ``` 4. Older backups still hold the old key beside the old ciphertext. Treat them as containing plaintext credentials: purge them once a post-cutover backup is - verified (the backup script prunes after 30 days on its own), and rotate the + verified (the backup script prunes its own dumps after 14 days, see + `docs/PRODUCTION_BACKUPS.md`; `coves_*.sql.gz` files from the old backup + script are not pruned and must be deleted by hand), and rotate the community PDS passwords and aggregator OAuth sessions if any old dump may have left the host. diff --git a/docs/PRODUCTION_BACKUPS.md b/docs/PRODUCTION_BACKUPS.md new file mode 100644 index 0000000..48c7ac9 --- /dev/null +++ b/docs/PRODUCTION_BACKUPS.md @@ -0,0 +1,184 @@ +# Production backups + +What is backed up on the production host, how it is scheduled, and how to +restore it. Tidepool's own database backup is documented in tidepool's +`DEPLOY.md` ("Backup and restore"); it runs on the same host on the same +pattern. + +## What runs + +| Unit | Time (UTC) | Script (installed copy) | Output in `/opt/coves/backups` | +|---|---|---|---| +| `tidepool-pg-backup.timer` | 02:17 | `/usr/local/sbin/tidepool-pg-backup` ← `/opt/tidepool/scripts/pg-backup.sh` | (writes to `/opt/tidepool/backups/tidepool-*.dump`) | +| `coves-pg-backup.timer` | 02:47 | `/usr/local/sbin/coves-pg-backup` ← `scripts/pg-backup.sh` | `coves-.dump` | +| `coves-pds-backup.timer` | 03:07 | `/usr/local/sbin/coves-pds-backup` ← `scripts/pds-backup.sh` | `pds-.tar.gz` | + +- **AppView Postgres**: `pg_dump -Fc`, verified with a full read + (`pg_restore -f /dev/null`) before it gets its final name. +- **PDS data directory**: every SQLite database copied with the online backup + API and checked with `PRAGMA integrity_check`, plus the per-actor signing + `key` files. This is the source of truth for every hosted community and + native account. Blobs are in object storage (`PDS_BLOBSTORE_S3_BUCKET`), not + in this archive. The run fails if any account in `account.sqlite` lacks its + store or signing key in the snapshot. +- Retention is 14 days (`RETENTION_DAYS`). +- A run that fails while the dump or archive is being written keeps only the + most recent failed partial, under a fixed name that each failed retry + overwrites: `coves-last-failed.dump.partial` or + `pds-last-failed.tar.gz.partial`. It is never deleted and is warned about + once it is older than a day. A PDS run that fails earlier leaves nothing + behind: check `systemctl --failed` and the journal. A `.pds-stage.*` staging + directory left by a killed PDS run is warned about once it is older than a + day. +- A failed run is retried every 5 minutes, up to 4 starts in 2 hours, then + left failed until the next night. A run missed while the host was down + fires at boot. +- Files are root-only (`umask 077`); the directory is mode 700. + +Output goes to the journal: `journalctl -u coves-pg-backup.service`. A failed +run shows in `systemctl --failed`. Nothing alerts on failure yet. + +## What is NOT backed up here + +- **Secrets in `.env` files**: `/opt/coves/.env` (`ENCRYPTION_KEY`, OAuth keys, + the PDS PLC rotation key and admin password, S3 keys), `/opt/tidepool/.env` + (`BRIDGE_KEK`), and the aggregators' `.env` files. A database dump without + `ENCRYPTION_KEY` restores community and aggregator credentials as unreadable + ciphertext. Keep these in a password manager, not on this host. +- **Off-host copies**: the dumps sit on the same disk array as the data they + protect. They cover a dropped database or a bad migration, not loss of the + host. +- **The S3 blob bucket**: one copy, in object storage only. + +## Install or update (on the production host) + +The scripts are installed outside `/opt/coves`, because deploys `git pull` +that checkout and a locally changed or untracked file blocks the pull. After a +change to either script, re-run the `install` lines. Tidepool's script is +installed the same way, as `/usr/local/sbin/tidepool-pg-backup`; its steps are +in tidepool's `DEPLOY.md`. + +```bash +sudo apt-get install -y sqlite3 +sudo chmod 700 /opt/coves/backups +cd /opt/coves +sudo install -m 755 scripts/pg-backup.sh /usr/local/sbin/coves-pg-backup +sudo install -m 755 scripts/pds-backup.sh /usr/local/sbin/coves-pds-backup +sudo cp scripts/systemd/coves-{pg,pds}-backup.{service,timer} /etc/systemd/system/ +sudo systemctl daemon-reload +sudo systemctl enable --now coves-pg-backup.timer coves-pds-backup.timer +sudo systemctl start coves-pg-backup.service coves-pds-backup.service # first backups now +systemctl list-timers 'coves-*' 'tidepool-*' +``` + +## Restore + +### AppView Postgres + +Restore into a freshly created database, not over the live one: +`pg_restore --clean` only drops objects that are in the archive, so tables or +columns added by a later migration survive and goose then fails on them. +`--exit-on-error --single-transaction` makes any error abort the whole restore +instead of being skipped. + +Set `DUMP` on the second line, then run the block. It stops at the first +failed step: the dump must exist and pass a full read before the AppView is +stopped or the database dropped, and the AppView is started only after the +restore succeeds (on startup it applies migrations, so it must not boot +against an empty database). The script is passed with `-c`, not on stdin, +because `pg_restore` reads the dump from stdin. The block is one +single-quoted string, so it must not contain a single quote. + +```bash +sudo bash -euo pipefail -c ' +DUMP=/opt/coves/backups/coves-.dump +POSTGRES=coves-prod-postgres +cd /opt/coves + +DB_USER="$(docker exec "$POSTGRES" printenv POSTGRES_USER)" +DB_NAME="$(docker exec "$POSTGRES" printenv POSTGRES_DB)" + +test -s "$DUMP" +docker exec -i "$POSTGRES" pg_restore -f /dev/null < "$DUMP" + +docker compose -f docker-compose.prod.yml stop appview +docker exec "$POSTGRES" dropdb -U "$DB_USER" "$DB_NAME" +docker exec "$POSTGRES" createdb -U "$DB_USER" "$DB_NAME" +docker exec -i "$POSTGRES" pg_restore --exit-on-error --single-transaction \ + --no-owner --no-acl -U "$DB_USER" -d "$DB_NAME" < "$DUMP" +docker compose -f docker-compose.prod.yml start appview +' +``` + +The AppView resumes Jetstream from the cursors in the dump. Everything indexed +after the dump was taken is lost, apart from what Jetstream can still replay. + +To check a dump without touching production, restore it into a throwaway +container: + +```bash +DUMP=/opt/coves/backups/coves-.dump # root-only: read it with sudo +sudo docker run -d --name coves-restore-drill --network none \ + -e POSTGRES_PASSWORD=drill postgres:15 +# TCP, not the socket: during first-boot init the socket answers for a +# temporary server that is about to restart. +until sudo docker exec coves-restore-drill pg_isready -h 127.0.0.1 -U postgres; do sleep 1; done +sudo cat "$DUMP" | sudo docker exec -i coves-restore-drill pg_restore --exit-on-error --no-owner --no-acl -U postgres -d postgres +sudo docker exec coves-restore-drill psql -U postgres -At \ + -c 'SELECT count(*) FROM posts' -c 'SELECT count(*) FROM communities' +# -v: the image's anonymous data volume holds a full copy of production data. +sudo docker rm -f -v coves-restore-drill +``` + +### PDS + +This procedure has not been rehearsed. + +First record the live sequencer's high-water mark (skip this if the live +database is the thing that is broken): + +```bash +sudo sqlite3 -readonly /var/lib/docker/volumes/coves-prod-pds-data/_data/sequencer.sqlite \ + 'SELECT max(seq) FROM repo_seq' +``` + +Set `ARCHIVE` on the second line, then run the block. It stops at the first +failed step: the PDS must be stopped and the archive must extract with an +`account.sqlite` before the live data is touched, and the live data is moved +aside, not deleted. As above, the script is passed with `-c`, not on stdin, +so no command inside it can swallow the rest of the script. + +```bash +sudo bash -euo pipefail -c ' +ARCHIVE=/opt/coves/backups/pds-.tar.gz +VOLUME=/var/lib/docker/volumes/coves-prod-pds-data +STAMP="$(date -u +%Y%m%dT%H%M%SZ)" + +cd /opt/coves +docker compose -f docker-compose.prod.yml stop pds +test "$(docker inspect -f "{{.State.Running}}" coves-prod-pds)" = false + +NEW="$(mktemp -d "${VOLUME}/_data.restore-${STAMP}.XXXXXX")" +tar -C "$NEW" -xzf "$ARCHIVE" +test -f "${NEW}/account.sqlite" +# mktemp creates the directory mode 700; match the live directory. +chown --reference="${VOLUME}/_data" "$NEW" +chmod --reference="${VOLUME}/_data" "$NEW" + +mv "${VOLUME}/_data" "${VOLUME}/_data.before-restore-${STAMP}" +mv "$NEW" "${VOLUME}/_data" +docker compose -f docker-compose.prod.yml start pds +' +``` + +Anything the PDS accepted after the archive was taken is lost, and the +restored PDS rewinds in two ways that relays and Jetstream notice: + +- Its sequencer restarts below the cursor they already hold for this PDS, so + new events can be skipped until `seq` passes the high-water mark recorded + above. +- Every repo's `rev` goes back, so relays may treat those repos as out of + sync. + +After the PDS is up, ask each relay that crawls it to recrawl +(`com.atproto.sync.requestCrawl` with this PDS's hostname). diff --git a/scripts/backup.sh b/scripts/backup.sh deleted file mode 100755 index c786db2..0000000 --- a/scripts/backup.sh +++ /dev/null @@ -1,57 +0,0 @@ -#!/bin/bash -# Coves Database Backup Script -# Usage: ./scripts/backup.sh -# -# Creates timestamped PostgreSQL backups in ./backups/ -# Retention: Keeps last 30 days of backups - -set -e - -SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)" -PROJECT_DIR="$(dirname "$SCRIPT_DIR")" -BACKUP_DIR="$PROJECT_DIR/backups" -COMPOSE_FILE="$PROJECT_DIR/docker-compose.prod.yml" - -# Load environment -set -a -source "$PROJECT_DIR/.env.prod" -set +a - -# Colors -GREEN='\033[0;32m' -YELLOW='\033[1;33m' -NC='\033[0m' - -log() { echo -e "${GREEN}[BACKUP]${NC} $1"; } -warn() { echo -e "${YELLOW}[WARN]${NC} $1"; } - -# Create backup directory -mkdir -p "$BACKUP_DIR" - -# Generate timestamp -TIMESTAMP=$(date +%Y%m%d_%H%M%S) -BACKUP_FILE="$BACKUP_DIR/coves_${TIMESTAMP}.sql.gz" - -log "Starting backup..." - -# Run pg_dump inside container -docker compose -f "$COMPOSE_FILE" exec -T postgres \ - pg_dump -U "$POSTGRES_USER" -d "$POSTGRES_DB" --clean --if-exists \ - | gzip > "$BACKUP_FILE" - -# Get file size -SIZE=$(du -h "$BACKUP_FILE" | cut -f1) - -log "✅ Backup complete: $BACKUP_FILE ($SIZE)" - -# Cleanup old backups (keep last 30 days) -log "Cleaning up backups older than 30 days..." -find "$BACKUP_DIR" -name "coves_*.sql.gz" -mtime +30 -delete - -# List recent backups -log "" -log "Recent backups:" -ls -lh "$BACKUP_DIR"/*.sql.gz 2>/dev/null | tail -5 - -log "" -log "To restore: gunzip -c $BACKUP_FILE | docker compose -f docker-compose.prod.yml exec -T postgres psql -U $POSTGRES_USER -d $POSTGRES_DB" diff --git a/scripts/pds-backup.sh b/scripts/pds-backup.sh new file mode 100755 index 0000000..82f99b1 --- /dev/null +++ b/scripts/pds-backup.sh @@ -0,0 +1,157 @@ +#!/usr/bin/env bash +# pds-backup.sh — nightly consistent copy of the production PDS data directory. +# +# The PDS is the source of truth for every account it hosts (communities, +# native users, aggregators): their repos live in the per-actor store.sqlite +# files and their signing keys in the per-actor `key` files. The AppView can be +# rebuilt from repos; these cannot be rebuilt from anything. Blobs are NOT in +# this directory — PDS_BLOBSTORE_S3_BUCKET puts them in object storage — so +# the archive stays small. +# +# Copying the live directory with cp or tar is not a backup: the databases run +# in WAL mode while the PDS writes to them. Each database is copied through +# SQLite's online backup API (sqlite3 .backup), which yields a consistent +# snapshot without stopping the PDS, and every copy must pass +# `PRAGMA integrity_check` before it is archived. Everything else (actor keys, +# the blocks directory) is copied as-is. +# +# Run as root (the volume is root-owned) by coves-pds-backup.timer from the +# installed copy at /usr/local/sbin/coves-pds-backup. Needs the host sqlite3 +# package. docs/PRODUCTION_BACKUPS.md has the install and restore steps. + +set -euo pipefail +umask 077 + +PDS_DATA="${PDS_DATA:-/var/lib/docker/volumes/coves-prod-pds-data/_data}" +BACKUP_DIR="${BACKUP_DIR:-/opt/coves/backups}" +RETENTION_DAYS="${RETENTION_DAYS:-14}" + +STAMP="$(date -u +%Y%m%dT%H%M%SZ)" +ARCHIVE="${BACKUP_DIR}/pds-${STAMP}.tar.gz" + +log() { echo "[$(date -u +%FT%TZ)] $*"; } + +# An empty or wrong PDS_DATA would otherwise archive nothing and report success. +if [[ ! -f "${PDS_DATA}/account.sqlite" ]]; then + log "ERROR: ${PDS_DATA}/account.sqlite not found; refusing to write an empty backup" >&2 + exit 1 +fi + +# WORK holds the NUL-delimited manifests next to STAGE so they stay out of the +# archive, which is built from STAGE alone. +# +# A failed run keeps its archive partial under one fixed name, so the timer's +# retries overwrite it instead of piling up a partial per attempt. The partial +# only exists between tar starting and the final mv, so its presence here +# means this run failed. +WORK="$(mktemp -d "${BACKUP_DIR}/.pds-stage.XXXXXX")" +FAILED_PARTIAL="${BACKUP_DIR}/pds-last-failed.tar.gz.partial" +on_exit() { + local status=$? + rm -rf "${WORK}" + if [[ -e "${ARCHIVE}.partial" ]]; then + mv -f "${ARCHIVE}.partial" "${FAILED_PARTIAL}" + log "ERROR: run failed; its partial is kept as ${FAILED_PARTIAL}" >&2 + fi + exit "${status}" +} +trap on_exit EXIT +# bash skips the EXIT trap on a signal it has no trap for; systemd stops and +# timeouts send SIGTERM, so turn it into an exit the EXIT trap sees. +trap 'exit 143' TERM +trap 'exit 130' INT +STAGE="${WORK}/data" +mkdir "${STAGE}" + +log "backup starting: ${ARCHIVE}" +cd "${PDS_DATA}" + +databases=0 +backup_db() { + local db="$1" tables check + mkdir -p "${STAGE}/$(dirname "${db}")" + # The PDS holds write locks briefly; wait for them instead of failing on SQLITE_BUSY. + # The CLI's .backup restarts the copy whenever the source is written + # mid-copy, so on a large, busy database (sequencer.sqlite) a slow or + # never-finishing run points here; `VACUUM INTO` is the alternative. + sqlite3 -readonly -cmd '.timeout 30000' "${db}" ".backup '${STAGE}/${db}'" + # integrity_check reports "ok" for an empty file, so a copy with no schema + # would otherwise pass as a valid backup. + if [[ ! -s "${STAGE}/${db}" ]]; then + log "ERROR: the copy of ${db} is empty" >&2 + exit 1 + fi + tables="$(sqlite3 "${STAGE}/${db}" 'SELECT count(*) FROM sqlite_master;')" + if (( tables == 0 )); then + log "ERROR: the copy of ${db} has no schema" >&2 + exit 1 + fi + check="$(sqlite3 "${STAGE}/${db}" 'PRAGMA integrity_check;')" + if [[ "${check}" != "ok" ]]; then + log "ERROR: integrity_check failed for the copy of ${db}: ${check}" >&2 + exit 1 + fi + databases=$((databases + 1)) +} + +# Each database is a separate snapshot, so the set is only consistent if the +# order is right. Stopping the PDS would make it exact but costs nightly +# downtime; instead account.sqlite is copied first, then the sequencer, then +# did_cache and the actor stores, then the key files, and the account list is +# checked against the staged stores and keys below. An account created mid-run +# only adds a store with no account row (harmless on restore); an account +# deleted mid-run leaves a row with no store or key, which fails the run and +# the next night retries. +backup_db ./account.sqlite +backup_db ./sequencer.sqlite + +find . -type f -name '*.sqlite' ! -path ./account.sqlite ! -path ./sequencer.sqlite -print0 > "${WORK}/databases" +while IFS= read -r -d '' db; do + backup_db "${db}" +done < "${WORK}/databases" + +# The -wal and -shm files belong to the live databases; each .backup copy +# above is already self-contained. +find . -type f ! -name '*.sqlite' ! -name '*.sqlite-wal' ! -name '*.sqlite-shm' -print0 > "${WORK}/files" +files=0 +while IFS= read -r -d '' file; do + mkdir -p "${STAGE}/$(dirname "${file}")" + cp -p "${file}" "${STAGE}/${file}" + files=$((files + 1)) +done < "${WORK}/files" + +# An account without its repo or signing key cannot be restored. +sqlite3 "${STAGE}/account.sqlite" 'SELECT did FROM account;' > "${WORK}/dids" +missing=0 +shopt -s nullglob +while IFS= read -r did; do + stores=("${STAGE}"/actors/*/"${did}"/store.sqlite) + keys=("${STAGE}"/actors/*/"${did}"/key) + if (( ${#stores[@]} == 0 || ${#keys[@]} == 0 )); then + log "ERROR: account ${did} has no staged store.sqlite or key" >&2 + missing=$((missing + 1)) + fi +done < "${WORK}/dids" +shopt -u nullglob +if (( missing > 0 )); then + log "ERROR: ${missing} account(s) incomplete; not writing an archive" >&2 + exit 1 +fi + +tar -C "${STAGE}" -czf "${ARCHIVE}.partial" . +tar -tzf "${ARCHIVE}.partial" > /dev/null +mv "${ARCHIVE}.partial" "${ARCHIVE}" +log "backup verified: ${ARCHIVE} (${databases} databases, ${files} other files, $(du -h "${ARCHIVE}" | cut -f1))" + +find "${BACKUP_DIR}" -maxdepth 1 -name 'pds-*.tar.gz' -mtime "+${RETENTION_DAYS}" -delete +# Matches the failed-run partial, which is never deleted. +find "${BACKUP_DIR}" -maxdepth 1 -name 'pds-*.tar.gz.partial' -mmin +1440 -print | while read -r stale; do + log "WARNING: stale partial from a failed run: ${stale}" >&2 +done +# A SIGKILL or power loss skips the EXIT trap (SIGTERM and SIGINT are +# trapped) and leaves plaintext database copies and signing keys behind. +find "${BACKUP_DIR}" -maxdepth 1 -type d -name '.pds-stage.*' -mmin +1440 -print | while read -r stale; do + log "WARNING: stale staging directory from a killed run (contains signing keys): ${stale}" >&2 +done + +log "backup done" diff --git a/scripts/pg-backup.sh b/scripts/pg-backup.sh new file mode 100755 index 0000000..bdb0a72 --- /dev/null +++ b/scripts/pg-backup.sh @@ -0,0 +1,69 @@ +#!/usr/bin/env bash +# pg-backup.sh — nightly logical backup of the production AppView Postgres. +# +# Run as root by coves-pg-backup.timer (scripts/systemd/) from the installed +# copy at /usr/local/sbin/coves-pg-backup, never from the /opt/coves checkout: +# deploys `git pull` that checkout, and an untracked or edited script there +# blocks the pull. docs/PRODUCTION_BACKUPS.md has the install steps. +# +# Custom format (-Fc) so a restore can be selective. The dump is streamed out +# of the container and written by the host under umask 077, so it is root-only +# from its first byte: it holds every hosted community's sealed PDS password +# and every aggregator's sealed OAuth session. +# +# NOT covered: ENCRYPTION_KEY and the rest of /opt/coves/.env. Without +# ENCRYPTION_KEY the credential columns in this dump are unreadable ciphertext. + +set -euo pipefail +umask 077 + +BACKUP_DIR="${BACKUP_DIR:-/opt/coves/backups}" +CONTAINER="${CONTAINER:-coves-prod-postgres}" +RETENTION_DAYS="${RETENTION_DAYS:-14}" + +STAMP="$(date -u +%Y%m%dT%H%M%SZ)" +DUMP="${BACKUP_DIR}/coves-${STAMP}.dump" + +log() { echo "[$(date -u +%FT%TZ)] $*"; } + +# A failed run keeps its partial under one fixed name, so the timer's retries +# overwrite it instead of piling up a full-size partial per attempt. The +# partial only exists between pg_dump starting and the final mv, so its +# presence here means this run failed. +FAILED_PARTIAL="${BACKUP_DIR}/coves-last-failed.dump.partial" +keep_failed_partial() { + local status=$? + if [[ -e "${DUMP}.partial" ]]; then + mv -f "${DUMP}.partial" "${FAILED_PARTIAL}" + log "ERROR: run failed; its partial is kept as ${FAILED_PARTIAL}" >&2 + fi + exit "${status}" +} +trap keep_failed_partial EXIT +# bash skips the EXIT trap on a signal it has no trap for; systemd stops and +# timeouts send SIGTERM, so turn it into an exit the EXIT trap sees. +trap 'exit 143' TERM +trap 'exit 130' INT + +log "backup starting: ${DUMP}" + +# The user and database names are not hardcoded in this script or its unit +# file: they come from the container's environment. +docker exec "${CONTAINER}" sh -c \ + 'pg_dump -Fc --no-owner --no-acl -U "$POSTGRES_USER" -d "$POSTGRES_DB"' \ + > "${DUMP}.partial" + +# Verify the archive BEFORE it gets the real name, with a full read: --list +# only reads the table of contents, so a truncated dump would pass it. +docker exec -i "${CONTAINER}" pg_restore -f /dev/null < "${DUMP}.partial" +mv "${DUMP}.partial" "${DUMP}" +log "backup verified: ${DUMP} ($(du -h "${DUMP}" | cut -f1))" + +# Retention ages out completed dumps only. The failed-run partial is never +# deleted, and is warned about once it is older than a day. +find "${BACKUP_DIR}" -maxdepth 1 -name 'coves-*.dump' -mtime "+${RETENTION_DAYS}" -delete +find "${BACKUP_DIR}" -maxdepth 1 -name 'coves-*.dump.partial' -mmin +1440 -print | while read -r stale; do + log "WARNING: stale partial from a failed run: ${stale}" >&2 +done + +log "backup done" diff --git a/scripts/setup-production.sh b/scripts/setup-production.sh index fde35c3..9d86c62 100755 --- a/scripts/setup-production.sh +++ b/scripts/setup-production.sh @@ -105,4 +105,4 @@ log "" log "Useful commands:" log " View logs: docker compose -f docker-compose.prod.yml logs -f" log " Deploy update: ./scripts/deploy.sh appview" -log " Backup DB: ./scripts/backup.sh" +log " Backups: docs/PRODUCTION_BACKUPS.md" diff --git a/scripts/systemd/coves-pds-backup.service b/scripts/systemd/coves-pds-backup.service new file mode 100644 index 0000000..4e80978 --- /dev/null +++ b/scripts/systemd/coves-pds-backup.service @@ -0,0 +1,14 @@ +[Unit] +Description=Coves nightly backup: PDS data directory +# A failed run retries every 5 minutes: a run missed while the host was down +# fires at boot, and after an unclean shutdown the PDS may not yet have +# recovered its WAL; other SQLite busy or I/O errors are transient too. Four +# failed starts in two hours stop the retries. +StartLimitIntervalSec=2h +StartLimitBurst=4 + +[Service] +Type=oneshot +ExecStart=/usr/local/sbin/coves-pds-backup +Restart=on-failure +RestartSec=5min diff --git a/scripts/systemd/coves-pds-backup.timer b/scripts/systemd/coves-pds-backup.timer new file mode 100644 index 0000000..8a91ef1 --- /dev/null +++ b/scripts/systemd/coves-pds-backup.timer @@ -0,0 +1,9 @@ +[Unit] +Description=Run coves-pds-backup.service nightly + +[Timer] +OnCalendar=*-*-* 03:07:00 UTC +Persistent=true + +[Install] +WantedBy=timers.target diff --git a/scripts/systemd/coves-pg-backup.service b/scripts/systemd/coves-pg-backup.service new file mode 100644 index 0000000..fb1e9f2 --- /dev/null +++ b/scripts/systemd/coves-pg-backup.service @@ -0,0 +1,15 @@ +[Unit] +Description=Coves nightly backup: AppView Postgres +Requires=docker.service +After=docker.service +# A failed run retries every 5 minutes (a run missed while the host was down +# fires at boot, possibly before the containers are up); four failed starts in +# two hours stop the retries. +StartLimitIntervalSec=2h +StartLimitBurst=4 + +[Service] +Type=oneshot +ExecStart=/usr/local/sbin/coves-pg-backup +Restart=on-failure +RestartSec=5min diff --git a/scripts/systemd/coves-pg-backup.timer b/scripts/systemd/coves-pg-backup.timer new file mode 100644 index 0000000..7a903e2 --- /dev/null +++ b/scripts/systemd/coves-pg-backup.timer @@ -0,0 +1,9 @@ +[Unit] +Description=Run coves-pg-backup.service nightly + +[Timer] +OnCalendar=*-*-* 02:47:00 UTC +Persistent=true + +[Install] +WantedBy=timers.target -- 2.51.2