diff --git a/AGENTS.md b/AGENTS.md index 5632729..91765b8 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -103,8 +103,8 @@ to its own Codeberg Pages repo (`atproto-website`, `pages` branch), which Co at the matching domain. No manual step — push `main` and it deploys. **Push reaches two remotes (both required).** `origin` fans out to Tangled (home base, where the -Spindle listens) AND the Codeberg source mirror `atprotonl-website` (source of truth + feeds the NAS -job). A single `git push origin main` hits both. Harden a fresh clone once: +Spindle listens) AND the Codeberg source mirror `atprotonl-website` (source of truth + feeds the +scheduled stats job). A single `git push origin main` hits both. Harden a fresh clone once: ```sh git remote set-url --add --push origin tangled.org:did:plc:aqrkn66k2alnwxfrhstp3unq/website @@ -127,7 +127,8 @@ Fetch stays on Tangled; `git remote -v` should show two `origin` push URLs. With - Verify what actually deployed: each Codeberg `pages` branch's latest commit is authored `atproto.nl deploy`, message `deploy `. Check: `curl -s "https://codeberg.org/api/v1/repos/gxjansen/atproto-website/commits?sha=pages&limit=1"`. - Author `nl-stats deploy` instead = the NAS deployed it, not the Spindle. + Author `nl-stats deploy` instead = an old direct deploy from the scheduled stats job (it no longer + publishes pages; only the Spindle does). - Failure shape: a fast (~10s) failure = an early step (install/test/copy-voice); a slower failure = a build/push inside the deploy loop. Two already-fixed early-step bugs (see git log): the shallow clone missing `HEAD~1` for copy-voice (now `git fetch --deepen 1`, skip if absent) and a wrong @@ -144,10 +145,11 @@ at the target commit, `for s in nl be eu no it ch hu; do SITE_SCOPE=$s bash scri content-diff almost never skips — so we don't bother diffing, just publish all seven (~2-min run vs the hosted spindle's 15-min cap). Builds ARE deterministic (`src/lib/seeded-shuffle.ts` seeds the list shuffles) so the display order stays stable between deploys. -- A **NAS920 6h job** (`~/nl-estimate/nl-stats-publish.sh`, git author `nl-stats deploy`) still - refreshes the stats JSON AND redeploys all scopes — redundant with the Spindle now, slated to be - trimmed to stats-only. Until then keep pushing `main` to both remotes (a stale-clone NAS run could - otherwise clobber). Realign the mirror to home base if they diverge: +- A **scheduled 6h stats job** (`scripts/nl-stats-publish.sh`, run on a separate host, git author + `nl-stats deploy`) refreshes the stats JSON, commits it, and pushes `main` to both remotes; the + Tangled push retriggers the Spindle, which republishes with the fresh stats. It no longer deploys + pages itself (removed so the Spindle is the sole publisher). Keep pushing `main` to both remotes so + the two never diverge; realign the mirror to home base if they do: `git push --force codeberg :main`. - Custom domain per scope needs the A/AAAA + `_git-pages-repository` TXT **and** the Codeberg webhook (the webhook is the piece people forget; without it Pages/TLS never updates). diff --git a/scripts/README.md b/scripts/README.md index a7d4e57..41e2334 100644 --- a/scripts/README.md +++ b/scripts/README.md @@ -236,16 +236,16 @@ leaves `nl-live.json` where the aggregator expects it. Cron it daily. 5 0 * * * /path/to/nl-collector-day.sh >> /path/to/collector.log 2>&1 ``` -**On a Synology NAS there is no user `crontab` and no passwordless sudo**, so use -`nl-collector-supervisor.sh` instead: it sleeps until the next `RUN_AT` (default 00:05), -runs one day, and repeats. +**On a host with no user `crontab` and no passwordless sudo** (as on some appliance +systems), use `nl-collector-supervisor.sh` instead: it sleeps until the next `RUN_AT` +(default 00:05), runs one day, and repeats. ```sh setsid nohup ./nl-collector-supervisor.sh >> supervisor.log 2>&1 < /dev/null & ``` -It does not survive a reboot on its own — add a DSM Task Scheduler *triggered task* on -boot-up running the same command as the same user. Running it alongside a real cron entry +It does not survive a reboot on its own — add a boot-time *triggered task* in the host's +task scheduler running the same command as the same user. Running it alongside a real cron entry is harmless: the `flock` in `nl-collector-day.sh` still allows only one collector at a time. Replaces the old "one long capped run" pattern. A 30-day capped run has a cliff: when it diff --git a/scripts/collect-countries.sh b/scripts/collect-countries.sh index 170b8fa..4977f82 100755 --- a/scripts/collect-countries.sh +++ b/scripts/collect-countries.sh @@ -12,7 +12,7 @@ # Usage: # bash scripts/collect-countries.sh # every country in the map # COUNTRIES="nl be de" bash scripts/collect-countries.sh -# DIR=/volume1/homes/sshuser/nl-estimate bash scripts/collect-countries.sh +# DIR=/path/to/nl-estimate bash scripts/collect-countries.sh # # Output per country, in $DIR: # sifa-.dids accounts with a Sifa location in that country diff --git a/scripts/estimate-nl-users.py b/scripts/estimate-nl-users.py index 1a6ae8f..9213ead 100644 --- a/scripts/estimate-nl-users.py +++ b/scripts/estimate-nl-users.py @@ -41,7 +41,7 @@ the full (with handle resolution) every RESOLVE_EVERY seconds, so a live t read it. Every checkpoint it also writes: .dids distinct Dutch-poster DIDs (unchanged; the NL aggregator's feed) ..dids distinct DIDs per tracked language -Keep the .dids files private (NAS only). Prefer a capped duration (e.g. 30 days) over +Keep the .dids files private (host only, never committed). Prefer a capped duration (e.g. 30 days) over 0/forever so it can't run unbounded. Run it under cron/systemd or in a container; see scripts/README.md. """ @@ -76,7 +76,7 @@ RESOLVE_EVERY = 3600 # in forever mode, resolve handles + write the full res def parse_args(argv): """Positional [seconds] [out.json], plus optional --langs / --author-langs. Positional - order is kept for compatibility with the deployed NAS job and the docs.""" + order is kept for compatibility with the deployed scheduled job and the docs.""" langs = DEFAULT_LANGS author_langs = None rest = [] @@ -218,8 +218,8 @@ def write_dids(authors, path): These are WINDOWED sets: DIDs seen posting a language during this run. The aggregator unions them across runs into a master file, so the "ever seen" floor grows and never forgets. The windowed count stays separate in the JSON, so "distinct this window" is - never conflated with "ever seen". Keep the .dids files private (NAS only); only the - aggregated counts go public. + never conflated with "ever seen". Keep the .dids files private (host only, never + committed); only the aggregated counts go public. """ if not path: return diff --git a/scripts/eu-stats-aggregate.py b/scripts/eu-stats-aggregate.py index 3d4132f..1934179 100644 --- a/scripts/eu-stats-aggregate.py +++ b/scripts/eu-stats-aggregate.py @@ -21,7 +21,7 @@ Inputs, per country cc, in --dir: verifier..dids country-community verifier list members -> verifier starterpack.dids curated pack members (NL only today) -> starterpack -Output is COUNTS ONLY. DID lists never leave the NAS. +Output is COUNTS ONLY. DID lists never leave the host that generates them. Usage: python3 scripts/eu-stats-aggregate.py --dir . --out europe-stats.json diff --git a/scripts/nl-collector-supervisor.sh b/scripts/nl-collector-supervisor.sh index 8b15ede..4a3273c 100755 --- a/scripts/nl-collector-supervisor.sh +++ b/scripts/nl-collector-supervisor.sh @@ -1,17 +1,17 @@ #!/usr/bin/env bash # Run one collector day, every day, starting at RUN_AT local time. # -# Why this exists: Synology DSM has no user `crontab` and no passwordless sudo, so the -# normal "cron it daily" instruction in nl-collector-day.sh cannot be followed on the -# NAS. This is the same cadence implemented in userspace: sleep until the next RUN_AT, +# Why this exists: some hosts have no user `crontab` and no passwordless sudo, so the +# normal "cron it daily" instruction in nl-collector-day.sh cannot be followed there. This +# is the same cadence implemented in userspace: sleep until the next RUN_AT, # run one day, repeat. `flock` inside nl-collector-day.sh still guarantees one collector -# at a time, so this is safe to run alongside a DSM task if one is ever added. +# at a time, so this is safe to run alongside a scheduled task if one is ever added. # # Start it detached: # setsid nohup ./nl-collector-supervisor.sh >> supervisor.log 2>&1 < /dev/null & # -# It does NOT survive a reboot on its own. For that, add a DSM Task Scheduler -# "triggered task" on boot-up that runs this script as the same user. +# It does NOT survive a reboot on its own. For that, add a boot-time "triggered task" +# in the host's task scheduler that runs this script as the same user. # # Config via env: # RUN_AT local HH:MM to start each day's run (default 00:05) diff --git a/scripts/nl-stats-aggregate.py b/scripts/nl-stats-aggregate.py index 9783835..3226602 100644 --- a/scripts/nl-stats-aggregate.py +++ b/scripts/nl-stats-aggregate.py @@ -15,8 +15,8 @@ on infra we control: This unions those DIDs so overlap is deduplicated, not double-counted: `identified.total` is the size of the union (the headline floor), and `bySource` holds the per-feeder counts (non-exclusive -- one DID can be in several). It also folds in the activity estimate from -the collector's full result. Output is COUNTS ONLY: the DID lists stay private on the NAS, -and only `nl-stats.json` (which carries no DIDs) is committed to the public repo. +the collector's full result. Output is COUNTS ONLY: the DID lists stay private on the host +that generates them, and only `nl-stats.json` (which carries no DIDs) is committed to the public repo. The output matches src/data/nl-stats.ts: total is always >= the largest single source and <= the summed sources (both hold for any real union), so the site's build-time guard passes. @@ -27,7 +27,7 @@ Usage: [--starterpack P] [--sifa-nl P] uv run scripts/nl-stats-aggregate.py --self-test -Run it on the NAS (where the .dids live), then commit the resulting nl-stats.json to +Run it where the .dids live, then commit the resulting nl-stats.json to src/data/nl-stats.json. Idempotent; safe to re-run. Missing feeder files count as 0. """ import argparse diff --git a/scripts/nl-stats-publish.sh b/scripts/nl-stats-publish.sh index e64bc8e..fdf441f 100644 --- a/scripts/nl-stats-publish.sh +++ b/scripts/nl-stats-publish.sh @@ -24,8 +24,8 @@ DEPLOY="${DEPLOY:-1}" DOCKER="${DOCKER_BIN:-docker}" REPO_URL="https://gxjansen:$(cat "$DIR/.codeberg-token")@codeberg.org/gxjansen/atprotonl-website.git" -# The full job (feeders + aggregate + deploy of all seven sites) runs only on 6-hour slots -# (00/06/12/18 local). The DSM trigger stays hourly, so 5 of every 6 invocations exit here. The +# The full job (feeders + aggregate + stats commit/push) runs only on 6-hour slots +# (00/06/12/18 local). The scheduler trigger stays hourly, so 5 of every 6 invocations exit here. The # stats move slowly, so this keeps every site fresh enough at a quarter of the build load. FORCE=1 # overrides; a manual DEPLOY=0 test always runs (only the scheduled DEPLOY=1 job is gated). if [ "$DEPLOY" = "1" ] && [ -z "${FORCE:-}" ] && [ "$(( 10#$(date +%H) % 6 ))" -ne 0 ]; then @@ -202,7 +202,7 @@ $DOCKER run --rm -e REPO_URL="$REPO_URL" -e BRANCH="$BRANCH" -e DEPLOY="$DEPLOY" # the Spindle and could clobber it from a stale clone. Stats still reach the live sites because # the stats commit + push retriggers the Spindle, which rebuilds with the fresh JSON. (This also # retires the old "only .nl auto-deploys" trap: the per-scope loop needed a token with write to - # all seven Pages repos, but the NAS .codeberg-token only has atprotonl-website, so .be/.eu/.no + # all seven Pages repos, but the .codeberg-token used here only has atprotonl-website, so .be/.eu/.no # silently returned 403. Moot now -- the Spindle deploys with its own all-seven secret; the token is # still used, but only to clone + push the source mirror above, which it can do.) else diff --git a/src/data/europe-stats.ts b/src/data/europe-stats.ts index 56572b4..9c54ed4 100644 --- a/src/data/europe-stats.ts +++ b/src/data/europe-stats.ts @@ -1,7 +1,7 @@ // Per-country AT Protocol figures for atproto.eu — the aggregate the whole domain exists for. // // Produced by scripts/eu-stats-aggregate.py on infra we control and committed here as counts -// only; DID lists never leave the NAS. Every account is attributed to exactly ONE country by +// only; DID lists never leave the host. Every account is attributed to exactly ONE country by // the precedence in scripts/attribution.py, so the per-country numbers are people, not rows, // and they sum to the total instead of exceeding it. // diff --git a/src/data/nl-stats.ts b/src/data/nl-stats.ts index 59434fd..209cf95 100644 --- a/src/data/nl-stats.ts +++ b/src/data/nl-stats.ts @@ -5,7 +5,7 @@ // - estimate: the modelled activity figure, mirrors nl-estimate.ts. Clearly a model. // // The aggregate is produced off-site (on infra we control) and committed here as JSON, same -// model as nl-estimate.ts. The NAS never appears publicly and DID lists never leave it; only +// model as nl-estimate.ts. The host never appears publicly and DID lists never leave it; only // these counts do. This module just types + light-checks the committed JSON at build time. import raw from './nl-stats.json';