# Scripts ## estimate-nl-users.py Approximates how many people in the Netherlands are on AT Protocol by sampling the public Jetstream firehose. See the docstring at the top of the file, and the methodology on `/community-size`. ```sh uv run scripts/estimate-nl-users.py [seconds] [out.json] ``` Output is JSON: total posts, Dutch posts and their share, distinct Dutch-posting accounts, resolved handle suffixes (.nl / .be / other), and the current Bluesky account total from bsky.jazco.dev/stats. Feed the numbers into `src/data/nl-estimate.ts`, rebuild, and redeploy. ### Why run it long Dutch activity peaks in the evening (UTC+2). A short window over- or under-counts depending on the time of day, so a representative figure needs a full day or more. The script auto-reconnects across the whole window and writes `.partial` every five minutes so you can watch progress and survive a mid-run drop. ### Running a full day on a server With uv installed: ```sh uv run scripts/estimate-nl-users.py 86400 ~/atproto-nl/nl-24h.json ``` Or in a container (no local Python needed): ```sh docker run --rm -v "$PWD:/w" -w /w python:3.12-slim sh -c \ "pip install -q websockets && python scripts/estimate-nl-users.py 86400 nl-24h.json" ``` Schedule a daily run with cron (adjust paths): ```cron # 03:00 daily: sample the next 24h, write a dated result 0 3 * * * cd ~/atproto-nl/website && uv run scripts/estimate-nl-users.py 86400 "nl-$(date +\%F).json" ``` Averaging several daily results smooths out day-to-day and time-of-day variation. ## nl-stats-aggregate.py Unions the feeder `.dids` files into the public `nl-stats.json` the `/stats` page reads. ```sh uv run scripts/nl-stats-aggregate.py [--dir DIR] [--out nl-stats.json] uv run scripts/nl-stats-aggregate.py --self-test ``` Each feeder writes a `.dids` file (one DID per line) into a shared private directory: `nl-live.json.dids` (Dutch posters), `nl-handles.dids` (`.nl` handles), `starterpack.dids`, `sifa-nl.dids`. The aggregator deduplicates them into a single union so overlap is not double-counted, folds in the activity estimate from the collector's full result, and writes counts only. The DID lists stay private; only `nl-stats.json` (no DIDs) is committed to `src/data/nl-stats.json`. `identified.total` is the union size (always at least the largest single source and at most the summed sources), so the build-time guard in `src/data/nl-stats.ts` passes. Idempotent; run it after the feeders refresh, then commit the JSON. Missing feeders count as 0. ## nl-sifa-location.py Pulls accounts that declare a Netherlands location in their Sifa ID profile into `sifa-nl.dids`, a feeder for the aggregator. Calls the public Sifa API (`GET /api/accounts/by-location?country=NL`, paginated), which already excludes deactivated and non-discoverable accounts. ```sh uv run scripts/nl-sifa-location.py [--base https://sifa.id] [--country NL] uv run scripts/nl-sifa-location.py --self-test ``` ## nl-stats-publish.sh The scheduled STATS job that runs ON the collector host (a thin host bootstrap fetches this file from Tangled raw and execs it). It regenerates the feeders + aggregators, refreshes the stats JSON (`nl-stats.json`, `europe-stats.json`, `europe-facts.json`, per-country facts; counts only, no DID lists), then in a container clones Tangled fresh, commits the refreshed stats, and pushes to Tangled `main`. That push re-triggers the Spindle CI (`.tangled/workflows/deploy.yml`), the SOLE publisher of every scope to Codeberg Pages. This job is NOT a deployer and never touches the abandoned Codeberg source mirror. As one of three writers to `main`, its push rebase-and-retries so a concurrent dev/PR commit cannot silently drop the stats deploy; a hard failure exits non-zero. The NL floor logic is untouched: `europe-stats.json` is produced independently (no `--emit-dids`/`--attributed` convergence here — that stays a separate, deliberate step). (The older `publish-nl-stats.sh`, a local-checkout script that pushed BOTH Tangled and the Codeberg source mirror, was removed 2026-09-22: writing that mirror is what caused the two-way divergence that froze deploys for ~11h. Tangled is the single source of truth; only `nl-stats-publish.sh` runs now.) The European aggregate now ends with two linked steps on the host: ```sh python3 eu-stats-aggregate.py --dir . --out europe-stats.json --emit-placed eu-placed.dids python3 eu-infra-pds.py --stats europe-stats.json --placed eu-placed.dids ``` `--emit-placed` writes the DIDs behind the published confirmed total; `--placed` lets the PDS counter report how many European-infrastructure accounts a country already claims (`europeanInfra.placedInCountry`) and how many it does not (`europeanInfra.countryless`). Only the second number can be added to the per-country figure, which is what lets /countries publish one European total instead of two numbers a reader might add themselves and double-count. `eu-placed.dids` stays on the host, and is deliberately not named `attributed.nl.dids`, which `nl-stats-aggregate.py` would auto-adopt and shift the Dutch floor. Without `--placed` the fields are absent and the page falls back to keeping the numbers separate. ```sh NL_STATS_DIR=/path/on/host DEPLOY=1 bash scripts/nl-stats-publish.sh # DEPLOY=0 = build-verify only ``` All host config is via env vars (no infra in the repo). The scheduler triggers it hourly; the full feeder+commit+push runs only on 6-hour slots (`FORCE=1` overrides). `DEPLOY=0` builds and verifies without committing or pushing. It pushes only to Tangled (`origin`), never the Codeberg source mirror. The push rebase-and-retries (3 attempts) so a dev/PR commit landing during the multi-minute feeder+scan window cannot reject it; after 3 failed attempts it exits non-zero so the cron wrapper can alert. ## nl-starterpacks.py Unions the members of curated Dutch starter packs into `starterpack.dids`, a feeder for the NL-stats aggregator. Someone curating a pack of Dutch accounts is a decent "this account is NL" signal. ```sh uv run scripts/nl-starterpacks.py [--packs FILE] [--out-prefix starterpack] uv run scripts/nl-starterpacks.py --self-test ``` The curated pack list lives in `scripts/nl-starterpacks.txt` (one `/` per line with a description, so it stays auditable). For each pack it resolves the member list via the public AppView (`getStarterPack` -> `getList`) and unions the member DIDs. Output is `starterpack.dids` (private) + `starterpack.json` (count + packs). Keep the list tight and clearly-NL: union membership just means "curated as NL", which is fine for the identified floor. Add packs from blueskystarterpack.com's `/dutch`, `/netherlands`, `/nederland`, `/dutch-artists` category pages. ## eu-verifier-lists.py Unions country-community verifier lists into per-country `verifier..dids` feeders for the European aggregator (`eu-stats-aggregate.py`). The Eurosky verification programme publishes one `app.bsky.graph.list` per verifier ("Verified by ", from the bot `@eurosky.bskycheck.com`); a country community's verifier vouching for an account is a deliberate, vetted signal that it belongs to that country's atproto community. ```sh uv run scripts/eu-verifier-lists.py [--lists FILE] [--out-prefix verifier] uv run scripts/eu-verifier-lists.py --self-test ``` The config lives in `scripts/eu-verifier-lists.txt` (one ` # desc` per line). Only geographic country-community verifiers (`belgium-atmosphe.re`, `france-atmosphe.re`, `spain-atmosphe.re`, ...) belong there — topic/org verifier lists (`philosophy.eurosky.social`, `nodeconf.eu`, `medsky.network`, ...) are not a country signal. For each list it resolves members via the public AppView (`getList`, paginated) and unions them per country. Output is `verifier..dids` (private) + `verifier.json` (per-country counts). Attribution treats it as rank 4 (`verifier`): a vetted community vouch — stronger than a starter pack, below a declared place. Discover new lists via `app.bsky.graph.getLists?actor=did:plc:tbz4faycgvx2frofy5xiy5fc`. ## nl-company-candidates.py + nl-company-triage.py Two steps that together turn "which of these NL accounts are companies/orgs?" into a list a human can mark off quickly. The result feeds Sifa's entity DB, which is what `nl-sifa-company.py` reads — so better classification here shows up as a more accurate Companies number on `/community-size`. ```sh # 1. collect + score candidates (search the network, and/or re-check DIDs we already have) uv run scripts/nl-company-candidates.py --search --out candidates.json uv run scripts/nl-company-candidates.py --dids nl-handles.dids --out candidates.json uv run scripts/nl-company-candidates.py --search --dids nl-handles.dids \ --exclude-dids already-classified.dids --out candidates.json # 2. build the offline triage page, then open it in a browser uv run scripts/nl-company-triage.py --in candidates.json --out triage.html ``` Two candidate sources: - `--search` queries the public AppView (`app.bsky.actor.searchActors`, no auth) for Dutch terms (`--terms`, defaults to nederland / dutch / holland / big-city names). This reaches accounts we do not have yet, so it can also grow the identified floor. - `--dids FILE` hydrates a feeder list we already have — `nl-handles.dids` is the interesting one. Adds no unique accounts by definition; it makes the Companies count more accurate. DIDs whose handle no longer resolves come back as `handle.invalid` and are dropped by default (`--keep-invalid-handles` to keep them). There is nothing to classify — no handle, usually no profile. In the first `nl-handles.dids` pass that was **169 of 1.781 (9,5%)**, which is also a data-quality signal: those DIDs still count toward the `nlHandle` source in the floor even though their Dutch domain is gone. `*.bsky.social` handles are dropped by default (`--keep-bsky-social` to keep them): an org that cared enough to set a domain handle is the population we're after, and it cuts the list by roughly an order of magnitude. Scoring is a transparent heuristic, not a verdict: legal forms (BV/NV/stichting), org language ("wij zijn", agency, hiring, role email), an apex-domain handle, minus person signals (pronouns, family words, "opinions are my own"). Every hit is listed on the card as a reason, and the list sorts by score so the obvious orgs come first. The triage page is one self-contained HTML file — no server, `file://` works. The call is a boolean: **is this one specific human?** If not — company, foundation, usergroup, meetup, museum, media outlet, gemeente — it is an org. That matches Sifa's entity DB, which distinguishes entities from personal accounts, not company from non-profit. Mark each candidate org / person / not-NL / unsure (keys `o` / `p` / `x` / `u`, `j`/`k` to move). "Not NL" is for candidates that should never have been in the list — `dutchtownstl.org` is a town in Missouri. Those are detection false positives, not classifications, and the export lists them separately as `notNl` so they can be fed back as an exclusion list (`nl-company-candidates.py --exclude-dids`) instead of polluting the NL floor. Two accelerators for long lists: a **score filter** in the header (5+, 3+, 1+, or 0 and below) so the obvious orgs can be cleared in one sweep, and an **Accept N person suggestions** button. A suggestion requires *positive* evidence of a person — at least one person signal (pronouns, family words, "opinions are my own") and no net org signal. A bare score of 0 means no signals were found at all, and absence of evidence is not evidence, so those are never suggested. Accepted suggestions are flagged `auto` on the card, filterable via "Auto-marked (review)", cleared the moment you touch the card, and exported with an `auto` column so a spot-check is possible after the fact. Decisions persist in `localStorage` as you go, and Export writes `nl-org-decisions.json` / `nl-org-decisions.csv` keyed by DID, each row carrying an `is_org` boolean (Import reloads an export, so a run survives a browser or machine change). DIDs are the identity key throughout; handles are re-resolved on every run and never stored as identity. ## nl-collector-day.sh One collector day: samples the firehose for ~23h, archives the result to `days/`, and leaves `nl-live.json` where the aggregator expects it. Cron it daily. ```cron 5 0 * * * /path/to/nl-collector-day.sh >> /path/to/collector.log 2>&1 ``` **On a host with no user `crontab` and no passwordless sudo** (as on some appliance systems), use `nl-collector-supervisor.sh` instead: it sleeps until the next `RUN_AT` (default 00:05), runs one day, and repeats. ```sh setsid nohup ./nl-collector-supervisor.sh >> supervisor.log 2>&1 < /dev/null & ``` It does not survive a reboot on its own — add a boot-time *triggered task* in the host's task scheduler running the same command as the same user. Running it alongside a real cron entry is harmless: the `flock` in `nl-collector-day.sh` still allows only one collector at a time. Replaces the old "one long capped run" pattern. A 30-day capped run has a cliff: when it ends the floor stops growing and the estimate goes stale with nothing saying so. Daily runs also produce exactly what the error bounds need — one Dutch-post-share sample per day — and a reboot costs one day instead of the whole series. `flock` keeps two runs from overlapping, which would double-count posts into one window and poison that day's sample. **Gaps are normal and stay visible.** A day the host is down has no sample and no history line. Nothing backfills it, nothing interpolates across it, and the accumulated DID master never shrinks: a missing day is missing evidence, not a drop to zero. Anything plotting this data must break the line at a gap rather than draw through it. ## History, error bounds and overlap (nl-stats-aggregate.py, v2) Three additions, all counts-only so they are safe to publish: - **`identified.exclusiveBySource`** — per source, the DIDs *only* that source found, plus `summedSources`. `summedSources - total` is the overlap. This is what lets the page show overlap as a bar segment instead of an unreadable 5-set Venn diagram. Currently the sources are near-disjoint: 9.735 summed against a 9.363 union, so 372 DIDs (3,8%) overlap. - **`nl-share-samples.jsonl`** — one line per completed collector day (`date`, `sharePct`, `windowMinutes`). A window shorter than 24h is not a sample of a day and is skipped; a re-run of the same date replaces that date's line. Malformed lines are ignored rather than fatal, because a cron job on another host appends to this file. - **`estimate.modelledLo` / `modelledHi`** — derived from the observed day-to-day spread of the share (mean ±1.96σ pushed through the same model), not guessed. Two hard rules: the low end is clamped to `identified.total`, because no model may claim fewer accounts than we can actually point at; and below `MIN_SHARE_SAMPLES` days no range is published at all. A missing range renders as "no published margin yet", which is honest; a fabricated one would not be. - **`nl-first-seen.jsonl`** — one append-only line per DID, ever: the date we first identified that account as Dutch, and which detector got there first. A DID is dated once and never re-dated. This is what keeps a future graph honest. A rising floor tangles two different claims together: more Dutch accounts exist, or we simply got better at spotting them. A cumulative count cannot separate them, and plotting it as growth would be wrong. With first-seen dates you can plot **arrivals per day** instead of a running total, and the crediting detector tells you which kind of arrival it was: a DID first seen by the firehose posted Dutch that day (closer to a real arrival), while one first seen by the PLC handle scan is almost always catch-up, an account that existed for years until we looked. Same +1 on the floor, opposite meanings. When several detectors see a DID in the same run the credit is a tie, broken by a fixed precedence (`FIRST_SEEN_PRECEDENCE`): the firehose first, because it is the only real-time signal; the standing-state scans after. The seeding run is a cliff by construction: every account already known was dated the day this started (2026-08-29), so treat that date as backlog, not arrivals. `nl-stats-history.jsonl` gets one append-only line per aggregator run: the counts, the estimate, a `coverage` block (collector window length and whether it completed, how many starter packs were in the list, which feeders were present, `scriptVersion`), and `newBySource` / `newTotal`: first-time identifications in that run rather than the running total. Without coverage a jump in the floor is uninterpretable — community growth, a finished PLC scan, and "we added starter packs" all look identical. The stats job pulls this file back alongside `nl-stats.json` and commits it to `src/data/nl-stats-history.jsonl`. ## nl-triage-ledger.py The running record of every triage decision, so no account is ever judged twice. ```sh # after a round, fold its export in uv run scripts/nl-triage-ledger.py add --export ~/Downloads/nl-org-decisions.json \ --ledger nl-triage-ledger.jsonl --round "nl-handles-2026-08-30" # regenerate the exclusion list, then sweep with it uv run scripts/nl-triage-ledger.py dids --ledger nl-triage-ledger.jsonl --out judged.dids uv run scripts/nl-company-candidates.py --search --exclude-dids judged.dids --out next.json uv run scripts/nl-triage-ledger.py stats --ledger nl-triage-ledger.jsonl ``` Decisions live in a browser's `localStorage` until exported, which is fine for one pass and useless across rounds: the next sweep re-offers every person already rejected. The ledger is the memory between rounds, and `judged.dids` is what `--exclude-dids` reads. `unsure` is deliberately NOT suppressed: it means "come back to this one", so those accounts are offered again. `org`, `person` and `notnl` are. Append-only, newest decision per DID wins. A revised call adds a line instead of rewriting one, so a mistaken bulk accept can be traced and reversed. Malformed lines are skipped rather than fatal. **The ledger is private data** (it lists individual accounts and how they were judged), so it lives beside the `.dids` files on the collector host, never in this repo. The canonical copy is `~/nl-estimate/nl-triage-ledger.jsonl`. ## nl-handle-class.py Shared handle classifier: `apex` / `subdomain` / `pds` / `bridged` / `invalid`. Used by the candidate sweep and by `nl-handle-revalidate.py`, because both need the same distinctions and got them wrong independently. The two uses differ, and the difference matters: - **For the NL floor**, a subdomain counts. `smb.phys.tue.nl` is as Dutch as `tue.nl`; the only question is "is this account Dutch". PDS-issued (`*.bsky.social`, `*.eurosky.social`) and bridged (`*.ap.brid.gy`) handles never count — that trailing domain belongs to infrastructure and says nothing about the holder. - **For binding an account to an organisation** (Sifa's entity DB), only `apex` is safe. Binding `nl.esa.int` would mint a wrong entity or claim the parent's domain. Sifa's own `isBindableHandleDomain` guard rejected **65 of the first 155 orgs** we handed over on exactly this: 22 bridged, 15 invalid, 10 PDS-issued, 18 non-apex subdomains. Candidates now carry `handleClass` and `bindable`, the triage page shows them on the card, and the export has an `orgsBindable` list plus `handle_class` / `bindable` CSV columns — so a hand-off contains only rows the receiving system can actually use. Apex detection uses a small suffix table, not the full Public Suffix List: it runs offline inside feeders and only has to be right for the TLDs Dutch accounts use. ## nl-handle-revalidate.py Re-checks that every DID in the `nlHandle` feeder still earns its place in the floor. ```sh uv run scripts/nl-handle-revalidate.py --dids nl-handles.dids --out nl-handles-clean.dids \ --details nl-handles-revalidate.json uv run scripts/nl-handle-revalidate.py --dids nl-handles.dids --report-only ``` `nl-handles-scan.py` records a DID the moment its handle ends in a Dutch domain, and nothing ever removes it. But that feeder is a **current-state snapshot** by design — the aggregator's own docstring says a lapsed `.nl` handle *should* drop out — so without a pruning pass the floor accumulates accounts that no longer meet the criterion that admitted them. First real run, 2.027 DIDs: **342 (16,9%) no longer qualify** — 178 whose handle no longer resolves, 109 the AppView will not return at all (deleted / deactivated / taken down), and 55 that resolve but have moved off their Dutch domain. Only 13 of the 342 are caught by another detector, so the union drops from 9.441 to 9.112 (−3,5%). Non-destructive: it writes a cleaned file, never the live feeder, unless `--out` points there. ## nl-did-facts.py + nl-lexicons.py + facts-aggregate.py History we can read today instead of accumulating forward. Every `did:plc` identity has a public audit log, and every repo lists the collections it holds. ```sh uv run scripts/nl-did-facts.py --dids nl-first-seen.jsonl --out nl-did-facts.jsonl uv run scripts/nl-lexicons.py --facts nl-did-facts.jsonl --out nl-lexicons.jsonl --stale-days 30 uv run scripts/facts-aggregate.py --dir . --out nl-facts.json ``` - **`nl-did-facts.py`** — one audit-log request per DID yields three things: `createdAt` (the first operation), every handle the DID has ever had with the date it took it, and the PDS it lives on now. Supersedes the creation-date-only version. - **`nl-lexicons.py`** — `com.atproto.repo.describeRepo` against each account's own PDS lists its collections, which is a direct answer to "which apps does this account use". It counts **accounts, not records**: a repo listing says a collection exists, not how much is in it. `--stale-days 30` re-reads a repo monthly, which is how new adoption surfaces. - **`facts-aggregate.py`** — folds both into `nl-facts.json` (counts only, no DIDs, no handles) for `src/data/nl-facts.json`. **Unreachable accounts.** `describeRepo` failures are not all the same, and treating them alike quietly undercounts. A `400`/`404` is the PDS *answering* that no repo is there (deactivated, taken down, deleted) and is recorded as `gone`. A rate limit or timeout is not an answer at all, so it is recorded as `error` and retried on the next run. Only after `MAX_FAILED_RUNS` separate runs with no answer is an account called `abandoned`, and one successful read clears the streak. Why it matters: the first full pass came back with 201 unreadable repos, and the failure rate climbed steadily as the run went on. That is the signature of rate limiting, not of 201 accounts dying in DID order. Recording those as settled would have removed them permanently. `--abandoned-out nl-abandoned.dids` writes the given-up-on DIDs, and `nl-stats-aggregate.py` subtracts them from every source and from the union: the floor is what we can point at today, and an account that is not there cannot be pointed at. The DIDs stay listed in `nl-lexicons.jsonl` precisely so the weekly run stops asking. Both collectors are resumable and append-only, so an interrupted run costs nothing and a weekly re-run only fetches what is new. `nl-collector-day.sh` runs them on Sundays; creation dates and PDS moves are rare and app adoption is slow, so daily would be wasted requests against other people's infrastructure. Two things the aggregate deliberately does: - Bluesky's hosted fleet (`*.host.bsky.network`) collapses to one operator. Counting each shard separately would say "spread across 40 hosts", which is false. - A lexicon used by fewer than three accounts is dropped. Naming a one-account lexicon invites over-reading a coincidence. **What this population is:** every account OUR detectors identified as Dutch. A creation curve from it shows when the Dutch accounts we can see joined, not when Dutch people joined atproto, and detection favours the visible. The page says so where it renders it.