Large file storage for AT Protocol: an IPFS server & gateway with XRPC upload & admin. atfs.dev
ipfs atproto xrpc
Go 78%
Svelte 7%
TypeScript 7%
Shell 4%
CSS 2%
JavaScript <1%
Nix <1%
Makefile <1%
HTML <1%
Dockerfile <1%
<1%

README.md

atfs #

A tiny, CGO-free Go server that lets a preconfigured set of atproto accounts upload files to it — over dev.atfs.repo.uploadFile, an XRPC call also served under the same com.atproto.repo.uploadBlob NSID every atproto client already speaks, wire-identical either way — and serves that content back out over IPFS (a full Bitswap + DHT participant, not a gateway bolted onto someone else's node) and over plain HTTP. Extremely simple setup is the point: local configuration is a couple of environment variables, and everything else lives in an atproto record you create once, on your own PDS.

How it works #

  • Blessed CIDs. Uploaded blobs are addressed with atproto's own "blessed" CID format — CIDv1, raw multicodec (0x55), sha-256 multihash, base32 multibase — exactly what a PDS's own uploadBlob would hand back. Because it's a raw block, it's directly fetchable over IPFS with no translation step: the CID your upload gets back is its IPFS CID.
  • Dual-CID for large blobs. Bitswap can't move a block bigger than roughly 1–2 MiB, so anything larger than that also gets a UnixFS DAG built for it at upload time — a second, dag-pb root CID over the same bytes, fetchable block-by-block by ordinary IPFS peers. The blessed raw CID stays canonical: it's what the XRPC APIs return, and atfs's own HTTP paths serve it at any size regardless (the block-size limit only constrains Bitswap transport, not a local read). The UnixFS CID exists purely so large blobs can cross the IPFS network at all.
  • Configuration in two places. Local config — who owns the instance, how it's reached — is a handful of environment variables (ATFS_*). Everything else — the complete upload allowlist, the service DID uploaders authenticate against (which doubles as this instance's own address), which other instances to mirror — lives in an atproto record, dev.atfs.server, at at://{owner-did}/dev.atfs.server/{peer-id}. atfs reads it at boot and then watches the owner's repo, so later edits apply within seconds without a restart. A missing record isn't fatal: the instance comes up anyway, logs the exact at:// URI it went looking for, and runs with defaults — uploads disabled — until you create one.
  • Serves only what it has been asked to. An instance is never a general cache or an open relay: everything it serves has been claimed by name — uploaded to it by an allowlisted account, pinned by one, or mirrored from an instance the owner has explicitly listed in follows. Nothing arrives because it happened to pass through.

Quickstart #

# atcr.io serves pulls to authenticated atproto identities only — log in
# once with any Bluesky handle + an app password
# (bsky.app/settings/app-passwords):
docker login atcr.io -u your.handle.example.com

docker run -d \
  -e ATFS_OWNER_DID=did:plc:yourowndid \
  -v atfs-data:/data \
  -p 2837:80 -p 4001:4001/tcp -p 4001:4001/udp \
  atcr.io/atfs.dev/atfs

ATFS_OWNER_DID and a volume on /data are the only two things a first run needs. On boot, before any dev.atfs.server record exists, the log tells you exactly what to do next: it prints this instance's peer ID and the at:// URI a record needs to exist at, then carries on serving with uploads disabled. GET / returns a one-line identity string, handy for confirming the instance is actually up.

To enable uploads, create that record on the owner's own PDS at the rkey the log gave you, with a bare-domain did:web serviceDid and at least your own DID listed in accounts — the owner isn't implicitly included, so this step is what actually grants yourself upload rights (see Configuration below). No restart needed: the instance subscribes to the owner's repo, so the record appearing (or changing later) is picked up within seconds and logged.

The container runs as a non-root user #

The image runs as uid/gid 65532:65532 and serves plain HTTP on port 80.

Port 80 needs no extra privilege, because Docker sets net.ipv4.ip_unprivileged_port_start=0 inside every container: the usual "below 1024 is root-only" rule doesn't apply there, so no capability is granted and none is needed. On a runtime that keeps the restriction — some hardened Kubernetes setups — set ATFS_HTTP_PORT above 1024 and remap it on the outside.

/data has to be writable by uid 65532, or atfsd exits immediately with mkdir data/blobs: permission denied. A named volume (-v atfs-data:/data, as in the quickstart) is seeded from the image and inherits that ownership for free. A bind mount keeps the host directory's own ownership, so chown it first:

mkdir -p ./atfs-data && sudo chown 65532:65532 ./atfs-data
docker run -d -e ATFS_OWNER_DID=did:plc:yourowndid \
  -v "$PWD/atfs-data:/data" \
  -p 2837:80 -p 4001:4001/tcp -p 4001:4001/udp \
  atcr.io/atfs.dev/atfs

Configuration #

Local: ATFS_* environment variables #

ATFS_OWNER_DID is the only one a first run needs — see the Quickstart above. For the full list, every default, and how to set them on Docker or an SD-card image, see atfs.dev/docs/reference/settings/ and atfs.dev/docs/configuration/required-settings/.

The dev.atfs.server record #

accounts is required; every other field is optional, and unknown fields are ignored, so a running instance is never broken by a field it predates.

Field Default Purpose
accounts (required) The complete upload allowlist: atproto accounts (DIDs), permitted to upload. The owner is not implicitly included — list it explicitly if it should be able to upload to its own instance. DIDs rather than handles, so an entry survives a handle change and can be reverse-indexed to find which instances a person can use.
serviceDid (empty — uploads disabled) This instance's own identity and address: the aud uploaders must address in their inter-service auth JWTs, and — when it's a bare-domain did:web (see "did:web self-resolution" below) — also this instance's single HTTPS base URL. Uploads stay off until this is a bare-domain did:web and accounts names at least one account.
follows [] Other atfs instances to mirror, each named by the at-uri of that instance's own dev.atfs.server record. See "Following another instance".

Edits to this record apply live. The instance subscribes to the owner's PDS repo stream and re-reads the record whenever it changes — and again on every reconnect, so an edit made while the instance was offline isn't missed — logging one line saying what changed. accounts and serviceDid take effect on the very next request; follows reaches the running sync loop immediately. One thing still needs a restart, and says so in the log when it changes: the /.well-known/did.json document atfs serves for a bare-domain did:web serviceDid (built once at startup — see "did:web self-resolution" below). An instance that can't reach the owner's PDS just keeps running on the configuration it read at boot.

Removing an account from accounts is a revocation, not just a closed door. Everything that account uploaded or pinned here is released, and any of it nothing else claims is deleted — content, IPFS artifacts, DHT announcements and all — because an appliance that publishes content to a global DHT from the operator's own address has to be able to stop. Content another account (or a followed instance's mirror) still claims stays, minus the delisted account's claim. Three deliberate limits: emptying accounts altogether turns uploads off and revokes nothing (a record that was deleted, or that a PDS lost, looks exactly the same from here, and that mistake would be unrecoverable — so revoke by removing an account while at least one other stays listed); an account whose identity can't be resolved at that moment keeps its content until the next edit; and only an edit the running instance actually sees revokes, so if it was offline for the removal, re-add and re-remove the account.

Uploads are authenticated with standard atproto inter-service auth JWTs — signed by the uploading account's own signing key, verified against its DID document — exactly like talking to a PDS. There's no atfs-specific credential.

The uploadFile/uploadBlob response is the PDS's exact shape plus one extra top-level field, ipfsRoot: the CID to fetch over the IPFS network (the UnixFS root for a large blob, the blessed CID itself for a small one). Standard atproto clients ignore it; clients writing dev.atfs.file records want it.

Trust model #

The owner's PDS is a full, unverified authority over every instance configured from it. The record is read over a plain com.atproto.repo.getRecord call — no signed commit, no MST proof checked — so it's exactly what the PDS chooses to hand back, and (per "Edits to this record apply live" above) an edit takes effect within seconds with no restart and no human in the loop. Whoever controls that PDS therefore controls, for every instance pointed at it: the upload allowlist (accounts), the audience uploaders authenticate against (serviceDid), and the set of other instances mirrored (follows).

Concretely, a compromised PDS — or one that's mis-issued a certificate, or is simply operated by someone untrustworthy — can add itself to accounts and gain upload/pin rights, repoint serviceDid so the instance starts accepting tokens minted for a different audience, or add a follows entry and have the instance fetch, store, serve, and DHT-announce content of its choosing.

None of that is a bug in atfs: it's atproto's ordinary "you trust your PDS" model, applied to config-in-a-record on purpose (see "Configuration in two places" above). It's worth stating plainly anyway, because the stakes are higher than they are for a PDS's usual job — this record grants write access to an appliance's disk and its public serving surface, not just read access to some posts. A compromised owner PDS is equivalent to a compromised instance. Picking which PDS an atfs owner account lives on is picking who gets to hold that.

CLI #

atfs (cmd/atfs, built alongside the server — see "Building from source") is a small client for the three calls above: it does the login → mint-a-scoped-token → call dance itself, instead of four hand-rolled curls.

Getting it #

The container image doesn't carry the CLI — it ships atfsd alone — so there are two routes.

This repo is a nix flake whose only package is the CLI, which makes it installable and runnable directly:

nix run 'git+https://tangled.org/byjp.me/atfs?ref=refs/tags/vX.Y.Z#atfs' -- --version
nix profile install 'git+https://tangled.org/byjp.me/atfs?ref=refs/tags/vX.Y.Z#atfs'

Substitute the newest release tag for vX.Y.Z. The flake landed after v0.2.2, so pin something newer than that — an older tag has no flake.nix to find.

Or build it from a checkout with make build (see "Building from source" below), which is what a contributor wants.

In a Tangled workflow #

Tangled's CI resolves a workflow's dependencies: with nix and takes flake refs, so naming this repo puts atfs on PATH. The syntax differs by engine. On nixery, dependencies is a map keyed by flake ref, whose values are the attributes to take from it:

engine: "nixery"

dependencies:
  nixpkgs:
    - git
  "git+https://tangled.org/byjp.me/atfs?ref=refs/tags/vX.Y.Z":
    - atfs

On microvm it's a flat list, and the attribute rides on the ref itself:

engine: "microvm"
image: nixos

dependencies:
  - git
  - "git+https://tangled.org/byjp.me/atfs?ref=refs/tags/vX.Y.Z#atfs"

environment:
  ATFS_DOMAIN: myatfs.example.com
  ATFS_USER: you.example.com
  # ATFS_APP_PASSWORD comes from a repo secret — never a flag, never inline.

steps:
  - name: "upload a release artifact"
    command: "atfs upload --tag myproject ./dist/app.tar.gz > app.atfs.json"

Spell the pin refs/tags/, not just the tag, for two independent reasons that happen to give the same instruction. Nix expands a bare ?ref= to refs/heads/, so ?ref=vX.Y.Z hunts for a branch of that name and fails with couldn't find remote ref. And the runner's nix cache keys on the literal ref string, so anything unpinned is resolved once and then served stale forever — that's not hypothetical, it's what fed this repo's own image builds a months-old gosd (see .tangled/workflows/image.yml). Bumping the version should be a deliberate edit to that line.

atfs --version reports what a given pin actually resolved to. For a worked example of the CLI doing this job in anger, see .tangled/workflows/publish-sbc-images.yml, which uploads every release's images with it.

Every subcommand takes --user (your handle or DID) plus enough to address the instance. The easy way is --domain (its bare domain, e.g. myatfs.example.com, also settable as ATFS_DOMAIN): for any instance you can actually upload to, serviceDid is always a bare-domain did:web and the instance's own address is always https://<that domain> (uploads stay off otherwise — see "Configuration" above), so one domain fills both --server and --aud for you.

Explicit --server (the instance's base URL) and --aud (its service DID — the dev.atfs.server record's serviceDid), also settable as ATFS_SERVER/ATFS_AUD, still win individually over --domain/ATFS_DOMAIN when set. That pair is the override for reaching an instance at an address that isn't its canonical domain — over a LAN, through a port-forward, at a staging name — while --aud keeps naming the instance's real identity, since the aud in a minted token has to match who you mean to address, not the route you took to get there. The CLI never derives --aud by asking the instance itself: a host that answered with somebody else's did would get a token minted for that victim and handed straight to it, so the domain has to come from something the caller already trusts.

Your app password comes from ATFS_APP_PASSWORD only — never a flag, since argv is visible to every other process on the machine — and is never printed or logged.

export ATFS_DOMAIN=myatfs.example.com
export ATFS_USER=you.example.com
export ATFS_APP_PASSWORD=xxxx-xxxx-xxxx-xxxx

Reaching that same instance at an address that isn't its canonical domain:

export ATFS_SERVER=http://192.168.1.201:2837
export ATFS_AUD=did:web:myatfs.example.com   # still the instance's real serviceDid
export ATFS_USER=you.example.com
export ATFS_APP_PASSWORD=xxxx-xxxx-xxxx-xxxx

atfs upload FILE uploads a file and prints a ready-to-embed dev.atfs.file JSON object; repeat --tag to label it (see "Tagging" below):

$ atfs upload --tag board-x --tag v1.2.3 --tag atfs ./photo.jpg > photo.atfs.json
$ cat photo.atfs.json
{
  "$type": "dev.atfs.file",
  "cid": { "$link": "bafkrei..." },
  "ipfsRoot": { "$link": "bafkrei..." },
  "size": 483821,
  "mimeType": "image/jpeg",
  "providers": ["https://myatfs.example.com"],
  "tags": ["atfs", "board-x", "v1.2.3"]
}

atfs pin FILE|- reads a dev.atfs.file reference — a file, or - for stdin, so it composes with upload's own output — and asks the instance to fetch and serve it, printing where the pin stands:

$ atfs pin photo.atfs.json
{
  "state": "seeking"
}

Call it again with the same reference to check progress; it settles at "pinned" (or "failed" — see "Pinning content from elsewhere" below).

atfs delete CID releases your claim on a blob; the bytes themselves go too, once every claimant has done the same:

$ atfs delete bafkrei...
bafkrei...: content deleted (last claim released)

atfs init [HANDLE|DID] is for setting up a brand new instance, not for talking to one that already exists — it takes none of the common flags above, and for the env-var half below never touches an atfs instance at all, just ordinary atproto identity resolution (a DID still round-trips through it, so a typo is caught here rather than at the new instance's first boot). It mints a fresh libp2p identity key and prints the two ATFS_* environment variables the instance needs (see "Configuration" above — there are only ever these two):

$ atfs init jp.example.com
# atfs settings for did:plc:... (resolved from "jp.example.com") — see README's "Configuration"
ATFS_OWNER_DID=did:plc:...
ATFS_IDENTITY_KEY=CAESQ...

Run at a real terminal (not piped, not scripted), it goes further, prompting through a small charmbracelet/huh form: for a handle if none was given, then — after printing the env vars — for the domain the server will be reachable at and any other accounts to allow, with a checklist to confirm the final list (you're on it by default; uncheck yourself if you don't want to be), before offering to write the resulting dev.atfs.server record straight to your account, given your password or an app password. Press enter instead of a password to print the record rather than write it; a wrong password does the same rather than failing outright. Roughly (the real thing is a bordered, styled form, not plain lines — this is just the conversation it asks):

$ atfs init
Your server needs an owning account, and a peer ID. To generate them, please enter your atproto handle:
> jp.example.com                                   eg. you.eurosky.social

# atfs settings for did:plc:... (resolved from "jp.example.com") — see README's "Configuration"
ATFS_OWNER_DID=did:plc:...
ATFS_IDENTITY_KEY=CAESQ...

What domain will this server be reachable at?
Leave blank if you don't know yet
> myatfs.example.com                            eg. myatfs.example.com

Which other atproto accounts should be able to upload to this server?
A comma-separated list of handles or DIDs — leave blank if there are none
>                                     eg. friend.bsky.social, did:plc:...

Confirm which accounts can upload to this server
Everyone just entered starts checked — including you — uncheck any you don't want
> [x] jp.example.com
  [x] friend.bsky.social

Your server requires a record in your atproto account. To create it, please enter your account password (or an app password).
>                            Just press enter to print the record instead
Wrote at://did:plc:.../dev.atfs.server/12D3KooW...

Piped or scripted — atfs init jp.example.com > server.env in a script or CI — it's exactly the old behaviour: a handle or DID argument is required, and only the env vars are printed. Nothing here is a breaking change to that.

Worked examples #

The CLI above is the easy way to drive any of this — it does the login → mint-a-token → call dance for you. What follows is the tooling-free equivalent: goat (for identity and token minting) plus plain curl, so it's clear exactly what's on the wire for each of atfs's five XRPC calls.

Two traps before copying any of this:

  • A token's lxm claim must name the NSID you actually call. Every authenticated call below is an atproto inter-service auth JWT (see "Configuration" above), and atfs checks its lxm claim against the path the request hit — a token minted with --lxm dev.atfs.repo.uploadFile is rejected on /xrpc/com.atproto.repo.uploadBlob, even though the two paths do exactly the same thing (see "The alias trade-off" below).
  • Every token is single-use. atfs tracks each token's jti claim and rejects a repeat within that token's own validity window (auth: token already used (jti replay)) — so retrying a call, or polling pinFile for status, means minting a fresh token each time, never reusing the last one.

Set up once per shell session:

export ATFS_SERVER=https://myatfs.example.com
export ATFS_AUD=did:web:myatfs.example.com   # the instance's serviceDid
goat account login -u you.example.com -p "$ATFS_APP_PASSWORD"

goat account login persists a session that every goat account service-auth call below reuses — it's the identity you're minting tokens as, unrelated to $ATFS_SERVER.

uploadFile #

TOKEN=$(goat account service-auth --lxm dev.atfs.repo.uploadFile --aud "$ATFS_AUD")
curl -sS -X POST "$ATFS_SERVER/xrpc/dev.atfs.repo.uploadFile?tag=board-x&tag=v1.2.3&tag=atfs" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: image/jpeg" \
  --data-binary @photo.jpg
{"blob":{"$type":"blob","ref":{"$link":"bafkrei..."},"mimeType":"image/jpeg","size":483821},"ipfsRoot":{"$link":"bafkrei..."}}

goat's default token lifetime is 60 seconds (--duration-sec); bump it for a large, slow upload — atfs's own CLI asks for 10 minutes.

Tagging #

tag is a repeatable query parameter on uploadFile (and atfs upload --tag, above) — free-form text, up to 16 tags per upload, each 1-128 bytes with no control characters; an invalid tag is rejected with InvalidRequest rather than silently dropped or truncated, so a scripted upload that mistypes one fails loudly. A tag belongs to the claim your upload made — (this account, this cid) — not to the blob. Releasing a claim (deleteFile) takes its tags with it. Re-uploading bytes that already exist here unions your new tags into whatever that claim already carried rather than replacing them — content addressing means it's the same claim either way.

listFiles never says who applied a tag. By default it reports each file's tags as a flat union — every tag any claim on that cid carries here, account-class and mirrored-server alike — with no way to tell which claimant applied which; enumeration is public and unauthenticated, so it discloses nothing about who claims or tagged anything. To ask a claimant-scoped question, pass ?did= (below): the listing narrows to files that DID claims, and the reported tags narrow to that DID's own tags too.

Filtering follows the same rule: without did, ?tag= (below) matches a file when any claimant applied the named tag — a caller filtering by tag doesn't need to know, or learn, who applied it. Add did and the filter narrows to that DID's own tags, so ?did=X&tag=index means "files X claims where X applied the tag index".

One limitation remains deliberate for now: nothing removes or renames a tag once set — a mistyped tag is permanent short of deleting and re-uploading the content. pinFile (below) is no longer one of these limitations: it adopts the tags on the dev.atfs.file reference it's handed and they become the caller's own.

This is a labelling aid, not an index: no uniqueness, no guaranteed meaning, no lookup beyond scanning listFiles with a filter. It exists for exactly the kind of bookkeeping a release pipeline needs — tag every artifact of a build with its version and its role, then ask for them back by tag instead of keeping a separate manifest.

getFile #

Public — no token:

curl -sS "$ATFS_SERVER/xrpc/dev.atfs.repo.getFile?cid=bafkrei..." -o photo.jpg

Add &did=<did> to restrict the fetch to a file that DID holds a pin on; a miss (wrong cid, or a did that doesn't) comes back 400 BlobNotFound, not 404 (see "Retrieval" above).

deleteFile #

TOKEN=$(goat account service-auth --lxm dev.atfs.repo.deleteFile --aud "$ATFS_AUD")
curl -sS -X POST "$ATFS_SERVER/xrpc/dev.atfs.repo.deleteFile" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"cid": "bafkrei..."}'
{"contentDeleted": true}

pinFile #

The request body wraps a whole dev.atfs.file reference — the same JSON atfs upload prints (see "CLI" above) — under a file key:

FILE_REF='{
  "$type": "dev.atfs.file",
  "cid": {"$link": "bafkrei..."},
  "ipfsRoot": {"$link": "bafkrei..."},
  "size": 483821,
  "mimeType": "image/jpeg",
  "providers": ["https://myatfs.example.com"]
}'

TOKEN=$(goat account service-auth --lxm dev.atfs.repo.pinFile --aud "$ATFS_AUD")
curl -sS -X POST "$ATFS_SERVER/xrpc/dev.atfs.repo.pinFile" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d "{\"file\": $FILE_REF}"
{"state": "seeking"}

Re-calling with the same body is the status check — poll it, minting a fresh token each time since the last one is now spent:

while :; do
  TOKEN=$(goat account service-auth --lxm dev.atfs.repo.pinFile --aud "$ATFS_AUD")
  RESP=$(curl -sS -X POST "$ATFS_SERVER/xrpc/dev.atfs.repo.pinFile" \
    -H "Authorization: Bearer $TOKEN" \
    -H "Content-Type: application/json" \
    -d "{\"file\": $FILE_REF}")
  STATE=$(jq -r .state <<<"$RESP")
  echo "$STATE"
  case "$STATE" in
    pinned|failed) break ;;
  esac
  sleep "$(jq -r '.nextAttemptIn // 5' <<<"$RESP")"
done

nextAttemptIn is a relative second count, never a deadline (an appliance's clock can't be trusted — see "Pinning content from elsewhere" above); it's absent while a fetch is actually in flight (state: "fetching", which carries a progress object instead) or once the pin has settled, so the fallback above just avoids a tight loop in that gap.

listFiles #

Public, no token, cursor-paginated:

CURSOR=
while :; do
  PAGE=$(curl -sS "$ATFS_SERVER/xrpc/dev.atfs.repo.listFiles?limit=500&cursor=$CURSOR")
  jq -c '.files[]' <<<"$PAGE"
  CURSOR=$(jq -r '.cursor // empty' <<<"$PAGE")
  [ -z "$CURSOR" ] && break
done

cursor is present in a response only when that page filled to limit; its absence marks the last page.

Add &tag= (repeatable) to filter to files carrying every named tag — ?tag=index&tag=v1.2.3 lists only the index file tagged with that version. It ANDs, not ORs, and matches against any claimant's tags regardless of who applied them — deliberately, since neither the filter nor the output's tags field says who applied what (see "Tagging" above). Add &did= to scope both the listing and the reported tags to one claimant: ?did=X&tag=index means "files X claims where X applied the tag index". The cursor doesn't encode which tags or did produced it — it's a plain lexical cid comparison, applied before any filtering — so a paginated, filtered walk has to resend the same tag set (and did) on every page, exactly as it already has to keep limit stable; changing the filter mid-walk just silently changes what the rest of the pages return.

The alias trade-off #

dev.atfs.repo.uploadFile is also mounted, wire-identical, as com.atproto.repo.uploadBlob — so any existing client that already speaks uploadBlob via inter-service auth works against atfs unchanged, once it's pointed at the atfs instance and minting its token with the alias's own NSID. That doesn't come for free: a token minted with --lxm dev.atfs.repo.uploadFile is rejected on the com.atproto.repo.uploadBlob path, and vice versa — the two names share a wire contract, not a token. dev.atfs.repo.getFile's com.atproto.sync.getBlob alias is wire-identical too, but since retrieval needs no token at all, there's nothing there for an lxm to mismatch. deleteFile and pinFile have no com.atproto alias to begin with — atproto has no equivalent call for either to mirror.

Serving #

Plain HTTP — ATFS_HTTP_PORT (default 2837, or 80 in the Docker image; 0 to disable). Answers any Host header by default; set ATFS_HTTP_HOSTNAMES to a comma-separated allowlist and anything else gets 421 Misdirected Request. This is the only transport atfs itself runs, in-process.

Ingress, request-body caps, and custom domains — getting a public HTTPS hostname pointed at that port, and everything that comes with it (an ingress's own upload-size cap, registering an extra domain), is someone else's job: gosd's baked-in Cloudflare Tunnel and Tailscale Funnel agents on an SD-card image, or your own reverse proxy in Docker. See atfs.dev/docs/configuration/expose-to-the-internet/ for the full walkthrough on either platform. One fact worth knowing here regardless of platform: atfs itself never answers with a bare size rejection — its own ceiling (see "Storage ceiling" below) is a named 400 BlobTooLarge. A 413 means an ingress refused the body on sight, before atfs ever saw it.

did:web self-resolution — when serviceDid is a bare-domain did:web (e.g. did:web:atfs.example.com, nothing further after the domain), atfs serves its own GET /.well-known/did.json for exactly that domain's Host, with a service entry whose serviceEndpoint is derived from that same domain — one name, one address, nothing extra to configure or let drift out of sync. It's implicitly allowed through ATFS_HTTP_HOSTNAMES even if you forget to list it, and is cross-origin readable like the rest of the public retrieval surface (see "Retrieval" below). This makes the instance a resolvable atproto service with no external identity hosting — but the domain still has to actually terminate public TLS somewhere (gosd ingress or your own reverse proxy) for resolution to work. A did:plc serviceDid, or a did:web naming a path or port, isn't just missing this document — uploads stay disabled for it too (see "Configuration" above).

Retrieval #

  • GET /ipfs/<cid> — a standard IPFS gateway path, for either a blessed raw CID or a UnixFS root CID.
  • GET /xrpc/dev.atfs.repo.getFile?cid=<cid> (did optional; also served, did-required, as com.atproto.sync.getBlob) — including returning 400 BlobNotFound (not 404) for a miss, matching a PDS's own getBlob.
  • Any ordinary IPFS client — kubo, helia, anything speaking Bitswap — since atfs is a real DHT participant that announces what it holds, not a gateway-only node.

Both HTTP paths serve a blob's Content-Type exactly as its uploader declared it — the operator's own domain will render whatever an allowlisted account uploads. Every response also carries X-Content-Type-Options: nosniff and a locked-down Content-Security-Policy: default-src 'none'; sandbox (what a Bluesky PDS sends on getBlob and what kubo sends on its own path gateway), and anything outside a small inline-safe allowlist (images, video, audio, PDF, plain text) gets Content-Disposition: attachment, so a declared text/html downloads instead of rendering.

Both paths also send Access-Control-Allow-Origin: *, so a browser page on another origin can fetch() them — including a ranged read, which they answer with Access-Control-Expose-Headers covering Content-Range, Accept-Ranges and ETag so a caller can actually read the result back. listFiles (see "Enumeration" below) and the did:web document (see "did:web self-resolution" above) get the same treatment, for the same reason: every one of these is a public read already, so there's no narrower origin worth restricting to. uploadFile/uploadBlob, pinFile and deleteFile deliberately get none of this — they're authenticated by a caller's bearer token, and letting an arbitrary origin initiate one of those from a browser is exactly what CORS exists to prevent. This is what lets the browser setup flow fetch a card image and its settings manifest cross-origin before either touches disk.

Pinning content from elsewhere #

POST /xrpc/dev.atfs.repo.pinFile (JSON body {"file": {…a dev.atfs.file reference…}}) asks the instance to fetch content that already exists on the IPFS network and serve it as if you had uploaded it. The reference is the whole request: the cid the fetched bytes must hash to (so the fetch is trustless whatever the source), the ipfsRoot to fetch, the size to refuse before fetching a byte, and any providers to try over plain HTTPS before the IPFS network. Same allowlist as uploading, because a pin is an upload — just one where atfs goes and gets the bytes itself. If the reference carries a tags field, pinning adopts it: those tags land on your resulting claim, under your own DID, as if you'd supplied them yourself (see "Tagging" above). No reference tags just means the resulting claim starts untagged.

It answers immediately with a state — seeking, fetching, pinned or failed — because the bytes usually aren't here yet. Retrying is the node's job, not yours: a cid whose provider records haven't propagated (kubo's re-provide cycle runs 12–24 hours) is retried on a backoff doubling from a minute to an hour, for about 48 hours of the instance's uptime, and the request survives reboots. Only content that's provably wrong — bytes that don't hash to the cid, a file over the size limit — fails straight away. Call it again with the same file to check progress; call it again after a failure to start over. To stop waiting, deleteFile the same cid.

A providers list is someone else's instruction about where your instance should make connections, so it's treated as one: origins must be https, and anything resolving to an address that isn't routable on the public internet — loopback, your LAN, link-local, CGNAT — is never dialled, checked on the resolved address rather than the name. lastError is one of a small set of reasons (no-source-responded, content-mismatch, …; see the lexicon) rather than the error itself, which stays in the instance's logs where it can't map your network for whoever asked. The same applies to instances you follow: they name their own address, and it gets the same treatment.

Deletion #

POST /xrpc/dev.atfs.repo.deleteFile (JSON body {"cid": "<cid>"}) releases the caller's pin on a blob; the content itself is only physically removed once every claimant has released it. The same goes for content that hasn't arrived yet: a pending pinFile request is a claim before the fact, so deleting it withdraws your interest, and the fetch is abandoned only when the last account waiting on it has done the same. There's no com.atproto alias here — atproto has no equivalent blob-deletion call to mirror, since a PDS just garbage-collects a blob once no record references it. Authorization is by pin-set membership rather than the upload allowlist, so an account can always release its own claims, even after being removed from the allowlist.

Storage ceiling #

A single file is capped at 1 GiB too, whatever the volume looks like: uploadFile and pinFile both refuse anything larger with a named 400 BlobTooLarge (see "Request-body caps" above for how to tell that apart from an ingress's own rejection). atfs's primary payload is SD-card images a few hundred MB each, so the figure leaves headroom without leaving the cap unbounded — and, same as the reserve below, there's no config knob for it. It's deliberately absent from the published lexicons too, which say only that a blob exceeds "this instance's maximum blob size": a client has no business hardcoding a number that's really a property of the instance it's talking to.

An instance keeps a slice of its data volume permanently free — a twentieth of it, never less than 512 MiB and never more than 4 GiB (see internal/store) — because a genuinely full volume takes the metadata, pin records and IPFS index down with the content, and on an SD-card appliance there's no shell to clear it from. There's nothing to configure: the figure is derived from the volume itself.

Once what's left above that reserve is spoken for, uploadFile and pinFile answer 400 InsufficientStorage and store nothing; retrieval, enumeration and deletion carry on untouched, and deleteFile is how you make room. Space is reserved for the duration of each upload and each pin fetch rather than merely checked, so several large transfers at once can't all be waved through on the same free space. There's a ceiling on outstanding pinFile requests too (thousands — far past allowlist scale), since a pin for content that never turns up costs a small file and a retry timer indefinitely.

A follower that runs out of room stops adding for that poll and picks up where it left off an hour later. It never releases anything over it: the origin still lists what it holds, and being full is this instance's problem, not evidence the origin dropped a file.

Enumeration #

GET /xrpc/dev.atfs.repo.listFiles?limit=&cursor=&did=&tag= lists this instance's committed, directly claimed files, a page at a time — dev.atfs.file-shaped entries plus a firstSeen (see "First seen" below), so an entry still hands straight back to pinFile, which ignores the extra field. limit (default 500, max 1000) bounds the page; cursor resumes from a previous response's cursor field, which is present only when the page filled (its absence marks the last page). tag (repeatable) ANDs the listing down to files carrying every named tag — see "Tagging" above, including why a filtered walk has to resend the same tags (and did) on every page. did restricts the listing to files that DID holds a claim on, and scopes the reported tags (and what tag matches against) to just that DID's own claims — must look like a DID or the call is InvalidRequest. Unauthenticated, unlike the upload/pin/delete calls — every cid it lists is already public, announced to the DHT and served at /ipfs/<cid> — and, like /ipfs/ and getFile, cross-origin readable (see "Retrieval" above).

"Directly claimed" means uploaded here or pinned here by one of this instance's accounts. Content held only because this instance follows another is served and announced as normal but never listed here — a mirror doesn't re-export what it mirrors. That's what stops two instances following each other from echoing forever, lets an origin's deletions travel outward, and keeps mirroring non-transitive: follow each origin you actually want. A file mid-deletion or not yet indexed for IPFS is left off the listing too, so a follower should expect the set to shift slightly between polls and should never read a single absence as a deletion. Each entry's tags is a flat union, never attributed: every tag any of that file's claims carries, account-class and mirrored-server alike, with no indication of which claimant applied which — listFiles never discloses who pinned or tagged anything (see "Tagging" above). did above scoping the listing also scopes each entry's tags down to that one claimant. Naming a followed server's DID here doesn't surface mirror-only content either — listing eligibility is unchanged, so it still needs a local account-class claim too.

First seen #

Every listed file carries a firstSeen timestamp: when this instance first held those bytes. It's the instance's own observation, never the content's age — the same cid on two instances will report two different times, and neither is wrong. Nothing can set it, and no API changes it.

It survives claims coming and going. Re-uploading identical bytes, or another account pinning them, leaves it exactly where it was — content addressing means those are the same bytes either way, so the instance has held them since the first time. The one thing that resets it is the content actually leaving: once the last claim is released and the bytes are deleted (see "Deletion" above), a later upload or pin of the same cid starts the clock again. Mirroring from an instance you follow records when you first held the bytes, not when the origin did; a dev.atfs.file reference carries no such field, so there's nothing to adopt.

firstSeen is absent when the instance has no trustworthy reading to report, and that's deliberate rather than a gap. An appliance starts before networking and NTP are up — its clock reads 1970 for the first seconds of every boot — and a 1970 timestamp would be indistinguishable from a real one forever after. So atfs records nothing at all below a sanity floor, and sweeps those files once the clock is set, stamping them then. An absent firstSeen is therefore usually temporary: it means "this box couldn't tell you yet", not "this file has no history". Files stored before this existed at all get the same treatment on the next boot, since their 1970 stamps read as no stamp.

The same reading backs Last-Modified on /ipfs/<cid> and getFile — so a file with no trustworthy time is served without that header rather than with a 1970 one, which is also why firstSeen lives in listFiles' own output shape rather than in dev.atfs.file: a portable reference travels to other instances, where when this one first saw something means nothing.

Following another instance #

Add another instance's dev.atfs.server at-uri to your own record's follows list and this instance mirrors everything that instance directly claims:

{
  "$type": "dev.atfs.server",
  "accounts": ["did:plc:youraccountdid"],
  "serviceDid": "did:web:myatfs.example.com",
  "follows": ["at://did:plc:theirowner.../dev.atfs.server/12D3KooWtheirpeerid"]
}

It's the record's at-uri rather than a URL on purpose: the identity is what's followed, so the instance you follow can change hostname without you noticing (atfs just re-resolves and picks up the new one). atfs resolves that record to find where to poll — derived from its serviceDid, exactly like this instance's own address (see "did:web self-resolution" above) — and its serviceDid (whose name to hold the mirrored claims under — an instance with no serviceDid gets a did:key derived from its peer ID instead). Never your own DID: mirrored claims stay distinct from your uploads, so releasing one never touches the other, even when both happen to belong to the same DID.

Following is unilateral and needs no permission from the instance being followed. listFiles is public and every pinned cid is already a DHT provider record, so an opt-in would be theatre — anyone able to mirror your content can already do so.

Roughly hourly (fixed, no knob), atfs walks the followed instance's listFiles, pins anything new through the same machinery as pinFile, and gives up anything that has vanished. Two details worth knowing:

  • Releases need two consecutive polls. A file that drops out of one listing and comes back in the next is left alone — listFiles legitimately omits files mid-GC or mid-index, and releasing then re-fetching content is worse than an hour of staleness.
  • An unreachable instance is never an empty one. If the record can't be resolved, no endpoint answers, or a walk dies halfway, that poll releases nothing at all and tries again next time. Only a listing that completed can say a file is gone.
  • Tags travel with a first mirror, under the origin's own DID. The origin's listFiles already reports each file's tags as a union (see "Tagging" above), and that union lands on the mirror claim held under its serviceDid — never yours, so an operator can tell mirrored content apart from what they claimed themselves. A followed instance is a stranger, so its tags are treated defensively: anything that wouldn't pass validation here is dropped, and the set is capped at 16, rather than letting a bad or oversized tag list fail the mirror. This only happens the first time a cid is mirrored — follow diffs by cid, not by tags, so re-tagging a file at the origin after it's already mirrored doesn't propagate; the mirror keeps whatever labels it first saw. Travelled tags stay visible to this instance's own operator (and sit alongside any local account claim's tags on the same file), never re-exported: listFiles still never lists mirror-only content, so they don't leak onward to whoever follows you — the same rule that makes an A-follows-B-follows-A cycle converge instead of echoing.

Removing an entry from follows releases all of that instance's claims at once — no waiting — and the content is deleted once nothing else claims it. Anything you also uploaded or pinned yourself stays, since every claimant has to release before bytes go.

Reachability: what to expect #

HTTPS reachability comes from gosd's ingress agents (or your own reverse proxy in Docker) — they only ever dial outward, so any NAT, including CGNAT, still works. Whatever your network, your content is globally fetchable by URL.

The IPFS side — being fetchable by CID, from the swarm — is best-effort by the nature of NAT, and atfs climbs the whole ladder automatically:

  1. Router port-mapping (UPnP/NAT-PMP): atfs asks your router for a public port. Routers that honour it make the instance truly public — DHT server mode, gateways fetch on the first try — with no human involvement.
  2. Relay + hole-punching: otherwise atfs holds relay reservations and coordinates hole-punches, which pierce most home NATs (best over QUIC). Patient clients (ipfs get) fetch fine; a busy public gateway's short timeout may lose the race on a cold fetch.
  3. Relay only: symmetric NAT and CGNAT defeat hole-punching by construction — no software fixes that. Peers can still trickle through the relay, but treat HTTPS as the distribution channel there.

AutoNAT measures which rung you're on (the reachability changed log lines); it doesn't open anything. If you control the router and want gateway-grade p2p latency, forward TCP+UDP 4001 to the instance — it notices the change, promotes itself, and re-announces its content automatically.

Running on an SBC #

atfs also builds as flashable SD-card images via gosd, for the boards listed in .tangled/workflows/image.yml today (rock-4se, nanopi-zero2, pi-zero-2w, cubie-a5e) — adding another board is just adding its gosd board id there.

Flash the image, then open the config/ folder on the card's small atfs-boot partition (it mounts like any other FAT32 volume). Every setting is its own file, beside a .explain.md saying in plain language what to type into it, so setting config/env/ATFS_OWNER_DID is normally the only edit a new card needs — see packaging/config for the ones atfs adds, and gosd's docs/config.md for the tree as a whole. Blobs and this node's identity live on a separate atfs-data partition, mounted at the fixed /data; if that partition is missing or fails to mount, gosd falls back to an empty read-only volume instead, so a broken data partition fails loudly rather than silently losing writes. To expose the instance publicly, fill in one of the config/ingress/* settings groups — see "Serving" above.

Boards with onboard eMMC (NanoPi Zero2, Radxa Zero 3E) use it for storage automatically instead: atfsd formats and mounts a blank eMMC on first boot, then just mounts it on every boot after that. If the eMMC already holds something else, atfsd halts with instructions on the serial console rather than wiping it, and the same happens if the eMMC carries atfs's own volume but has become unmountable — gosd refuses to silently reformat either case. Set ATFS_DATA_DIR to opt out and use the SD /data partition instead, leaving the eMMC untouched.

To authorize either — wiping a non-blank eMMC, or rebuilding atfs's own volume if it stops mounting — set config/env/ATFS_FORMAT_DRIVES_IF_NOT_ATFS on the card to today's date, in UTC as YYYY-MM-DD: there's no shell to run mkfs by hand, so this is how an operator authorizes either without needing a serial adapter first. It has to be a date rather than true/false because gosd boards have no battery-backed clock — every boot starts at 1970 until the network comes up and SNTP syncs — so atfs waits (briefly, with a bounded timeout) for that sync before checking it, and only accepts today's date or yesterday's, so editing the card at 23:55 and booting at 00:05 still works. That also means the value can't be left on the card and silently authorize a wipe or rebuild again weeks or months later the way a bare true could, so there's no need to remove it afterwards. Which of the two actually happens depends only on what the eMMC needs: adopting foreign content is the harmless, expected case, but rebuilding atfs's own unhealthy volume destroys every blob on the drive and the node's identity key, so it comes back with a new peer ID and its dev.atfs.server record (whose rkey is the old peer ID) has to be recreated by hand — the serial console names which one happened.

Cards that arrive already configured #

An image can also be handed its settings before it's ever flashed. Every setting's file is padded to a fixed size, and atfs-<board>.inject.json alongside the .img records the byte ranges each one occupies; whoever distributes the image can overwrite exactly those bytes with real values — gosd's image-injection docs cover the mechanics, and it needs no understanding of FAT32, so a browser can do it between the CDN and the user's disk.

Because the reserved space is the settings file itself, injected settings are ordinary card settings: atfs needs no code to read them, they survive an upgrade reflash (gosd keeps a copy on /data and puts it back), crash reports redact them, and whoever holds the card can read and edit them beside the same explanation as every other setting. A card nobody injected anything into ships the same defaults it always did.

This is what lets a setup flow create the dev.atfs.server record and the card in one sitting: injecting ATFS_IDENTITY_KEY fixes the instance's peer ID up front, and the record's rkey is that peer ID, so it no longer has to wait for a first boot to report an identity it invented. One consequence to know: since the snapshot carries those values into /data and re-renders them onto the next card, reflashing is not a way to clear a device's identity key — only wiping /data is.

That flow is atfs.dev/setup/ — a Svelte app whose source is web/src/setup, one of the three pages the web/ Vite package builds, landing at site/setup when the site deploys (make web does the same locally, offline included — see the next paragraph). (atfs.dev/downloads/ still exists, as a redirect to the new location.) It signs in with atproto OAuth, mints the identity in the browser, writes the record, and then — behind the same SD-card/Docker toggle the docs site uses — either splices ATFS_OWNER_DID and ATFS_IDENTITY_KEY into the .img's reserved ranges (for an image it downloads, or one already on your disk) or fills both into a ready-to-run docker run command. Downloading the .img means fetching it and its .inject.json manifest with the browser's own fetch(), so wherever they're hosted has to send the CORS headers described under "Retrieval" above — an atfs instance itself does, out of the box. The key is generated in the tab and sent nowhere, so it also asks you to save it: pasting it back is how a rebuilt server keeps the identity its record is keyed by. The splicing is @jphastings/gosd; the identity and record logic is atfs's own, and webtest/ checks its arithmetic against go-libp2p's on shared fixtures, so a browser-minted key and the peer ID atfsd derives from it cannot drift apart.

Plain downloads — no browser, no OAuth — live at atfs.dev/downloads/raw/: every board's .img and .inject.json, with sizes, for whichever release is newest. Both that page and the setup flow's in-browser catalogue (images.json) are generated at deploy time by hack/bake-site-downloads.sh, which queries the hosting instance live for the newest release's manifest rather than reading anything from this repo (see releases/README.md and that script's own header). That page also publishes the index it was built from, at atfs.dev/setup/index.json — which is what a later deploy reads back when the instance is unreachable, so it republishes exactly what's already advertised rather than failing or blanking the page. A deploy fails only when neither the instance nor that copy answers, and until a first release exists anywhere there is nothing to read: deliberately, since publishing an empty downloads page over a good one is the worse outcome. make web is exempt — it degrades to a "nothing released yet" page rather than requiring the network, so make check still works offline. The two halves have different requirements: a plain download is a navigation and works against anything, while the setup flow's own fetch of a catalogued image is a cross-origin read, so it needs the hosting instance to send CORS headers — which atfs does (see "Retrieval" above), so an atfs-hosted release works in the browser flow as well as by plain download.

There's no shell and no SSH on a gosd image: atfsd — the same binary as the container, built with gosd's gosd build tag — is the only binary gosd-init runs, restarted forever with backoff if it ever exits, and the serial console is the entire interface. That's where boot shows which release the card is running, this instance's peer ID, the at:// URI for its config record, and whether uploads are enabled.

Release images #

Every board's .img and its .inject.json (see above) are also published on an atfs instance — dogfooding atfs's own upload API for its own release hosting, rather than reaching for object storage. Tagging a release (see "Building from source" below) triggers .tangled/workflows/publish-sbc-images.yml, which builds every board's image, uploads each file with the atfs CLI (see "CLI" above), and assembles and uploads a manifest of where everything landed, tagged index (see releases/README.md for its shape). Each entry is a dev.atfs.file reference plus a direct /ipfs/<cid> URL, served by the hosting instance itself at any size (see "Retrieval" above) — also fetchable over the IPFS network by any Bitswap client, subject to that instance's own reachability (see "Reachability" above). "At any size" is true only of that download: getting each image onto the hosting instance in the first place is the ordinary upload flow, so it still has to clear the 1 GiB blob cap and whatever the publishing workflow's own ingress carries (see "Storage ceiling" and "Request-body caps" above) — atfs's own images, a few hundred MB each, comfortably do.

atfs.dev's own site deploy (.tangled/workflows/deploy-site.yml, and this same workflow once it's uploaded a release) picks up the current release live: hack/bake-site-downloads.sh queries the hosting instance for dev.atfs.repo.listFiles?tag=index&tag=atfs, takes the newest vX.Y.Z tag it finds, and fetches that entry's manifest — nothing about a release is committed to this repo. A site build fails outright, before anything is published, if no valid index can be found (see that script's header for every case that covers) — which is why this repo's own site deploy can't go live until a first real release has been published.

Building from source #

make build   # -> bin/atfsd, bin/atfs
make test
make vet
make check   # vet + test

go.mod pins the minimum Go toolchain, and CI sets GOTOOLCHAIN=auto so an older local toolchain fetches the right one automatically rather than failing. Build with make build or go build ./cmd/atfsd rather than a bare go build ./...: with more than one main package in this module, that discards its output instead of producing a binary you can run.

With nix, nix build .#atfs builds the CLI against the toolchain and dependency set flake.lock pins, and nix develop gives a shell holding the Go version go.mod asks for without installing anything system-wide. The flake packages only atfs: atfsd already ships as a container image and as an SD-card image (both above), and go mod vendor works over whatever source is present — so packaging the daemon here would pull the entire libp2p dependency tree into every CLI build, including the ones in other people's CI. .tangled/workflows/flake.yml builds and smoke-tests the flake on every pull request.

The whole module is CGO-free by design (CGO_ENABLED=0) — load-bearing for both the gosd arm64 cross-compile and the scratch-based Docker image, which ships no libc, shell, or other runtime dependency at all.

Status #

atfs is young — pre-first-release, and its shape may still move. The dev.atfs.server and dev.atfs.file lexicons are drafted in lexicons/ but not yet published to the network under the dev.atfs authority. A project site at atfs.dev is planned but doesn't exist yet.