diff --git a/devlog/014-post-shaped-files.md b/devlog/014-post-shaped-files.md new file mode 100644 index 0000000..9c40feb --- /dev/null +++ b/devlog/014-post-shaped-files.md @@ -0,0 +1,100 @@ +# post shaped files + +Dan Abramov spent part of June 30 trying to explain why atproto still feels worth arguing for. the good posts were not about Bluesky feature work. they were about the shape under it. he wrote that he was "just putting post shaped files into my pds," that the app can be forked "together with all of its users and data," and that the topology makes other things changeable later. later in the thread he tried a shorter handle: atproto as a "model layer" for the web, "web for json." ([one](https://bsky.app/profile/danabra.mov/post/3mphos3a2d22o), [two](https://bsky.app/profile/danabra.mov/post/3mphq66fkec2o), [three](https://bsky.app/profile/danabra.mov/post/3mphqcfwres2o), [four](https://bsky.app/profile/danabra.mov/post/3mpjcy44xoc2z)) + +that gets close to why zat exists. + +most atproto projects should never care about CAR block boundaries or secp256k1 high-S signatures. if this record web is going to grow past Bluesky-the-app, people need small pieces of infrastructure they can actually run: relays, PDSs, indexers, backfill workers, verifiers, exporters, search systems, weird appviews, boring CLIs. those programs live below product taste. they move bytes, check signatures, walk repos, and decide whether a piece of network state deserves to be trusted. + +zat is for that layer. + +![records first, apps after](https://zat.dev/docs/devlog/img/record-projections.svg) + +## rss was right + +Paul Frazee put the same idea another way: RSS was an underpowered database replication protocol. Google Open Source replied with the joke query, `SELECT , FROM ?`, and Paul answered with the atproto database he has been playing with: + +```js +atdb.query({ + from: {collection: "app.bsky.feed.post"}, + where: [{eq: {field: "/lang", value: "en"}}], + order: {by: "@indexedAt", dir: "desc"}, + limit: 50 +}) +``` + +the joke works because RSS really did have part of the shape: publish something at a URL, let other programs fetch and transform it, stop asking one website to be every interface at once. atproto keeps that instinct and adds the parts RSS never had: portable identity, signed repositories, typed record collections, content-addressed blocks, a firehose, and a repo format that can express mutations instead of only snapshots. + +Dan's gloss was less SQL and more files: the old Web 2.0 dream of services talking to each other broke because every service had to expose an API and any service could break the chain. "here there's no API because we just aggregate from files. RSS was right." ([post](https://bsky.app/profile/danabra.mov/post/3mpjdarnoc22z)) + +that line is slightly too clean. there are APIs all over atproto. XRPC exists. relays and PDSs have behavior, policy, rate limits, bugs, and bills. but the important thing is where the durable state lives. a post is not only a row in Bluesky's database. it is a record in a repo, under a DID, with a key, a collection, a CID, and a signature chain. + +apps can project that state differently. one app can care about replies. another can care about bookmarks. another can index `site.standard.document` records. another can rebuild a feed from records whose lexicons Bluesky does not understand. people can disagree about ranking, moderation, logged-out discovery, or blocking semantics without needing the underlying data to move first. + +that part is worth being excited about. + +## why another sdk + +there are already atproto SDKs. the TypeScript packages are the reference surface for app developers. Indigo and Atmos prove a lot of production infrastructure in Go. Python is good glue. Rust has serious pieces for people already in that ecosystem. + +zat is not chasing the "most pleasant way to call `app.bsky.feed.getTimeline`" trophy. the repo and network machinery are the focus: + +- parse DIDs, handles, NSIDs, TIDs, record keys, and AT URIs +- encode and decode DAG-CBOR +- read and write CARs +- verify CID hashes and commit signatures +- load and walk Merkle Search Trees +- consume Jetstream and raw firehose events +- resolve identity without turning untrusted handles into SSRF +- make checked XRPC calls where errors keep their protocol envelope +- help programs author records, commits, and firehose frames + +some of that is user-facing. most of it is plumbing. downstream programs should not each re-learn the same unpleasant facts about canonical CBOR, repo completeness, rate-limit headers, DID documents, websocket frames, or high-S JOSE signatures. + +the current consumers make that less theoretical. + +[`zlay`](https://tangled.org/zzstoatzz.io/zlay) uses zat's bytes-and-trust layer to run relay-shaped work: decode event streams, track hosts, validate commits, and fan out frames. [`zds`](https://tangled.org/zat.dev/zds) uses zat to build a real PDS: write records, update MSTs, sign commits, emit `subscribeRepos`, import repos, and handle OAuth-adjacent protocol work. [`pub-search`](https://tangled.org/zzstoatzz.io/pub-search) grew out of leaflet search and uses the network as a corpus rather than treating Bluesky as the product boundary. [`atproto-bench`](https://tangled.org/zzstoatzz.io/atproto-bench) keeps us honest by making implementation claims measurable. the docs publisher in this repo is smaller, but I still like it: zat publishes its own docs as `site.standard.document` records, through zat. + +those are different programs. they want the same substrate. + +![zat as substrate](https://zat.dev/docs/devlog/img/zat-substrate.svg) + +## why zig + +this part is easy to make corny, so the narrow version first: Zig is a good fit for parsers, verifiers, stream processors, and little deployable tools. + +atproto has a lot of work where allocation shape matters. a firehose frame is DAG-CBOR plus an embedded CAR. repo verification walks content-addressed blocks and MST nodes. search and backfill systems may read huge numbers of records whose schema they mostly do not care about. a PDS needs to author the same structures without smuggling in a runtime that dominates the service. + +Zig lets zat keep those paths close to the bytes. CBOR strings and byte strings can be slices into the input. CIDs can be small references until a caller asks for fields. arena allocation can match frame and repo lifetimes. the same package can expose a library and also build tiny tools: smoke tests, decode benchmarks, a docs publisher. shipping one binary is not a personality trait. it is handy when the programs are infrastructure-shaped. + +the dependency story matters too. as of this writing, zat has one runtime dependency, [`websocket.zig`](https://tangled.org/zzstoatzz.io/websocket.zig), and a lazy test dependency on the official [`atproto-interop-tests`](https://github.com/bluesky-social/atproto-interop-tests) fixtures. the rest is stdlib and repo-local code. that is not moral purity. it makes the protocol behavior easier to inspect. when an MST node with duplicate keys is accepted or rejected, I want the answer to be in a file we can read. + +Zig's own culture matters here, but only if we keep the claim narrow. the Zig project has a strict no-LLM/no-AI rule for its own project spaces: no generated code or prose, no LLM editing, no LLM-assisted bug finding, no LLM brainstorming fed back into Zig project work. the same page is explicit that these rules govern the Zig-owned Codeberg, IRC, and Zulip spaces, not every community or application written in Zig. ([policy](https://ziglang.org/code-of-conduct/#strict-no-llm-no-ai-policy)) + +that distinction matters. people build Zig applications with all sorts of workflows. Ghostty, Mitchell Hashimoto's terminal emulator, is one visible example of a large Zig application with heavy AI-assisted development, while Zig-the-language keeps a much stricter process for its own core work. ([Ghostty](https://ghostty.org/), [Mitchell Hashimoto](https://mitchellh.com/)) + +the lesson for zat is narrow: the language is being designed in a space that strongly protects direct technical judgment, small surfaces, and maintainable invariants. Zig is not perfect. the async story has churned more than once, and churn has costs. but the ethos still lines up with this library: know what allocates, know what can fail, keep dependencies boring, make the binary portable, read the code when the edge case matters. + +zat is written inside that tension. I am using AI while writing this devlog. the library still has to earn trust the old way: interop tests, production consumers, benchmarks, code review, and a habit of refusing malformed network input. + +## fast is part of the deal + +performance is not a trophy here. if reading and verifying the network requires a large team and expensive managed infrastructure, the record web drifts back toward platform gravity. fewer people build indexers. fewer people keep backups. fewer people run alternate appviews. fewer experiments survive contact with the firehose. + +so zat spends time on boring speed. CBOR values got smaller. CID parsing became lazy. CAR lookup became O(1). secp256k1 verification got its own optimized path. MST lookup copied the plain Atmos shape after benchmarks showed our clever version was worse. full repo verification learned to reject incomplete CARs that parsed cleanly. XRPC stopped retrying non-idempotent POSTs on 5xx. handle DNS resolution stopped accepting ambiguous `did=` TXT records. + +that last list sounds like release notes because it is. but the shape behind it is broader: a replicated record network needs cheap readers that distrust what they read. + +Jim Calabro's [`Atmos`](https://github.com/jcalabro/atmos) has been useful for exactly that reason. it comes from production PDS work, so its refusals are interesting. when Atmos rejects non-canonical MST nodes, incomplete repo CARs, ambiguous DNS identity records, or unsafe POST retries, that is not pedantry. that is a production implementation saying where the floor is thin. + +zat should learn from that. + +## what I want from this + +I want atproto projects to get to the weird part faster. + +if someone wants to make a search engine for long-form records, they should not start by writing a CAR parser. if someone wants to run a tiny PDS, they should not rediscover repo signing and MST block collection from scratch. if someone wants to benchmark a relay path, they should not first spend a week finding out that a syntactically valid repo export can be missing blocks. if someone wants to publish a new record type, they should not need Bluesky-the-app to understand it before the network can carry it. + +that is the bet in zat: a small, fast, fairly complete Zig implementation makes more experiments possible at the layer where atproto is most interesting. + +posts are files now, sort of. files with DIDs, CIDs, collections, signatures, lexicons, and awkward edge cases. that is enough of a new primitive to be worth building good tools around. diff --git a/devlog/img/record-projections.svg b/devlog/img/record-projections.svg new file mode 100644 index 0000000..4d8b66d --- /dev/null +++ b/devlog/img/record-projections.svg @@ -0,0 +1,86 @@ + + From app databases to atproto record projections + Diagram comparing an app-owned database exposed through APIs with atproto user repos feeding multiple projections. + + + + records first, apps after + the useful atproto trick is topology: signed repos can outlive any one projection + + + usual app shape + + + + app + + + + + private database + schema, ranking, policy + + + + + API clients + + + + atproto shape + + + + + user repo + signed records + + + + user repo + signed records + + + + user repo + signed records + + + + + + + + relay / firehose + + + + + + + + social app + + + + search + + + + archive + + + + + apps still matter. the point is that their data source can be shared, signed, copied, indexed, and verified. + + + + + + + + + + + diff --git a/devlog/img/zat-substrate.svg b/devlog/img/zat-substrate.svg new file mode 100644 index 0000000..357c872 --- /dev/null +++ b/devlog/img/zat-substrate.svg @@ -0,0 +1,81 @@ + + Zat as a Zig substrate for atproto infrastructure + Layered diagram showing Zat protocol primitives underneath downstream consumers such as zlay, zds, pub-search, docs publishing, and atproto-bench. + + + + why zig here + zat tries to keep the protocol substrate small enough to read and fast enough to run hot + + + + downstream programs + + + + + + + + + zlay + zds + pub-search + publish-docs + atproto-bench + + + + + + zat + protocol pieces written close to the bytes + + + + syntax + DID, handle, NSID, AT URI + + + + repo data + CBOR, CID, CAR, MST + + + + trust + keys, JWT, repo verification + + + + network + XRPC, OAuth, firehose + + + + + one runtime dependency, lazy official interop fixtures for tests, binaries for tools + + + + + + small dependency surface + + arena-friendly parsing + + portable little binaries + + + + + + + + + + + + + +