STreaming ARchive format: simpler, stricter, smaller than CARs for a subset of use-cases.
star car archive atproto
README.md

Fuzz targets #

Requires nightly and cargo-fuzz:

cargo install cargo-fuzz

Targets #

target what it covers
read_arbitrary Reader over arbitrary bytes: never panics, never yields more items than the input has bytes, and any error is terminal
verify_arbitrary verify_archive over arbitrary bytes: header parse, MST reconstruction, and the LtHash fold
roundtrip_space structured: anything encode_space_archive accepts must verify and read back identically
flaky_source Reader over a source that short-reads and errors on a fuzzer-chosen schedule

Running #

The two byte-oriented targets need the seed corpus to get past the header — random bytes essentially never form a valid archive, and without seeds they only ever exercise "reject bad magic". Pass seeds/<target> as a second corpus directory:

cargo +nightly fuzz run read_arbitrary   corpus/read_arbitrary   seeds/read_arbitrary
cargo +nightly fuzz run verify_arbitrary corpus/verify_arbitrary seeds/verify_arbitrary

The seeds are worth roughly 30x the edge coverage on verify_arbitrary.

The structured targets generate their own valid inputs, so they need no seeds:

cargo +nightly fuzz run roundtrip_space
cargo +nightly fuzz run flaky_source

Why flaky_source exists #

The other targets read from a slice, which only ever succeeds or hits clean EOF. Real sources — sockets, pipes, failing disks — return Interrupted, short reads, and persistent errors, and a reader is most likely to spin or lose its place on exactly those paths.

This is not hypothetical: the reader once had a bug where an error left the iterator un-finished, so for entry in reader looped forever. It was not reachable from a truncated byte slice (the next read hits clean EOF and ends the iteration) — only from a source that keeps erroring. A slice-based fuzzer would never have found it.

seeds/ is checked in; corpus/ and artifacts/ are generated and ignored.