diff --git a/README.md b/README.md index 972949a..0780b08 100644 --- a/README.md +++ b/README.md @@ -1,129 +1,67 @@ @okikio/opfs ============ -`@okikio/opfs` gives application code one filesystem programming model across browser OPFS, Deno, Bun, Node, key-value stores, document databases, and SQL databases. +`@okikio/opfs` is an OPFS-shaped filesystem programming model that can sit on top of browser OPFS, host filesystems, +object stores, key-value stores, browser storage, document databases, and SQL databases. -The frontend can use either path-based filesystem methods or OPFS-shaped file and directory handles. The backend is selected with an adapter. +The public filesystem owns the semantics that application code should not have to rebuild: canonical virtual paths, +OPFS-shaped handles, recursive copy and move, cancellation, staged writable files, bounded stream fallbacks, coordination, +normalized failures, and resource ownership. An adapter translates those operations into one concrete backend. ```text -application code - | - | path API handle API - | readFile('/a.txt') root.getFileHandle('a.txt') - | writeFile('/a.txt', bytes) file.createWritable() - +-------------------+-------------------------+ - | - v - FileSystemType - | - v - AdapterType - | - +------------------+-------------------+ - | | | - v v v - native filesystem RecordStoreType custom adapter - | | - | +---------+---------+---------+ - | | | | | - v v v v v - OPFS/Deno/ unstorage RxDB db0 Drizzle - Bun/Node -``` - -This means OPFS-style application code does not have to know whether the bytes are in the browser's Origin Private File System, a server directory, an unstorage mount, an RxDB collection, a db0 database, or a Drizzle table. - -The reverse direction is also supported. `@okikio/opfs/driver/unstorage` exposes any `FileSystemType` as an unstorage driver. An application can therefore mount an OPFS, Deno, Bun, Node, RxDB, db0, or Drizzle-backed filesystem inside unstorage. - -Why the adapter is the important abstraction --------------------------------------------- - -The browser File System API is a useful frontend contract, but OPFS is only one persistence system. A library that hard-codes `navigator.storage.getDirectory()` cannot reuse that filesystem code on a server or over a database. - -This package separates the two responsibilities: - -- `FileSystemType` owns virtual paths, OPFS-shaped handles, recursive operations, cancellation, coordination, errors, and resource lifecycle. -- `AdapterType` owns the smallest backend primitive set needed to persist files and directories. -- `RecordStoreType` maps the filesystem primitives onto value, document, or SQL records when a backend is not naturally file-based. - -The separation is deliberate. It keeps backend-specific behavior out of application code without pretending that every backend has the same performance or durability characteristics. - -Pre-release use ---------------- - -This source tree is being prepared for JSR and npm publication. Until the package is published, consume it through the workspace or another explicit local source reference instead of assuming the registry entry exists. After release, the intended imports are: - -```ts -import { createFileSystem, openFileSystem } from "jsr:@okikio/opfs"; -``` - -and the intended npm-compatible install is `npm install @okikio/opfs`. Release validation must prove both registry artifacts before these forms are treated as available. - -Drizzle integration also needs the optional peer dependency: - -```sh -npm install drizzle-orm +application + | + +-- path API -------------------+ + | readFile / writeFile | + | copy / move / walk | + | v + +-- OPFS-shaped handles --> FileSystemType + | + canonical adapter operations + | + +---------------------+---------------------+ + | | | + v v v + native files record stores object stores + OPFS / Node / Deno KV / DB rows S3 / Azure Blob + / Bun | + +-- localStorage + +-- IndexedDB + +-- Cache Storage + +-- Deno KV + +-- unstorage + +-- RxDB + +-- db0 / SQLite + `-- Drizzle ``` -The root module is browser-safe and does not import Node, Bun, Deno, RxDB, unstorage, db0, or Drizzle at import time. Runtime-specific integrations live on explicit subpaths. +The reverse direction is useful too. `@okikio/opfs/driver/kv` exposes any `FileSystemType` as a small hierarchical key-value +store, and `@okikio/opfs/driver/unstorage` adapts that view to unstorage. This means an application can give unstorage an OPFS, +Node, Deno, Bun, S3, Azure Blob, IndexedDB, Deno KV, SQLite, or another custom OPFS backend without a second provider matrix. -Use native browser OPFS ------------------------ +Install and start with the backend you actually own +---------------------------------------------------- -`openFileSystem()` is the shortest path when the browser's OPFS is the intended backend. +Deno and JSR can import the package directly: ```ts -import { openFileSystem } from "@okikio/opfs"; +import { openFileSystem } from "jsr:@okikio/opfs"; const fileSystem = await openFileSystem(); - -await fileSystem.writeFile( - "/projects/kaiju/settings.json", - JSON.stringify({ capture: true }), - { parents: true }, -); - -const settings = JSON.parse( - await fileSystem.readText("/projects/kaiju/settings.json"), -); -``` - -`openFileSystem()` is equivalent to opening the native OPFS root, creating the OPFS adapter, and passing it to `createFileSystem()`. - -```ts -import { createFileSystem } from "@okikio/opfs"; -import { createOpfsAdapter } from "@okikio/opfs/adapter/opfs"; - -const root = await navigator.storage.getDirectory(); -const fileSystem = createFileSystem(createOpfsAdapter(root)); -``` - -Use one long-lived positional file ------------------------------------ - -Media muxers and database engines can rewrite earlier byte ranges while output is still open. Use `openWritableFile()` for that access pattern instead of issuing one `writeFile(update)` operation per chunk. - -```ts -const file = await fileSystem.openWritableFile("/output.mp4", { - create: true, - parents: true, -}); - try { - await file.write(header, { at: 0 }); - await file.write(mediaChunk, { at: offset }); - await file.flush(); - await file.close(); -} catch (error) { - await file.abort(error); - throw error; + await fileSystem.writeFile("/state/app.json", "{}", { parents: true }); +} finally { + await fileSystem.close(); } ``` -The OPFS, Node, Deno, and Bun adapters advertise this capability. Record-backed adapters such as memory, unstorage, RxDB, db0, and Drizzle do not. They remain appropriate for small records and ordinary bounded writes, but the facade will not disguise repeated record replacement as a native positional-file resource. +npm-compatible runtimes use the same TypeScript API: -Use OPFS-shaped handles over Node ---------------------------------- +```sh +npm install @okikio/opfs +``` + +Server code selects a concrete adapter instead of importing a different filesystem API: ```ts import { createFileSystem } from "@okikio/opfs"; @@ -133,284 +71,253 @@ const fileSystem = createFileSystem( createNodeAdapter({ root: "./data" }), { coordination: "local" }, ); - -const projects = await fileSystem.root.getDirectoryHandle("projects", { - create: true, -}); -const file = await projects.getFileHandle("state.json", { create: true }); -const writable = await file.createWritable(); - -await writable.write(JSON.stringify({ ready: true })); -await writable.close(); ``` -The application uses a File System API-shaped frontend. The bytes are written with Node filesystem APIs below `./data`. - -Deno and Bun use the same frontend: +The root entrypoint is intentionally import-safe in browsers, workers, Deno, Bun, and Node. Runtime-specific dependencies stay +on explicit subpaths. Importing `@okikio/opfs` does not import `node:fs`, inspect environment variables, connect to databases, +or configure application logging. + +The first-party backend set is deliberately broad, but the layers stay small: + +| Subpath | Backend or role | Important behavior | +| --- | --- | --- | +| `adapter/opfs` | native browser OPFS | native handles, streams, sync access when exposed | +| `adapter/node` | `node:fs` | streams, ranges, copy/rename, sync random access | +| `adapter/deno` | Deno filesystem | streams, ranges, copy/rename, sync random access | +| `adapter/bun` | Bun + Node-compatible fs | Bun read/write fast paths plus host filesystem operations | +| `adapter/memory` | in-memory records | deterministic tests and temporary state | +| `adapter/record` | `RecordStoreType` | common translation for value/document/SQL stores | +| `adapter/object` | `ObjectStoreType` | common translation for object stores without hiding object semantics | +| `adapter/s3` | `S3ClientType` | direct S3/S3-compatible storage, no AWS SDK | +| `adapter/azure` | `AzureClientType` | direct Azure Blob REST storage, no Azure SDK | +| `adapter/localstorage` | Web Storage | synchronous string store translated through records | +| `adapter/indexeddb` | IndexedDB | indexed record persistence with caller-controlled database ownership | +| `adapter/cache` | Cache Storage | record persistence in an injected Cache | +| `adapter/deno-kv` | Deno KV | record persistence over a caller-owned KV database | +| `adapter/sqlite` | connected SQLite | focused SQLite view over the same SQL record contract as db0 | +| `adapter/unstorage` | unstorage `Storage` | forward bridge above the selected unstorage driver | +| `adapter/rxdb` | RxDB collection | forward bridge above the selected RxStorage | +| `adapter/db0` | db0 `Database` | SQL bridge across db0 dialects/connectors | +| `adapter/drizzle` | Drizzle database + table | caller-owned schema and common CRUD bridge | +| `driver/kv` | reverse key-value view | collision-safe key hierarchy over any filesystem | +| `driver/unstorage` | reverse unstorage driver | lets unstorage consume any `FileSystemType` | + +Drizzle is an optional peer dependency because the integration is only loaded through its explicit subpath. + +Object storage is not flattened into a fake local disk +------------------------------------------------------- + +S3 and Azure Blob can both back the filesystem facade, but they remain object stores underneath. That distinction affects +performance and correctness. + +A complete replacement can stream to multipart/block upload. Append and update cannot normally mutate object bytes in place, +so the object adapter performs a read-modify-write operation. When the provider exposes conditional writes, the previous ETag +is used as an optimistic precondition so a concurrent writer fails rather than being silently overwritten. + +Native copy is also a separate capability. The filesystem asks the adapter to copy before it opens a source stream, so +provider-side copy stays inside S3/Azure instead of becoming an accidental download and re-upload. -```ts -import { createFileSystem } from "@okikio/opfs"; -import { createDenoAdapter } from "@okikio/opfs/adapter/deno"; -// or: import { createBunAdapter } from "@okikio/opfs/adapter/bun"; - -const fileSystem = createFileSystem( - createDenoAdapter({ root: "./data" }), - { coordination: "local" }, -); +```text +filesystem.copy() + | + +-- nativeCopy ----> provider/server-side copy + | + `-- fallback ------> source stream -> bounded transfer -> destination ``` -Use unstorage as the backend ----------------------------- - -The adapter receives an already-created high-level unstorage `Storage` object. It therefore works above the individual unstorage driver choice. +The direct S3 client implements Signature Version 4, range reads, ListObjectsV2, multipart upload, conditional completion, +CopyObject, and multipart UploadPartCopy for objects above CopyObject's 5 GB source limit. It also checks S3's unusual +success-with-error-body responses for copy and multipart completion. ```ts -import { createStorage } from "unstorage"; -import memoryDriver from "unstorage/drivers/memory"; import { createFileSystem } from "@okikio/opfs"; -import { createUnstorageAdapter } from "@okikio/opfs/adapter/unstorage"; - -const storage = createStorage({ driver: memoryDriver() }); -const fileSystem = createFileSystem(createUnstorageAdapter(storage)); +import { createS3Adapter } from "@okikio/opfs/adapter/s3"; +import { createS3Client } from "@okikio/opfs/s3"; + +const client = createS3Client({ + endpoint: "https://s3.us-east-1.amazonaws.com", + bucket: "my-bucket", + region: "us-east-1", + credentials: { accessKeyId, secretAccessKey }, +}); -await fileSystem.writeFile("/cache/report.json", "{}", { parents: true }); +const fileSystem = createFileSystem(createS3Adapter(client)); ``` -The same bridge can sit above compatible unstorage fs, Redis, S3, MongoDB, IndexedDB, Cloudflare, Vercel, db0, Deno KV, and other current unstorage drivers. Some upstream drivers are read-only. Use `{ readOnly: true }` when mutations cannot be supported. - -Use RxDB as the backend ------------------------ - -The RxDB integration targets an `RxCollection`, not one particular `RxStorage`. Create one collection from the exported schema and then use whichever RxStorage configuration is appropriate for that database. - -```ts -import { createFileSystem } from "@okikio/opfs"; -import { - createRxDbAdapter, - RxDbRecordJsonSchema, -} from "@okikio/opfs/adapter/rxdb"; - -const database = await createRxDatabase({ - name: "app", - storage: selectedRxStorage, -}); +S3 compatibility is a protocol family, not one identical product. The client therefore accepts endpoint, region, addressing, +headers, and capability overrides. For example, an S3-compatible service that does not support multipart preconditions should +set `conditionalWrite: false` instead of pretending the safety property exists. `client.request()` remains available for S3 +features that do not belong in the portable filesystem contract. -await database.addCollections({ - files: { schema: RxDbRecordJsonSchema }, -}); +Azure uses its own REST model instead of being forced through an S3 abstraction. It supports SAS, Microsoft Entra bearer +tokens, Shared Key, and caller-defined authorization headers. Large server-side copies use Put Block From URL after Azure's +smaller synchronous Copy Blob From URL path is no longer sufficient. -const fileSystem = createFileSystem( - createRxDbAdapter(database.files), - { coordination: "local" }, -); -``` +The protocol clients are documented separately because their wire contracts are larger than the filesystem adapter surface: -This design preserves RxDB's own storage-engine abstraction. It does not clone each RxStorage implementation into this package. +- [S3 client protocol](./docs/s3.md) covers SigV4, request canonicalization, multipart upload/copy, conditions, limits, errors, + compatibility controls, and known non-goals. +- [Azure Blob client protocol](./docs/azure.md) covers REST versions, SAS/bearer/Shared Key authorization, block upload/copy, + conditions, limits, errors, and Azurite behavior. +- [Provider integration tests](./docs/providers.md) explains the Docker-backed SeaweedFS and Azurite test matrix and what those + emulators can and cannot prove. -Use db0 as the backend ----------------------- +The facade makes capability, limits, routing, and cost inspectable +---------------------------------------------------------------- -`createDb0Adapter()` targets db0's high-level `Database` contract. It selects portable SQL for the database's reported `sqlite`, `libsql`, `postgresql`, or `mysql` dialect. +`AdapterCapabilitiesType` describes immediate adapter behavior. `FileSystemType.inspect()` describes the configured stack after +facade fallbacks and optimization policy are applied. This distinction lets callers ask whether a route is `native`, `emulated`, +`partitioned`, or `unsupported` without guessing from the adapter name. ```ts -import { createFileSystem } from "@okikio/opfs"; -import { createDb0Adapter } from "@okikio/opfs/adapter/db0"; - -const adapter = await createDb0Adapter(database, { - table: "opfs_entries", - initialize: true, -}); - const fileSystem = createFileSystem(adapter, { - coordination: "local", + maxBufferedWriteBytes: 32 * 1024 * 1024, + metrics: "basic", + optimizations: { + nativeCopy: false, + }, }); + +console.log(fileSystem.inspect()); +console.log(fileSystem.plan({ + operation: "write", + source: "stream", + mode: "replace", + size: 512 * 1024 * 1024, +})); ``` -The table uses a SHA-256 path identifier as its primary key and stores the canonical path separately. This avoids requiring an arbitrary-length text primary key on MySQL while preserving the original virtual path. +`inspect()` includes native capabilities, effective support, hard limits known by the adapter, partition layout, resolved +optimization controls, the facade buffer ceiling, and a detached metrics snapshot. `plan()` is deterministic and does no I/O. +When size is known it can reject a request before work begins, show expected facade materialization, or explain the physical +part count selected by a partitioned adapter. Unknown provider limits remain unknown rather than being invented. -Use Drizzle as the backend --------------------------- +Write planning separates the resulting logical file from the bytes supplied by the current call. `size` checks logical +file/partition limits. `inputBytes` checks whether a non-native input stream fits under `maxBufferedWriteBytes`. Replace usually +needs only `size`; append/update should provide both values when they are known. -Drizzle deliberately keeps database dialects and schemas explicit. The package therefore does not invent one universal Drizzle table definition. The caller supplies a connected database and a dialect-correct table with the required columns. +Optimizations that select a materially different route are independently disableable: native stream read/write, direct range +read, native/server-side copy, and native move. The fallback is used only when it can preserve the portable filesystem contract. +For example, disabling provider-side copy can force bytes through this process and cannot reproduce provider-private control-plane +metadata such as every ACL, tag, lock policy, or checksum policy. Portable file bytes and `mediaType` are preserved. -```ts -import { sqliteTable, integer, text } from "drizzle-orm/sqlite-core"; -import { createFileSystem } from "@okikio/opfs"; -import { createDrizzleAdapter } from "@okikio/opfs/adapter/drizzle"; - -const files = sqliteTable("opfs_entries", { - path: text("path").primaryKey(), - parent: text("parent").notNull(), - name: text("name").notNull(), - kind: text("kind").notNull(), - data: text("data"), - size: integer("size").notNull(), - lastModified: integer("lastModified").notNull(), - mediaType: text("mediaType"), -}); +`metrics: "none"` removes facade counter updates for baseline benchmarks. `basic` counts operations, bytes, failures, route +selection, and peak facade materialization. `timing` adds monotonic durations. The direct S3 and Azure clients expose separate HTTP +request/retry metrics so protocol overhead and facade overhead can be measured independently. -const fileSystem = createFileSystem( - createDrizzleAdapter({ database, table: files }), - { coordination: "local" }, -); -``` +Large values are a backend capability, not a promise that every value store is unlimited. `adapter/deno-kv` is the first record +backend with a physical partition layout. Small files stay inline; large files use raw binary parts and a manifest-last commit. +Metadata lookup, directory listing, byte ranges, and stream reads do not reconstruct the complete logical file. The partition +policy is `never | auto | always`, so applications that do not want a changed durable layout can disable it explicitly. +Materialized append/update writes also build a new generation part-by-part, so the existing logical file is not joined into one +large base64 record before a small patch can be applied. -`path` must be unique. The current bridge replaces a record with delete-then-insert so it can stay on Drizzle's common CRUD surface. That replacement is serialized by this library inside one JavaScript realm. If multiple server processes can write the same path, the application must add database-level serialization or a transaction appropriate for its dialect. +`streamWriteModes` remains mode-specific. A simple record adapter can have no native stream lane, while Deno KV can advertise a +partitioned replacement stream and an object store can advertise native replacement streaming. Append and update can still be +emulated or unsupported independently. Deno KV's materialized append/update lane is direct, but streamed append/update remains +emulated because the incoming stream must first fit under the facade buffer ceiling. -Expose a filesystem as an unstorage driver ------------------------------------------- +Bridges make both integration directions explicit +------------------------------------------------- -This is the reverse adapter direction. +Adapters remain `ecosystem -> OPFS`. Drivers remain `OPFS -> ecosystem`. A bridge groups both directions and records an explicit +reason when one direction cannot honestly exist. ```ts -import { createStorage } from "unstorage"; -import { createFileSystem } from "@okikio/opfs"; -import { createNodeAdapter } from "@okikio/opfs/adapter/node"; -import { createUnstorageDriver } from "@okikio/opfs/driver/unstorage"; +import { UnstorageBridge, RxDbBridge } from "@okikio/opfs/bridge"; -const fileSystem = createFileSystem( - createNodeAdapter({ root: "./data" }), -); +console.log(UnstorageBridge.directions); +// { toOpfs: { supported: true }, fromOpfs: { supported: true } } -const storage = createStorage({ - driver: createUnstorageDriver(fileSystem), -}); - -await storage.setItem("cache:result", { ready: true }); +console.log(RxDbBridge.directions.fromOpfs); +// unsupported: a filesystem is not an RxStorage query/conflict/change-stream engine ``` -The driver reversibly maps unstorage's `:` key hierarchy to private virtual directories. Each logical key owns a dedicated `value` file inside its encoded key directory. This lets `foo` and `foo:bar` coexist even though a normal filesystem cannot make one path both a file and a directory. Characters such as `%`, `~`, `/`, spaces, and `?` are encoded so distinct keys do not collapse onto one filesystem entry. - -Streaming and record-backed adapters ------------------------------------- +The included bridge descriptors cover unstorage, RxDB, db0, Drizzle, and the generic reverse key-value view. Reverse KV and +unstorage drivers also expose the backing filesystem's `inspect()`, `plan()`, and `getMetrics()` methods so capability, size, +partition, optimization, and instrumentation decisions remain visible after the direction changes. Third parties can +use `defineBridge()` without a global registry. An unsupported direction must include a reason, which prevents a bridge from +silently pretending that asynchronous filesystem behavior can provide an unrelated synchronous or query-oriented contract. -Native filesystem adapters can stream without materializing the complete file: +Use schemas directly +-------------------- -| Adapter | stream read | stream write | range read | native move | sync random access | -| --- | --- | --- | --- | --- | --- | -| OPFS | yes | yes | yes | no portable native rename | DedicatedWorker when exposed | -| Deno | yes | yes | yes | yes | yes | -| Bun | yes | yes | yes | yes | yes | -| Node | yes | yes | yes | yes | yes | -| memory / record store | facade fallback | buffered | yes | no | no | -| unstorage | facade fallback | buffered | yes | no | no | -| RxDB | facade fallback | buffered | yes | no | no | -| db0 | facade fallback | buffered | yes | no | no | -| Drizzle | facade fallback | buffered | yes | no | no | - -Record-backed adapters use one validated record per file or directory. File data is base64 text. This is intentionally portable across JSON/document/SQL stores, but it increases byte storage by roughly one third before provider overhead. - -A streamed write to a non-streaming adapter is buffered by the facade. The default limit is 64 MiB: +Project-owned structural data is defined by Zod schemas and inferred TypeScript types. Schema constants end in `Schema`, and +project-owned serializable types normally end in `Type`. ```ts -const fileSystem = createFileSystem(adapter, { - maxBufferedWriteBytes: 16 * 1024 * 1024, -}); -``` - -If the stream crosses that limit, the producer is cancelled and the write fails with `FileSystemError` code `too-large`. The package does not silently consume unbounded memory. - -Paths and filesystem behavior ------------------------------ - -All backends receive the same canonical virtual paths: +import { PathSchema, type PathType } from "@okikio/opfs/schema"; -```text -input: projects/./kaiju/../state.json -result: /projects/state.json +const path: PathType = PathSchema.parse("/cache/result.bin"); ``` -The virtual root is `/`. Paths cannot escape above it. Backslashes and NUL characters are rejected so host-specific path rules do not leak into adapter behavior. +Zod 4 schemas implement Standard Schema, so consumers that accept Standard Schema can use these exported schemas directly. The +package does not maintain a parallel wrapper layer that could drift from the executable Zod contract. -`readDir()` and `walk()` are lazy async iterators. Recursive copy and directory clearing use bounded concurrency. Copy and move reject overlapping source and destination trees before an overwrite can destroy source data. +Test the semantics where they actually run +------------------------------------------ -When an adapter advertises `nativeMove`, the facade uses it. Otherwise `move()` is copy-then-remove and is explicitly non-atomic. +Portable filesystem contracts use `node:test` and `@std/expect`. Deno runs the same portable test source, Node runs the same +source, and Bun runs the same `node:test` API through its compatibility layer. Runtime-specific suites then prove the real host +filesystem adapters. -Coordination and ownership --------------------------- +Playwright Test owns the browser matrix. The same tests run in Chromium, Firefox, and WebKit and exercise Window OPFS, +DedicatedWorker, SharedWorker, ServiceWorker, same-origin and cross-origin iframes, opaque sandbox behavior, fresh-context +isolation, persistent-profile reopen, cancellation, and browser storage adapters. Tests probe capabilities instead of selecting +behavior from browser names. -The default coordination mode is `auto`: +Mitata benchmarks compare three layers where possible: ```text -file mutation - | - +-- shared tree lock - +-- exclusive path lock - -structural mutation - | - +-- exclusive tree lock +raw backend API + | + v +adapter primitive + | + v +FileSystemType facade + | + +-- coordination: none + `-- coordination: local ``` -`auto` uses Web Locks when available and otherwise falls back to one-realm FIFO locks. `web-locks` requires Web Locks. `local` forces the one-realm implementation. `none` means another subsystem owns coordination. - -Synchronous file resources keep the facade path lock for the full native file lifetime. Closing the sync file releases both the native resource and the facade lock. +Browser benchmarks compare raw native APIs, direct adapters, and the facade for OPFS, localStorage, IndexedDB, and Cache +Storage in Chromium, Firefox, and WebKit. Node, Deno, and Bun benchmarks compare their raw filesystem APIs with direct adapters and the facade. Bun additionally compares +Node-compatible `copyFile` with `Bun.write(destination, Bun.file(source))` rather than assuming one host copy path is faster. +Deno KV and SQLite have the same raw-to-adapter-to-facade measurements. -Injected adapters, databases, collections, and storage objects are borrowed by default. Ownership changes only when an option explicitly says so, such as `disposeAdapter`, `disposeStorage`, `disposeDatabase`, or `disposeFileSystem`. +Provider benchmarks use the same local provider fixture but keep each layer separate: official AWS/Azure SDK baseline, direct +protocol client, direct object adapter, facade with metrics disabled, and facade with basic metrics. A Bun-native S3 run compares +Bun's Rust-backed `S3Client` against the same project layers. Multipart/block cases are separate from single-request writes so a +different request plan is never presented as abstraction overhead. -Errors ------- +With the pinned mise toolchain: -Public operations use `FileSystemError` with a stable `code`, `operation`, optional `path`, and original `cause`. - -```text -unavailable -not-found -already-exists -type-mismatch -invalid-path -invalid-operation -not-supported -locked -quota-exceeded -permission-denied -aborted -too-large -unknown +```sh +mise install +mise run check +mise run test +mise run bench +mise run test-browser +mise run bench-browser +mise run bench-providers ``` -The error mapper understands browser exception names and Node-style error codes such as `ENOENT` and `EEXIST`. - -Browser execution contexts --------------------------- - -Native OPFS support is capability-based, not browser-name-based. `probeOpfs()` reports what the current context can actually do. It does not attempt to classify private/incognito mode. +GitHub Actions uses the same tool declarations and mise tasks. The workflow installs mise once per job, asks mise to install +only the runtimes that job needs, and then calls `mise run ...`. Runtime matrix jobs override one configured version with +`MISE__VERSION`; the Node matrix uses this to test Node 22, 24, and 26 without introducing a second tool-version +source. Third-party actions are pinned to immutable commit SHAs, and the mise binary version is pinned separately. -The async OPFS adapter is usable where the browser exposes `navigator.storage.getDirectory()`. Synchronous OPFS access is only used when the current file handle exposes `createSyncAccessHandle()`. - -Third-party iframe storage is opened normally through the iframe's current storage key. A separate `@okikio/opfs/iframe` entrypoint exposes the explicit Storage Access API request for browsers that support unpartitioned OPFS access. The package never requests that permission automatically. - -Service-worker code must still attach filesystem work to `event.waitUntil()` because storage I/O does not extend the service-worker event lifetime by itself. - -Public entrypoints ------------------- - -```text -@okikio/opfs -@okikio/opfs/adapter -@okikio/opfs/adapter/opfs -@okikio/opfs/adapter/memory -@okikio/opfs/adapter/deno -@okikio/opfs/adapter/node -@okikio/opfs/adapter/bun -@okikio/opfs/adapter/record -@okikio/opfs/adapter/unstorage -@okikio/opfs/adapter/rxdb -@okikio/opfs/adapter/db0 -@okikio/opfs/adapter/drizzle -@okikio/opfs/driver/unstorage -@okikio/opfs/iframe -@okikio/opfs/path -@okikio/opfs/schema -``` +The focused Deno tasks are documented in [docs/validation.md](./docs/validation.md). -Read next ---------- +Read the rest by the question you have +-------------------------------------- -- [`docs/api.md`](docs/api.md) explains every public API family and resource contract. -- [`docs/design.md`](docs/design.md) explains the adapter architecture, invariants, record model, coordination, and failure semantics. -- [`docs/adapters.md`](docs/adapters.md) is the implementation guide for native, record, and custom adapters. -- [`docs/ecosystems.md`](docs/ecosystems.md) explains RxDB, unstorage, db0, Drizzle, and their upstream integration points. -- [`docs/environments.md`](docs/environments.md) covers browser contexts plus Deno, Bun, and Node. -- [`docs/validation.md`](docs/validation.md) records what is tested, how it is tested, and what this host cannot verify. -- [`docs/sources.md`](docs/sources.md) records the standards, upstream source, Kaiju, Mediad, and research inputs used for this design. +- [Public API](./docs/api.md) explains the developer-facing filesystem and handle contracts. +- [Adapters](./docs/adapters.md) explains every first-party backend and the contracts for custom storage. +- [Architecture](./docs/design.md) explains invariants, streaming, copy/move, locks, ownership, and failure behavior. +- [Ecosystems](./docs/ecosystems.md) explains unstorage, RxDB, db0, Drizzle, S3-compatible services, and reverse drivers. +- [Environments](./docs/environments.md) explains Window, workers, iframes, Deno, Bun, Node, and provider clients. +- [Validation](./docs/validation.md) defines the canonical test and benchmark matrix. +- [Sources](./docs/sources.md) records the standards and upstream contracts that the implementation follows. +- [Releasing](./docs/releasing.md) explains JSR/npm packaging and release checks. diff --git a/docs/adapters.md b/docs/adapters.md index 5f5b6c2..70701ce 100644 --- a/docs/adapters.md +++ b/docs/adapters.md @@ -1,9 +1,46 @@ Adapter guide ============= -An adapter translates canonical virtual filesystem operations into one backend. This document describes the included adapters and the contract for new ones. +An adapter translates the package's canonical virtual filesystem operations into one concrete backend. The filesystem facade +owns filesystem semantics. The adapter owns backend mechanics. -Use `createFileSystem()` for every adapter: +That distinction lets the required backend contract stay small: + +```text +stat read one entry's metadata +readFile read one file or range +writeFile commit one materialized write +readDir lazily list direct children +createDir create exactly one directory +remove remove one file or empty directory +``` + +The facade builds parent creation, recursive walking, recursive copy/remove, OPFS-shaped handles, write-command staging, +coordination, and normalized errors on top. A backend can add native operations when it can do better than the facade fallback. + +```text +openReadStream native streaming read +writeStream native streaming for declared write modes +copy native/server-side file copy +move native rename/move +openWritableFile long-lived asynchronous positional writes +openSyncFile synchronous random access +``` + +`AdapterCapabilitiesType` must describe these native paths truthfully. `streamWriteModes` is a list rather than one boolean +because replacement, append, and update can have different backend costs. `nativeCopy` is separate from `nativeMove` because +object stores often copy efficiently but cannot rename an object atomically. + +Adapters can also expose `limits` and `partition`. Limits are hard facts known by the configured backend, such as maximum file, +value, key, part, batch, or concurrency sizes. Missing fields mean unknown, not unlimited. Partition describes a durable physical +layout used when one logical file spans multiple provider values. These fields feed `FileSystemType.inspect()` and `plan()` but +do not change the required adapter method set. + +Route-changing optimizations live on the facade, not inside capability flags. `optimizations.streamRead`, `streamWrite`, +`rangeRead`, `nativeCopy`, and `nativeMove` can force the safe fallback for differential testing or application policy. An adapter +should therefore implement the best native route it can and let the caller decide whether to use it. + +Use `createFileSystem()` to put the public API over any adapter: ```ts import { createFileSystem } from "@okikio/opfs"; @@ -14,29 +51,46 @@ const fileSystem = createFileSystem(adapter, { }); ``` -Included adapters ------------------ +The first-party adapters cover three different storage shapes +------------------------------------------------------------- -| Public subpath | Backend | Main use | +Native filesystems expose files and directories directly. Record stores expose values keyed by logical identity. Object stores +expose whole-object replacement, ranges, prefixes, and provider-side copy. Keeping those shapes separate is what prevents one +"universal" adapter from hiding important performance and consistency behavior. + +| Public subpath | Backend | Translation layer | | --- | --- | --- | -| `adapter/opfs` | browser OPFS root | native browser persistence | -| `adapter/deno` | `Deno.*` file APIs | Deno services and CLIs | -| `adapter/bun` | Bun + Bun's Node-compatible fs APIs | Bun services and CLIs | -| `adapter/node` | `node:fs` | Node services, Electron main process | -| `adapter/memory` | in-memory record map | tests, examples, temporary state | -| `adapter/record` | generic `RecordStoreType` | build a new value/document/SQL adapter | -| `adapter/unstorage` | unstorage `Storage` | use any compatible unstorage mount as filesystem persistence | -| `adapter/rxdb` | RxDB `RxCollection` | use RxDB and its selected RxStorage | -| `adapter/db0` | db0 `Database` | use db0 connector/dialect infrastructure | -| `adapter/drizzle` | Drizzle database + table | use an existing Drizzle schema/driver | - -### OPFS +| `adapter/opfs` | browser OPFS | native filesystem | +| `adapter/node` | Node `fs` | native filesystem | +| `adapter/deno` | Deno filesystem | native filesystem | +| `adapter/bun` | Bun + Node-compatible fs | native filesystem | +| `adapter/memory` | in-memory map | records | +| `adapter/record` | custom value/document store | records | +| `adapter/localstorage` | Web Storage | records | +| `adapter/indexeddb` | IndexedDB | records | +| `adapter/cache` | Cache Storage | records | +| `adapter/deno-kv` | Deno KV | records | +| `adapter/sqlite` | connected SQLite | records through db0-compatible SQL | +| `adapter/unstorage` | unstorage `Storage` | records | +| `adapter/rxdb` | RxDB `RxCollection` | records | +| `adapter/db0` | db0 `Database` | records | +| `adapter/drizzle` | Drizzle database + table | records | +| `adapter/object` | custom object store | objects | +| `adapter/s3` | direct S3/S3-compatible client | objects | +| `adapter/azure` | direct Azure Blob client | objects | + +Native browser and host filesystems +----------------------------------- + +`openFileSystem()` is the convenience path for native browser OPFS: ```ts import { openFileSystem } from "@okikio/opfs"; + +const fileSystem = await openFileSystem(); ``` -or explicitly: +The explicit form is useful when the caller already owns the native root: ```ts import { createFileSystem } from "@okikio/opfs"; @@ -46,198 +100,317 @@ const root = await navigator.storage.getDirectory(); const fileSystem = createFileSystem(createOpfsAdapter(root)); ``` -The explicit adapter retains `nativeRoot` for advanced browser interop. It exposes long-lived positional writes through one `FileSystemWritableFileStream`. Synchronous access is exposed only when the current native file handle actually provides `createSyncAccessHandle()`. - -The adapter does not attempt browser or incognito detection. +The OPFS adapter reports synchronous access only when an actual file handle exposes `createSyncAccessHandle()`. The package +does not infer the feature from a browser name or from "worker" alone. -### Deno +Node, Deno, and Bun map virtual `/` below one configured host directory: ```ts -import { createFileSystem } from "@okikio/opfs"; -import { createDenoAdapter } from "@okikio/opfs/adapter/deno"; +import { createNodeAdapter } from "@okikio/opfs/adapter/node"; -const fileSystem = createFileSystem( - createDenoAdapter({ root: "./data" }), - { coordination: "local" }, -); +const adapter = createNodeAdapter({ root: "./data" }); ``` -`root` is the host directory represented by virtual `/`. `createRoot` defaults to true. +The host path mapper resolves the configured root once and rejects every virtual path whose resolved host path would leave that +root. Host adapters expose ranges, streams, native copy, native move, and synchronous random access when the underlying runtime +provides them. -The adapter uses Deno filesystem APIs for data operations, rename, long-lived positional access, sync access, and flush. It uses Node's path compatibility module only to normalize the configured host root and to verify that a virtual path stays below it. +Bun uses Bun's file APIs where they provide a direct read/write path and uses Bun's Node-compatible filesystem surface for the +operations whose exact semantics already live there. Importing the Bun adapter does not require the `Bun` global until adapter +creation. -### Bun +Record stores start small and can add byte lanes +----------------------------------------------- + +The required `RecordStoreType` stays intentionally small: ```ts -import { createBunAdapter } from "@okikio/opfs/adapter/bun"; +interface RecordStoreType { + get(path): Promise; + set(record): Promise; + delete(path): Promise; + list(parent): AsyncIterableIterator; +} +``` + +This complete-record path is enough for memory, Web Storage, RxDB, unstorage, and SQL-backed integrations. File records use +base64 because the same durable shape must round-trip through JSON-oriented stores. Base64 is a compatibility format, not a claim +that every record backend is suitable for large binaries. -const adapter = createBunAdapter({ root: "./data" }); +A store with a more capable physical layout can add optional lanes without implementing the filesystem facade again: + +```text +stat metadata without file body +readFile direct/range byte read +openReadStream backpressure-preserving logical stream +writeFile selected direct materialized modes +writeStream selected direct stream modes ``` -The replace/read fast path uses Bun file APIs. Directory operations, update-mode writes, native rename, and synchronous random access use Bun's Node-compatible filesystem APIs. +`RecordStoreCapabilitiesType` declares `rangeRead`, `streamRead`, `writeModes`, and `streamWriteModes`. The record adapter turns +only those declared lanes into native adapter capabilities. If a lane is absent, the complete-record implementation remains the +fallback. This is the extension point for third-party KV/document stores that can do better than one large JSON-shaped record. -The module resolves `Bun` lazily during adapter creation. Importing the module does not require the Bun global to exist. +`createMemoryAdapter()` uses the complete-record path for deterministic tests and temporary data. -### Node +`createLocalStorageAdapter(storage)` accepts an injected Web Storage object. Web Storage is synchronous, quota-limited, and +string-only underneath the adapter. It does not claim a portable maximum item size or native streaming. -```ts -import { createNodeAdapter } from "@okikio/opfs/adapter/node"; +`openIndexedDbAdapter()` can open its own IndexedDB database, while `createIndexedDbAdapter(database)` can borrow an existing +one. The store is keyed by canonical `path` and indexed by `parent` so direct directory listing stays indexed. Ownership remains +with the caller unless the adapter option explicitly transfers it. -const adapter = createNodeAdapter({ root: "./data" }); -``` +`createCacheAdapter(cache)` stores records in an injected `Cache` using synthetic request URLs. No network request is made. Cache +Storage quota, eviction, and persistence policy remain browser decisions. -Node supports native streaming reads/writes, byte ranges, rename, long-lived asynchronous positional writes, and synchronous random access. +Deno KV uses an explicit partition layout +----------------------------------------- -The host root is created by default. The virtual path mapper rejects any resolved host path that would leave that root. +`createDenoKvAdapter(kv)` accepts an already-open Deno KV database. Deno KV has a 2 KiB serialized key limit and a 64 KiB +serialized value limit, so treating one filesystem file as one KV value would create a small and surprising file ceiling. The +default adapter policy is `partition: "auto"`. -### Memory +```text +logical entry key + [prefix, "entry", parentPath, name] -```ts -import { createMemoryAdapter } from "@okikio/opfs/adapter/memory"; +list one parent + prefix [prefix, "entry", parentPath] + -> direct children only + +small file + entry -> normal FileRecord + +large file + [prefix, "part", canonicalPath, generation, 0] + [prefix, "part", canonicalPath, generation, 1] + ... + entry -> manifest committed last ``` -The memory adapter uses the same record-store layer as database adapters. It is intentionally deterministic and dependency-free. It is suitable for tests and temporary state, not durable storage. +Default decoded sizes are 32 KiB inline and 48 KiB per raw binary part. `maxParts` defaults to 10,000 and part I/O concurrency +to 8. Callers can set `partition: "never" | "auto" | "always"`, `inlineBytes`, `partBytes`, `maxParts`, and `concurrency`. The +adapter exposes these as inspectable limits/partition policy. -The companion `createMemoryRecordStore()` is useful when testing another record-store wrapper directly. +Manifest-last publication is the visibility rule. A reader sees the previous complete generation until all new parts exist and +the new manifest is stored. A process crash before the manifest commit can leave unreachable new-generation parts. That is a +storage leak, not a partially visible logical file. The adapter does not currently run a global orphan scavenger because doing so +would require a separate ownership/retention policy. -RecordStoreType ---------------- +Large-file hot paths avoid generic reconstruction: -Use `RecordStoreType` when the backend is fundamentally value-based instead of filesystem-based. +- `stat()` reads the entry/manifest only. +- `list()` uses the direct-parent key prefix, so it reads direct-child metadata only and never scans descendant entry keys or body parts. +- range reads fetch only overlapping parts. +- stream reads load one physical part at a time under consumer backpressure. +- materialized append/update builds a new generation part-by-part and never joins the previous large file into one record. +- streamed replacement writes parts with bounded concurrency and publishes the manifest last. -```ts -import { - createRecordAdapter, - type RecordStoreType, -} from "@okikio/opfs/adapter/record"; - -const store: RecordStoreType = { - async get(path) { /* ... */ }, - async set(record) { /* ... */ }, - async delete(path) { /* ... */ }, - async *list(parent) { /* direct children only */ }, -}; - -const adapter = createRecordAdapter(store, { - name: "my-store", -}); -``` +Append/update still copy the untouched logical bytes into a new immutable generation because Deno KV has no provider-side range +copy primitive. The copy is bounded by `partBytes` and `concurrency`; the tradeoff is provider I/O proportional to the resulting +file size rather than JavaScript memory proportional to that size. Streamed append/update is not advertised as a direct lane, +so an incoming stream must still fit under the facade `maxBufferedWriteBytes` ceiling before this bounded patch path runs. -Important contract rules: +`partition: "never"` disables the partitioned streaming write lane. A large streamed write then follows the facade's normal +bounded materialization rule and fails `too-large` once it exceeds `maxBufferedWriteBytes`. This gives applications a deliberate +way to reject the changed durable layout. Deno KV remains an unstable Deno API, so real integration tests run with +`--unstable-kv`. -- `get()` returns one validated logical path. -- `set()` replaces one logical record as atomically as the provider permits. -- `delete()` removes one record only. Recursive behavior belongs to the filesystem facade. -- `list(parent)` yields direct children only. -- The store receives canonical paths. -- The store can expose `dispose()` for resources it explicitly owns. -- Use `readOnly: true` when the storage can read but cannot mutate. +`createSqliteAdapter(database)` accepts a small connected SQLite statement interface. It deliberately reuses the same SQL record +mapping used by the SQLite branch of `createDb0Adapter()` instead of creating a second schema and upsert implementation. -Record adapters do not claim native streaming. Their stream inputs are materialized under `maxBufferedWriteBytes`. +Existing ecosystem adapters stay above the abstraction the application already owns: -Custom AdapterType ------------------- +```text +unstorage Storage -> RecordStoreType +RxDB RxCollection -> RecordStoreType +db0 Database -> RecordStoreType +Drizzle DB + table -> RecordStoreType +``` -Use `AdapterType` directly when the backend can expose real file-like primitives. +Bridge descriptors group these forward adapters with reverse drivers when a real reverse contract exists. They do not fabricate +a reverse direction for RxDB, db0, or Drizzle. See [ecosystems.md](./ecosystems.md). -```ts -import { - defineAdapter, - type AdapterType, -} from "@okikio/opfs/adapter"; +Object stores keep object-store semantics visible +------------------------------------------------- -export const adapter = defineAdapter({ - name: "provider", - capabilities: { - read: true, - write: true, - streamRead: false, - streamWrite: false, - rangeRead: true, - nativeMove: false, - positionalWrite: false, - syncAccess: false, - }, +`ObjectStoreType` is the common client contract for S3, Azure Blob, and custom object storage. It models the operations those +systems actually have: - async stat(path, options) { /* ... */ }, - async readFile(path, options) { /* ... */ }, - async writeFile(path, bytes, options) { /* ... */ }, - async *readDir(path, options) { /* direct children */ }, - async createDir(path, options) { /* parent already exists */ }, - async remove(path, options) { /* file or empty directory */ }, -}); +```text +HEAD exact key +GET full object or range +PUT replacement +DELETE exact key +LIST prefix + delimiter +COPY inside provider, when supported +``` + +Its metadata includes byte size, media type, last modification time, ETag, provider version identity, and user metadata. Its +capabilities state whether range read, streaming read, streaming replacement, provider-side copy, and conditional writes are +really available. + +`createObjectAdapter()` maps that model to filesystem paths. A normal file maps to one object key. An empty directory maps to a +trailing-slash marker with private metadata, and prefix listing recognizes both those markers and foreign provider prefixes. + +```text +/photos -> photos/ +/photos/a.jpg -> photos/a.jpg +/photos/2026/b.jpg -> photos/2026/b.jpg ``` -`defineAdapter()` validates the adapter name and capability object at runtime. It does not register the adapter globally. +A raw object namespace can physically contain both `mixed` and `mixed/child`. The filesystem view resolves the exact `mixed` +object as the file because `stat("/mixed")` does the same. That rule keeps read, stat, and write behavior internally consistent +when foreign object layouts do not obey filesystem restrictions. -Required adapter semantics --------------------------- +Replacement can stream when the provider supports it. Append and update cannot normally mutate object bytes in place, so the +adapter performs: -`stat` -: Return `null` only for not-found. A file stat includes size, last-modified milliseconds, and a media type string. +```text +HEAD current object + | + v +GET current bytes + | + v +apply append/update in memory + | + v +conditional PUT replacement +``` -`readFile` -: Respect optional byte offset and maximum length. Return the bytes actually read. +When conditional writes are enabled and an existing object does not return an ETag, the adapter fails rather than quietly +performing an unsafe read-modify-write. When the file is being created through append/update, it uses create-only semantics where +the provider exposes them. -`writeFile` -: Preserve `replace`, `append`, and `update` semantics. Respect `truncate` at the final write cursor. +The S3 client implements the protocol directly +----------------------------------------------- -`readDir` -: Yield direct child names and kinds lazily. Do not recursively traverse here. +`createS3Client()` uses Web Fetch, Web Crypto, `@std/encoding`, and `@std/xml`. It does not depend on the AWS SDK. -`createDir` -: Create exactly one directory. The parent already exists when the facade calls this primitive. +```ts +import { createS3Client } from "@okikio/opfs/s3"; +import { createS3Adapter } from "@okikio/opfs/adapter/s3"; + +const client = createS3Client({ + endpoint: "https://s3.us-east-1.amazonaws.com", + bucket: "example", + region: "us-east-1", + credentials: async () => await credentials.get(), + addressing: "path", + concurrency: 4, +}); -`remove` -: Remove one file or one empty directory. Recursive behavior belongs to the facade. +const adapter = createS3Adapter(client, { prefix: "app" }); +``` -Optional native operations --------------------------- +Signature Version 4 includes the request authority in canonical headers. Browser Fetch forbids application code from setting +the `Host` header, so the client signs `url.host` while leaving actual Host or `:authority` transmission to Fetch. Static browser +credentials are usually a security mistake; browser deployments should use appropriately scoped short-lived credentials or a +trusted service design. -Only advertise a capability when the adapter implements the corresponding native method. +A streamed replacement uses multipart upload with bounded part concurrency: ```text -streamRead -> openReadStream -streamWrite -> writeStream -nativeMove -> move -positionalWrite -> openWritableFile -syncAccess -> openSyncFile +ReadableStream + | + v +fixed-size chunker + | + +--> UploadPart 1 --+ + +--> UploadPart 2 --+--> CompleteMultipartUpload + +--> UploadPart N --+ + | + failed operation + | + `--> wait active parts -> AbortMultipartUpload ``` -`positionalWrite` means the adapter can keep one asynchronous file resource open while callers write explicit byte positions. It is not inferred from ordinary `writeFile(update)` support. This distinction matters for media muxers and database files because repeated whole-record replacement can turn many chunk writes into nonlinear work. +The client applies `If-Match` and `If-None-Match` to multipart completion, which is the commit operation that current S3 exposes +for these preconditions. It waits for already-started parts before aborting so a late part cannot arrive after the abort request. -`rangeRead` describes whether the adapter can avoid materializing the complete file for a range. The facade still exposes ranged `readFile()` to all adapters. +S3 has two failure cases that are easy to miss. `CompleteMultipartUpload` can return HTTP 200 and then stream an XML error, and +`CopyObject` can also return an embedded error in HTTP 200. The client parses and rejects both bodies. -Cancellation ------------- +`CopyObject` has a 5 GB source limit. Larger copies use `UploadPartCopy` ranges into a multipart destination. The copy part size +increases when necessary to stay within S3's 10,000-part limit. Source bytes stay inside the object provider rather than crossing +JavaScript memory or network twice. -Every async operation that accepts an `AbortSignal` should check it before expensive work and between long-running chunks. When a stream write fails or aborts, cancel the source producer when possible so upstream work does not continue after the file operation is terminal. +The filesystem copy contract intentionally preserves file bytes, media type, and user metadata where the client can do so. It +does not claim to clone every S3 control-plane property such as ACLs, tags, object-lock state, or every checksum policy. Use the +low-level signed `client.request()` API when the S3 object itself, rather than its filesystem view, is the thing being managed. -Errors ------- +S3-compatible does not mean behavior-identical. Configure the client from the selected provider's current contract: -Adapters can throw native errors. The facade maps known browser and server error shapes through `toFileSystemError()`. +- Cloudflare R2 commonly uses the `auto` region and has provider-specific supported/unsupported S3 operations. +- DigitalOcean Spaces implements a compatible subset rather than every AWS S3 feature. +- Google Cloud Storage's XML multipart API documents different precondition behavior; disable `conditionalWrite` when the + selected path does not provide the safety contract expected by the object adapter. +- Other compatible providers should be treated the same way: verify endpoint, signing region, addressing, copy, conditional + requests, multipart limits, checksums, and error behavior before enabling a capability flag. -If an adapter itself must create a package error, use `FileSystemError` with a precise operation and canonical path. Do not return a boolean for exceptional filesystem states. +The Azure Blob client keeps Azure's own model +-------------------------------------------- -Ownership ---------- +`createAzureClient()` also uses Web Fetch and `@std/xml`, but it does not force Azure Blob through an S3-shaped client. -Adapters borrow injected resources unless their options explicitly transfer ownership. +```ts +import { createAzureClient } from "@okikio/opfs/azure"; +import { createAzureAdapter } from "@okikio/opfs/adapter/azure"; -Good: +const client = createAzureClient({ + endpoint: "https://account.blob.core.windows.net", + container: "example", + credential: { kind: "sas", token }, +}); -```ts -createUnstorageAdapter(storage, { disposeStorage: true }); -createDb0Adapter(database, { disposeDatabase: true }); -createFileSystem(adapter, { disposeAdapter: true }); +const adapter = createAzureAdapter(client, { prefix: "app" }); ``` -Avoid an adapter that always disposes a resource supplied by the caller. +Credentials can be SAS, refreshable bearer tokens, or a custom header callback. The service version is explicit and defaults to +the version pinned by this package. Block-size limits are selected from that service version rather than one timeless constant. + +Streamed replacements use Put Block followed by Put Block List. Azure has no abort call for uncommitted blocks, so a failure +waits for already-started requests, leaves the old committed blob untouched, and allows Azure to garbage-collect the uncommitted +blocks later. -Import safety -------------- +Copy Blob From URL has a smaller synchronous copy limit. Larger files use Put Block From URL ranges followed by Put Block List. +This keeps large copies server-side without hiding a size cliff behind `nativeCopy: true`. + +Custom adapters must preserve the same invariants +------------------------------------------------- + +Use `defineAdapter()` for a backend that already exposes filesystem-like primitives: + +```ts +import { defineAdapter } from "@okikio/opfs/adapter"; + +export const adapter = defineAdapter({ + name: "provider", + capabilities: { + read: true, + write: true, + streamRead: false, + streamWriteModes: [], + rangeRead: false, + nativeCopy: false, + nativeMove: false, + positionalWrite: false, + syncAccess: false, + }, + async stat(path) { /* ... */ }, + async readFile(path, options) { /* ... */ }, + async writeFile(path, bytes, options) { /* ... */ }, + async *readDir(path, options) { /* direct children only */ }, + async createDir(path, options) { /* parent already exists */ }, + async remove(path, options) { /* file or empty directory */ }, +}); +``` -A concrete adapter subpath can depend on its runtime, but importing unrelated entrypoints must not pull that runtime into the graph. +Every adapter receives canonical virtual paths. It must respect requested ranges and write modes. It must yield direct children, +not recursive descendants, from `readDir()`. It must not configure logging, inspect process environment, or open unrelated +resources during module evaluation. -Do not export server adapters from the root module. Do not probe environment variables, configure logs, connect to providers, or start workers at module evaluation time. +Injected resources are borrowed by default. Transfer ownership only through an explicit option such as `disposeDatabase`, +`disposeStore`, or `disposeAdapter`. A filesystem closing must never surprise another subsystem by disposing infrastructure that +it still owns. diff --git a/docs/api.md b/docs/api.md index 80c35a1..62a4832 100644 --- a/docs/api.md +++ b/docs/api.md @@ -27,6 +27,8 @@ const fileSystem = createFileSystem(adapter, { coordination: "auto", lockPrefix: "my-app:filesystem", maxBufferedWriteBytes: 64 * 1024 * 1024, + metrics: "basic", + optimizations: { nativeCopy: true }, disposeAdapter: false, }); ``` @@ -35,11 +37,68 @@ const fileSystem = createFileSystem(adapter, { - `coordination`: `auto`, `web-locks`, `local`, or `none`. - `lockPrefix`: stable lock namespace used for cooperating filesystem facades. -- `maxBufferedWriteBytes`: maximum stream size materialized for non-streaming adapters. +- `maxBufferedWriteBytes`: maximum stream/file size the facade may materialize for a fallback route. +- `metrics`: `none`, `basic`, or `timing`. `none` is intended for baseline overhead measurements. +- `optimizations`: partial override for `streamRead`, `streamWrite`, `rangeRead`, `nativeCopy`, and `nativeMove`. Every route defaults to enabled. - `disposeAdapter`: transfers adapter disposal ownership to the facade when true. `coordination` is runtime-validated by `CoordinationModeSchema`. + +Inspect and plan before I/O +--------------------------- + +### `inspect()` + +Returns a synchronous `InspectionType` for the configured filesystem stack. It contains: + +```text +adapter +native capabilities +effective support +known hard limits +optional partition layout +resolved optimization policy +maxBufferedWriteBytes +metrics mode and current metrics snapshot +``` + +Effective support uses `native | emulated | partitioned | unsupported`. Native capability flags never include facade emulation. +This makes a caller able to distinguish a fast provider range read from a full-read-and-slice fallback, or a Deno KV partitioned +stream from a complete-record materialization. + +### `plan(input)` + +Creates a deterministic `PlanType` without touching the backend. Supported operations are `read`, `write`, `copy`, and `move`. +For writes, `size` is the resulting logical file size and is used for `maxFileBytes` and partition-count checks. `inputBytes` is +the number of bytes supplied by the current write and is used for facade stream-buffer admission. For replace, omitted +`inputBytes` falls back to `size` because those values are normally equal. Append/update callers should provide both when known. + +```ts +const plan = fileSystem.plan({ + operation: "write", + source: "stream", + mode: "replace", + size: 200 * 1024 * 1024, + inputBytes: 200 * 1024 * 1024, +}); + +if (!plan.supported) throw new Error(plan.reasons.join(" ")); +``` + +An unknown limit remains unknown. The planner does not invent provider guarantees. For an unknown-size emulated stream, it warns +that `maxBufferedWriteBytes` remains the runtime admission limit. + +### `getMetrics()` + +Returns a detached `MetricsType`. `basic` tracks counts, failures, bytes, native/emulated/partitioned route counts, current +facade-owned buffered bytes, and peak buffered bytes. `timing` additionally records total and maximum durations. `none` keeps the +hot-path collector inactive. + +Optimization controls can force a fallback for differential tests or application policy. A disabled optimization is never +relabelled native. If the fallback cannot satisfy the request within `maxBufferedWriteBytes`, runtime execution and `plan()` both +fail with the same `too-large`/unsupported condition when size is known. + Path API -------- @@ -333,36 +392,127 @@ await writable.write({ type: "truncate", size: 100 }); Blob also has a `type` property, so the implementation identifies a command only when `type` is exactly `write`, `seek`, or `truncate`. -Adapter API ------------ +Adapter and storage API +----------------------- + +`@okikio/opfs/adapter` exports the backend contract consumed by `createFileSystem()`. The required primitive set stays small: + +```text +stat +readFile +writeFile +readDir +createDir +remove +``` + +Optional methods expose stronger native paths: + +```text +openReadStream -> capabilities.streamRead +writeStream -> mode is present in capabilities.streamWriteModes +copy -> capabilities.nativeCopy +move -> capabilities.nativeMove +openWritableFile -> capabilities.positionalWrite +openSyncFile -> capabilities.syncAccess +``` + +The exported contract includes `AdapterCopyOptionsType` as well as the signal, read, write, move, stat, directory-entry, +writable-file, sync-file, adapter, and filesystem option types. `defineAdapter()` validates the adapter name and capability +record. It does not add a global registry or change the adapter. -`@okikio/opfs/adapter` exports the complete backend contract: +The capability record describes what the backend performs natively. For example, an object adapter can expose +`streamWriteModes: ["replace"]` because replacement can stream to multipart/block upload while append and update still need a +read-modify-write cycle. The facade can emulate operations, but it does not relabel an emulation as native support. -- `AdapterSignalOptionsType` -- `AdapterReadOptionsType` -- `AdapterWriteOptionsType` -- `AdapterMoveOptionsType` -- `AdapterDirectoryEntryType` -- `AdapterFileStatType` -- `AdapterDirectoryStatType` -- `AdapterStatType` -- `AdapterSyncFileType` -- `AdapterType` -- `FileSystemOptionsType` -- `defineAdapter()` +### Record storage -`defineAdapter()` validates `AdapterNameSchema` and `AdapterCapabilitiesSchema` without adding a registry or global mutation. Adapter methods always receive canonical virtual paths. See [adapters.md](./adapters.md) for every primitive and first-party adapter. +`@okikio/opfs/adapter/record` is the common translation point for backends that naturally store values, documents, or SQL rows. +It exports `RecordStoreType`, `RecordAdapterOptionsType`, and `createRecordAdapter()`. -Record API ----------- +A record store always supports the complete logical `get/set/delete/list` contract. Simple stores can stop there. More capable +value stores can additionally expose `stat`, direct range `readFile`, `openReadStream`, selected direct materialized write modes, +and selected `writeStream` modes through `RecordStoreCapabilitiesType`. `createRecordAdapter()` translates only the declared +lanes into native adapter capabilities. -`@okikio/opfs/adapter/record` exports: +The portable complete-record fallback uses the versioned `RecordType` union and base64 file bytes so JSON, Web Storage, RxDB, +unstorage, and SQL text columns share one representation. A specialized store such as Deno KV can keep large body parts as raw +binary values and use metadata-only/range/stream/direct-write lanes so the generic base64 representation is not on its +large-file hot path. Deno KV declares materialized replace/append/update as direct store modes; append/update rebuild the next +immutable generation part-by-part instead of reconstructing the previous complete logical file. Its stream lane remains +replace-only, so streamed append/update can still require facade input buffering. -- `RecordStoreType` -- `RecordAdapterOptionsType` -- `createRecordAdapter()` +### Object storage -`@okikio/opfs/schema` exports the validated persistence schemas: +`@okikio/opfs/adapter/object` exports the provider-neutral object-storage layer: + +- `ObjectCapabilitiesSchema` / `ObjectCapabilitiesType` +- `ObjectStatType` and `ObjectEntryType` +- object GET, PUT, COPY, and LIST option types +- `ObjectStoreType` +- `ObjectAdapterOptionsType` +- `createObjectAdapter()` + +`ObjectStoreType` preserves object concepts such as ETags, provider version IDs, user metadata, prefix listing, conditional +writes, and server-side copy. `createObjectAdapter()` then maps that model into files and directories. Empty directories use +trailing-slash marker objects, while ordinary prefix listing also recognizes directories created outside this library. + +The direct S3 API lives at `@okikio/opfs/s3`: + +```ts +import { createS3Client } from "@okikio/opfs/s3"; +import { createS3Adapter } from "@okikio/opfs/adapter/s3"; + +const client = createS3Client({ + endpoint, + bucket, + region, + credentials, +}); +const adapter = createS3Adapter(client); +``` + +`S3ClientType` extends `ObjectStoreType` and also exposes signed `request()`, `createUpload()`, `uploadPart()`, +`completeUpload()`, and `abortUpload()`. The lower-level request method is the deliberate escape hatch for provider-specific +S3 features that do not belong in the filesystem API. + +The Azure counterpart lives at `@okikio/opfs/azure` and `@okikio/opfs/adapter/azure`: + +```ts +import { createAzureClient } from "@okikio/opfs/azure"; +import { createAzureAdapter } from "@okikio/opfs/adapter/azure"; + +const client = createAzureClient({ endpoint, container, credential }); +const adapter = createAzureAdapter(client); +``` + +`AzureClientType` retains the REST request escape hatch, provider request IDs, range access, block upload, and server-side copy. + +The object-store interfaces are intentionally smaller than either provider protocol. See [S3 client protocol](./s3.md) and +[Azure Blob client protocol](./azure.md) for signing, version gates, multipart/block lifecycles, limits, provider failures, and +known unsupported operations. + +### Reverse key-value APIs + +`@okikio/opfs/driver/kv` exports `createKeyValueDriver()`. It maps colon-delimited keys onto private directories so both `foo` +and `foo:bar` can exist at the same time: + +```text +foo -> /key-foo/value +foo:bar -> /key-foo/key-bar/value +``` + +The driver supports string/raw get and set, existence, metadata, hierarchical key enumeration, clear, and explicit filesystem +ownership transfer. It also exposes `inspect()`, `plan()`, and `getMetrics()`. Those methods delegate to the backing +`FileSystemType`, so a reverse ecosystem consumer sees the same effective routes, limits, partition policy, buffer ceiling, and +observed metrics instead of receiving a second approximation of storage capability. + +`@okikio/opfs/driver/unstorage` is a thin translation over this generic driver. It supplies unstorage method names and +`maxDepth` behavior while reusing the same collision-safe filesystem mapping. + +### Schemas + +`@okikio/opfs/schema` exports the executable project data contracts: - `PathSchema` / `PathType` - `AdapterNameSchema` / `AdapterNameType` @@ -379,6 +529,23 @@ Record API - `Db0DialectSchema` / `Db0DialectType` - `SqlIdentifierSchema` / `SqlIdentifierType` +The package exports Zod 4 schemas directly. Zod 4 implements Standard Schema, so a consumer that accepts that interface can use +the same schema value without a second OPFS-owned wrapper contract. + +### Bridge descriptors + +`@okikio/opfs/bridge` groups the adapter and reverse-driver directions for an ecosystem. It does not replace either primitive. + +- `UnstorageBridge`: both directions. +- `RxDbBridge`: collection to OPFS only. +- `Db0Bridge`: database to OPFS only. +- `DrizzleBridge`: database/table to OPFS only. +- `KeyValueBridge`: OPFS to generic asynchronous key-value only. + +`defineBridge()` validates that `directions.toOpfs/fromOpfs` agree with the constructors. Every unsupported direction must state a +reason. This is intentional for ecosystems where the reverse shape would require query, conflict, transaction, synchronous, or +other semantics a filesystem does not own. + Path utility API ---------------- diff --git a/docs/azure.md b/docs/azure.md new file mode 100644 index 0000000..43d187c --- /dev/null +++ b/docs/azure.md @@ -0,0 +1,492 @@ +Azure Blob client protocol guide +================================ + +Purpose +------- + +This document defines the Azure Blob Storage REST contract implemented by +`@okikio/opfs/azure`. It is intended for maintainers changing authentication, +service-version behavior, block upload, copy, conditional replacement, listing, +or Azurite interoperability. + +The client uses the Blob REST API directly rather than wrapping the Azure SDK. +That keeps the dependency graph small, but it also means this repository owns +the protocol work it chooses to implement. The source and tests must therefore +make the exact REST contract explicit. + +```text +ObjectStoreType / AzureClientType + | + v +Blob REST request construction + | + +--> SAS query authorization + +--> Microsoft Entra bearer authorization + +--> Shared Key signing + `--> caller-defined headers + | + v +Web Fetch + | + +--> Azure Blob Storage + `--> Azurite +``` + +Web Crypto owns HMAC-SHA256 for Shared Key. `@std/encoding` owns Base64, +`@std/async/pool` owns bounded block concurrency, and `@std/xml` owns list and +block-list documents. + +This guide uses these evidence classes: + + - **Implemented** means current source contains the behavior. + - **Protocol** means current Microsoft REST documentation defines the behavior. + - **Emulator** means Azurite reproduces enough of the contract for local + integration tests but is not treated as complete Azure parity. + +The current implementation was reviewed against Microsoft Learn and current +Azurite documentation on August 14, 2026. + + +The client targets one container +-------------------------------- + +`createAzureClient()` binds one endpoint and one container. Blob keys supplied +to `head()`, `get()`, `put()`, `delete()`, `copy()`, and `list()` are relative to +that container. + +A normal cloud endpoint looks like: + +```text +https://account.blob.core.windows.net +``` + +The client constructs: + +```text +https://account.blob.core.windows.net/container/path/to/blob +``` + +Azurite uses an account name in the endpoint path: + +```text +http://127.0.0.1:10000/devstoreaccount1 +``` + +which becomes: + +```text +http://127.0.0.1:10000/devstoreaccount1/container/path/to/blob +``` + +This difference matters to Shared Key canonicalization. Microsoft documents +that the emulator account segment appears once in the URL path and is prefixed +again by the signing account name. The implementation derives that duplicated +canonical-resource form from the URL rather than hard-coding an Azurite branch. + +`AZURE_STORAGE_VERSION` defaults to `2026-04-06`. A caller can select another +service version when it needs the size or authentication behavior of an older +REST contract. + + +Authorization is explicit +------------------------- + +`AzureCredentialType` supports four strategies: + +| Kind | Wire mechanism | Intended use | +| ---- | -------------- | ------------ | +| `sas` | SAS fields remain in the request query | Browser/server delegated access | +| `bearer` | `Authorization: Bearer ...` | Microsoft Entra access token | +| `shared-key` | Canonical Shared Key HMAC-SHA256 | Trusted server and Azurite | +| `headers` | Caller returns authorization headers | Provider/host integration not otherwise modeled | + +The client does not read environment variables. Credentials are supplied by +the caller, and bearer tokens can be refresh functions resolved immediately +before the request. + +Shared Key credentials contain an account name and Base64 account key. The +account key is a root-level storage credential. It should not be embedded in an +untrusted browser bundle. Browser applications normally use a scoped SAS or a +Microsoft Entra flow with suitable permissions. + +### Shared Key string to sign + +The implementation follows the Blob service Shared Key format: + +```text +HTTP verb +Content-Encoding +Content-Language +Content-Length +Content-MD5 +Content-Type +Date +If-Modified-Since +If-Match +If-None-Match +If-Unmodified-Since +Range +CanonicalizedHeaders +CanonicalizedResource +``` + +The signer supports the augmented Blob Shared Key format from service version +`2009-09-19` onward. Earlier versions are rejected rather than being signed +with modern rules that only look plausible. + +Two canonicalization rules change with the selected service version: + + - `2014-02-14` and earlier sign a zero byte `Content-Length` as the literal + `0`. Later versions contribute an empty line for the same header. + - Versions before `2016-05-31` omit empty `x-ms-*` headers from + `CanonicalizedHeaders`. Version `2016-05-31` and later retain them as + `name:\n`. + +Canonical `x-ms-*` headers that participate in the selected version are: + +1. converted to lowercase names; +2. normalized by collapsing linear whitespace outside quoted strings while + preserving whitespace inside quoted strings; +3. sorted by code-unit order; +4. emitted as `name:value\n`. + +The quoted-string rule is significant for metadata and other extension headers. +For example, `alpha beta` canonicalizes to `alpha beta`, while the two spaces +inside `alpha "beta gamma"` remain two spaces. Collapsing the quoted value +would sign different bytes from the value Azure receives. + +The canonical resource starts with: + +```text +/account-name/request-path +``` + +then appends lowercase query names in sorted order. Repeated values are sorted +and joined with commas. + +The signature is: + +```text +Base64(HMAC-SHA256(base64DecodedAccountKey, UTF8(stringToSign))) +``` + +and the HTTP header is: + +```text +Authorization: SharedKey account-name:signature +``` + +The deterministic unit suite signs Azurite requests with the documented +`devstoreaccount1` key and a fixed timestamp. It freezes exact signatures on +both sides of the `2014-02-14` zero-length change, verifies the `2016-05-31` +empty-header change, and rejects Shared Key versions older than `2009-09-19`. +This makes the tests independent from the implementation clock and catches +service-version canonicalization drift. + +A low-level streamed body using Shared Key must provide `content-length` because +the signer cannot know the stream length without consuming it. High-level +`put()` avoids this problem by splitting the stream into known-size block +requests. + + +The service version controls write limits +----------------------------------------- + +Azure Blob limits changed across REST service versions. The client resolves +the relevant limit from the selected version instead of assuming the newest +size everywhere. + +`AZURE_LIMITS` records the values used by planning: + +| Operation/era | Client limit | +| ------------- | ------------ | +| Maximum committed blocks | 50,000 | +| Maximum uncommitted blocks | 100,000 | +| `Copy Blob From URL` synchronous copy | 256 MiB | +| Old `Put Block` | 4 MiB | +| 2016-05-31 through 2019-era `Put Block` | 100 MiB | +| Current `Put Block` | 4,000 MiB | +| Old `Put Blob` | 64 MiB | +| 2016-05-31 through 2019-era `Put Blob` | 256 MiB | +| Current `Put Blob` | 5,000 MiB | + +The implementation selects: + +```text +Put Block limit + version >= 2019-12-12 -> 4,000 MiB + version >= 2016-05-31 -> 100 MiB + older -> 4 MiB + +Put Blob limit + version >= 2019-12-12 -> 5,000 MiB + version >= 2016-05-31 -> 256 MiB + older -> 64 MiB +``` + +`Put Block From URL` uses the version table published on the current REST page: +4,000 MiB from version `2020-04-08` onward and 100 MiB before that point. +Microsoft's same page currently contains a contradictory prose sentence that +still says the operation is limited to 100 MiB. The version table is also +consistent with the service's modern block-blob capacity model, so the client +follows the table. This contradiction is recorded as an upstream documentation +risk rather than hidden. The Docker suite uses small ranges and therefore does +not prove the 4,000 MiB ceiling; an opt-in real Azure test must protect that +limit before it is treated as independently verified. + +`blockSize` defaults to 8 MiB. The constructor rejects a configured block size +above the selected service-version limit. A known body can require a larger +block size to remain within 50,000 committed blocks; the planner chooses the +larger legal size and rejects an impossible request before starting the commit. + + +High-level upload has two paths +------------------------------ + +A materialized `Uint8Array` at or below the selected `Put Blob` limit uses one +`Put Blob` request with: + +```text +x-ms-blob-type: BlockBlob +``` + +A larger materialized body or a stream uses uncommitted blocks: + +```text +source bytes + | + v +fixed-size chunks + | + +--> Put Block A + +--> Put Block B bounded by `concurrency` + +--> ... + `--> Put Block N + | + v +Put Block List + | + v +Get Blob Properties +``` + +Block IDs are deterministic Base64 values derived from zero-padded sequential +numbers. Every request in one upload therefore has a stable order and the +commit document can list exactly the intended blocks. + +`Put Block` requests intentionally do not receive destination `If-Match` or +`If-None-Match`. Uncommitted blocks are not yet the authoritative destination +blob. The precondition and final metadata belong on `Put Block List`, which is +the operation that commits the new block blob. + +The XML block list is generated through `@std/xml/stringify`, so block IDs are +serialized by a real XML implementation rather than hand-escaped text. + +When `ObjectPutOptionsType.size` is supplied, the final streamed byte count must +match. A mismatch rejects the operation before final commit. + +Unlike S3 multipart uploads, Azure uncommitted blocks do not have a separate +abort REST operation. Failed uploads can leave uncommitted blocks until Azure +cleans them up according to service policy. Documentation and tests therefore +must not describe stream cancellation as an atomic remote rollback. + + +Range reads use the Blob range contract +--------------------------------------- + +`get()` maps package range fields to: + +```text +x-ms-range: bytes=start-end +``` + +`at` is the zero-based first byte. `length` controls the inclusive final byte. +When no length is supplied, the range remains open-ended. + +The client reports `rangeRead: true` because Azure Blob Storage can satisfy the +range at the provider rather than materializing the complete object in the +library first. + + +Server-side copy preserves the Azure model +------------------------------------------ + +Azure has more than one URL-based copy primitive. The client selects between +them rather than presenting one fictitious universal copy call. + +For a source up to 256 MiB, `copy()` uses synchronous `Copy Blob From URL`. +For a larger source, it performs ranged `Put Block From URL` operations and then +commits them with `Put Block List`. + +```text +Get Blob Properties(source) + | + +-- <= 256 MiB --> Copy Blob From URL + | + `-- > 256 MiB --> Put Block From URL range 1 + Put Block From URL range 2 + ... + | + v + Put Block List +``` + +The large-copy block size is increased when required to remain at or below +50,000 committed blocks. It is also constrained by the selected REST version's +`Put Block From URL` range limit. + +The copy feature is advertised only when the selected service version supports +URL-copy operations and the configured credential type lets the client derive a +source authorization strategy. + +For same-client copies: + + - SAS includes its authorization on the generated source URL; + - Shared Key can authorize the destination and a same-account source; + - bearer credentials can use `x-ms-copy-source-authorization` from service + version `2020-10-02` onward; + - custom header authorization is not assumed to work for the source, so the + portable `copy` capability is disabled. + +Cross-account copy has additional source-authorization requirements. A Shared +Key for the destination account cannot sign a different account's source. Use +a source SAS or a suitable bearer/source authorization design rather than +assuming one account key grants cross-account access. + +Source conditions map to the `x-ms-source-*` condition family where the selected +operation supports them. Destination `If-Match` / `If-None-Match` apply to the +single synchronous copy or to the final block-list commit for multipart copy. + + +Listing is container pagination, not a directory API +----------------------------------------------------- + +`list()` requests the container with `restype=container&comp=list` and maps: + +| Package field | Azure query field | +| ------------- | ----------------- | +| `prefix` | `prefix` | +| `delimiter` | `delimiter` | +| `limit` | `maxresults` | +| `cursor` | `marker` | + +`Blob` entries become `ObjectEntryType`. `BlobPrefix` entries become child +prefixes. `NextMarker` is returned as the next cursor. + +The filesystem adapter interprets directory markers and provider prefixes. The +Azure client itself retains object/blob terminology because Azure has no native +filesystem directory in the Blob service contract used here. + + +Errors retain Azure request evidence +------------------------------------ + +`AzureError` keeps: + +```text +HTTP status +Azure error code when present +x-ms-request-id when present +original Response +``` + +Azure can return XML or provider-specific text. The error parser uses structured +XML when available and retains the response even when a field is missing. + +The client has a configurable transport retry policy built on `@std/async/retry`. Client options control retry count, exponential +delay, jitter, and an optional per-attempt timeout. The policy retries 408, 429, 5xx, and transport failures for replayable +requests. Authorization is rebuilt on every attempt, which matters for refreshable bearer/custom credentials and Shared Key +dates. Redirects are manual so authorization is not silently carried to another authority. + +A one-shot `ReadableStream` receives one attempt. The low-level `request()` API also accepts `retry: false` because replayability +does not prove that a provider-specific operation is safe to repeat. `request: { retries: 0 }` disables automatic retry for the +client. Provider-specific `Retry-After` interpretation is not yet modeled. + +`getMetrics()` returns request, retry, terminal-failure, response, and optional Fetch-duration counters. `metrics: "none"` is the +baseline benchmark setting; `basic` counts; `timing` adds monotonic duration. + +`AbortSignal` reaches every Fetch operation. Cancellation ends local admission and HTTP work where Fetch can abort it. It does +not guarantee that Azure failed to accept a request before the signal reached the network stack. + + +Azurite is an integration target, not the specification +------------------------------------------------------- + +`tests/provider/compose.yml` runs the official Azurite Blob image with the +well-known development account: + +```text +account: devstoreaccount1 +endpoint: http://127.0.0.1:10000/devstoreaccount1 +``` + +The test uses Shared Key, creates a disposable container through the client's +signed low-level request, then exercises PUT, HEAD, range GET, conditional +create, block upload, server-side copy, list, delete, and the object-store +filesystem adapter. + +Azurite is intentionally treated as an emulator. Its documentation states that +it provides best-effort Azure Storage compatibility and can differ from the +cloud service. A green Azurite suite therefore proves real HTTP/authentication +interoperability, not complete conformance with every Azure version or feature. + +The deterministic unit suite remains responsible for exact Shared Key string +construction, version-dependent limits, condition placement, and source bearer +version gates. + +Before release, an opt-in real Azure Blob test should run against a disposable +container with short-lived CI credentials when organizational secret policy +permits it. + + +Known non-goals +--------------- + +The current direct client does not claim full Azure Storage coverage. Important +features outside this focused contract include: + + - hierarchical namespace/Data Lake Gen2 filesystem semantics; + - append blobs and page blobs; + - leases as a first-class high-level API; + - snapshots/version-ID aware filesystem paths; + - customer-provided encryption-key convenience APIs; + - immutability policies and legal holds; + - blob index tags; + - asynchronous `Copy Blob` polling workflows; + - batch operations; + - account/container administration beyond low-level requests; + - adaptive provider throttling beyond the configured exponential retry policy, including provider-specific `Retry-After` scheduling; + - Microsoft Entra token acquisition itself. + +`request()` can reach an unmodeled REST operation when a caller supplies the +correct method, query, headers, and body. A feature should become a typed public +operation only when its ownership, failure behavior, version gates, and tests +are explicit. + + +Primary specification sources +----------------------------- + +Review these Microsoft sources before changing protocol behavior: + + - Shared Key authorization: + https://learn.microsoft.com/rest/api/storageservices/authorize-with-shared-key + - Versioning for Azure Storage services: + https://learn.microsoft.com/rest/api/storageservices/versioning-for-the-azure-storage-services + - Put Blob: + https://learn.microsoft.com/rest/api/storageservices/put-blob + - Put Block: + https://learn.microsoft.com/rest/api/storageservices/put-block + - Put Block List: + https://learn.microsoft.com/rest/api/storageservices/put-block-list + - Put Block From URL: + https://learn.microsoft.com/rest/api/storageservices/put-block-from-url + - Copy Blob From URL: + https://learn.microsoft.com/rest/api/storageservices/copy-blob-from-url + - List Blobs: + https://learn.microsoft.com/rest/api/storageservices/list-blobs + - Azurite: + https://learn.microsoft.com/azure/storage/common/storage-use-azurite + +The REST documentation is authoritative for Azure. Azurite source and behavior +are integration evidence for the emulator only. diff --git a/docs/design.md b/docs/design.md index 65fe72e..4cee4b1 100644 --- a/docs/design.md +++ b/docs/design.md @@ -1,185 +1,115 @@ Architecture and invariants =========================== -This document explains why `@okikio/opfs` has two frontend styles, one adapter contract, and a second record-store contract. +The architecture starts from one rule: -The core rule is simple: +> The filesystem facade owns filesystem semantics. An adapter owns the mechanics of one backend. -> Filesystem semantics belong to the filesystem facade. Persistence mechanics belong to the adapter. +That rule matters because OPFS, Node files, Deno KV, SQLite, S3, and Azure Blob do not have the same native operations. A useful +portable library must make common application behavior consistent without hiding those differences from performance-sensitive +or correctness-sensitive code. -That rule keeps OPFS-style application code reusable without flattening meaningful differences between a browser filesystem, a host filesystem, a key-value store, a document collection, and a SQL database. - -The complete data path ----------------------- +The complete data path is: ```text - application - | - +------------+------------+ - | | - v v - path methods OPFS-shaped handles - readFile/writeFile FileHandle/DirectoryHandle - | | - +------------+------------+ - | - v - FileSystemType - | - path normalization / errors / cancellation - recursive copy / move / walk / remove - lock ownership / sync-file lifetime - stream fallback / buffer limit - | - v - AdapterType - | - +---------------+---------------+ - | | - v v - native adapters RecordStoreType - OPFS/Deno/Bun/Node | - | - +------------------+------------------+ - | | | - v v v - unstorage RxDB SQL rows - | - +------+------+ - | | - v v - db0 Drizzle + application + | + +---------------+---------------+ + | | + v v + path API OPFS-shaped handles + readFile / writeFile / walk DirectoryHandle / FileHandle + | | + +---------------+---------------+ + | + v + FileSystemType + | + normalize paths / normalize failures / cancellation + parent creation / recursive operations / staging + file locks / tree locks / sync-file lock lifetime + stream selection / bounded materialization / ownership + | + v + AdapterType + +----------------+----------------+ + | | | + v v v + native filesystem record storage object storage + OPFS/Node/Deno KV/doc/SQL S3/Azure Blob + | | | + | RecordStoreType ObjectStoreType + | | | + v v v + native bytes versioned row object key ``` -Why `AdapterType` is small --------------------------- +The frontend therefore has one filesystem contract, but the adapter capability record still tells the truth about how that +backend gets the work done. -The required primitive set is deliberately smaller than the public filesystem API: +The adapter contract stays deliberately small +--------------------------------------------- -```text -stat -readFile -writeFile -readDir -createDir -remove -``` +Every adapter implements six primitives: `stat`, `readFile`, `writeFile`, `readDir`, `createDir`, and `remove`. Everything else +is an optional acceleration or stronger native lifecycle. + +This avoids a common adapter failure mode where every backend reimplements recursive copy, walk, parent creation, handles, +locking, and error normalization separately. If those policies live in every adapter, semantics drift as soon as one backend gets +a bug fix that the others do not. -Optional native capabilities add faster or stronger paths: +Optional operations exist only when the backend can perform them natively: ```text -openReadStream -writeStream -move -openSyncFile +streamRead -> openReadStream +write mode -> writeStream when mode is in streamWriteModes +nativeCopy -> copy +nativeMove -> move +positionalWrite -> openWritableFile +syncAccess -> openSyncFile ``` -The facade builds higher-level behavior from these primitives. A custom backend therefore does not need to implement recursive copy, recursive remove, parent creation, OPFS-shaped handles, or lock orchestration independently. +`streamWriteModes` is intentionally a list. A local file can stream replace, append, and update. An object store can usually +stream a complete replacement but cannot append bytes to an existing object in place. One `streamWrite: true` flag would hide +that difference and make callers reason from a capability that was too broad. -This avoids a common adapter anti-pattern where every backend reimplements the full public API and gradually develops different semantics. +The adapter can additionally publish hard `limits` and a durable `partition` description. Those are facts about the configured +backend, not policy guesses. `FileSystemType.inspect()` combines them with resolved optimization controls and effective facade +support. `plan()` uses the same information before I/O, so runtime execution and preflight selection share one route model. -Capability flags describe native behavior ------------------------------------------ - -`AdapterCapabilitiesSchema` contains: +Route-changing optimizations are facade policy: ```text -read -write streamRead streamWrite rangeRead +nativeCopy nativeMove -positionalWrite -syncAccess -``` - -These values describe what the adapter itself can do. They do not describe everything the facade can emulate. - -For example, a record-store adapter reports `streamWrite: false`. The facade can still accept a `ReadableStream`, but it must buffer the stream before storing the record. The capability remains false because pretending that buffering is native streaming would hide an important memory and latency difference. - -The same rule applies to `positionalWrite`. A record adapter can implement one `writeFile(update)` operation by reading and replacing a record, but that does not mean it can keep a writable file open across thousands of positional chunks. Record adapters therefore report `positionalWrite: false`. Native OPFS, Node, Deno, and Bun adapters expose `openWritableFile()` when they can keep one underlying writable resource open. - -The record-store layer ----------------------- - -Value stores, document stores, and SQL databases do not naturally expose files and directories. `RecordStoreType` is the reusable translation point for those systems. - -```ts -interface RecordStoreType { - get(path): Promise; - set(record): Promise; - delete(path): Promise; - list(parent): AsyncIterableIterator; - dispose?(): void | Promise; -} ``` -The shared record is versioned and validated by Zod. +Each defaults to enabled and can be disabled independently. This is deliberately different from capability detection. The adapter +should expose the strongest implementation it has; the caller can force the safe fallback for differential tests, observability, +provider workarounds, or policy. A disabled route is never relabelled native. -Directory: - -```json -{ - "version": 1, - "path": "/projects", - "parent": "/", - "name": "projects", - "kind": "directory", - "lastModified": 1786550000000 -} -``` - -File: - -```json -{ - "version": 1, - "path": "/projects/state.bin", - "parent": "/projects", - "name": "state.bin", - "kind": "file", - "data": "AAECAwQ=", - "size": 5, - "lastModified": 1786550000000, - "mediaType": "application/octet-stream" -} -``` - -`path` is the durable logical identity. `parent` is stored independently because directory listing should not require parsing every stored path. Backends are free to index `parent` in the way that best fits the provider. - -File bytes are base64. The choice is not an assertion that base64 is the most storage-efficient format. It is the common representation that survives JSON, RxDB documents, unstorage values, and SQL text columns without backend-specific binary contracts. Native filesystem adapters do not pay this cost. - -Streaming policy ----------------- - -Native adapters stream when the backend gives a real streaming primitive. - -Record-backed adapters materialize one file record. A streamed input therefore passes through this sequence: +`nativeCopy` is also separate. A provider-side S3 or Azure copy can move terabytes without transferring the source through this +process, even though the same provider has no filesystem rename. The facade checks native copy before opening a source stream. +That ordering is a performance invariant, not an implementation detail: ```text -ReadableStream / AsyncIterable - | - v - bounded byte collector - | - +-------+-------+ - | | - under limit over limit - | | - v v - RecordStore.set cancel source - throw too-large -``` - -`maxBufferedWriteBytes` defaults to 64 MiB. The limit is part of `FileSystemOptionsType`, not a hidden constant inside each database adapter, so the application can choose a memory policy once. +correct selection +filesystem.copy() + | + +-- native copy available -> adapter.copy() + | + `-- no native copy -------> open source stream -> transfer -Path invariant --------------- +incorrect selection +open source stream -> discover native copy -> source GET was already wasted +``` -Every adapter receives canonical virtual paths. +Paths are virtual identities, not host paths +------------------------------------------- -Valid: +Every adapter receives a canonical `PathType`: ```text / @@ -187,7 +117,7 @@ Valid: /a/b.txt ``` -Rejected at the canonical adapter seam: +The adapter seam rejects or never receives forms such as: ```text a/b @@ -198,160 +128,270 @@ a/b /a\b ``` -Public path APIs may accept relative or non-canonical input. `normalizePath()` resolves it before the adapter sees it. +Public methods accept more convenient input and call `normalizePath()` first. This lets callers write ordinary path-like input +without making every adapter repeat normalization rules. + +The host filesystem adapters map this virtual namespace under one configured host root. A virtual path cannot escape that root +after host resolution. The virtual namespace does not expose symbolic-link identity, permission bits, or arbitrary host paths +as part of the portable contract. + +Record stores and object stores need different translation layers +----------------------------------------------------------------- -The virtual path namespace is not a host path namespace. `createLocalPath()` maps virtual paths below one configured host root and verifies that the result does not escape that root. +A value store naturally answers "what value is stored at this key?" It does not naturally answer filesystem questions such as +"what are the direct children of this directory?" `RecordStoreType` supplies the reusable record translation for that family. -Handle invariant ----------------- +The complete persisted record has a canonical `path` plus a separate `parent`. Direct directory listing can therefore use an +index or prefix query over `parent` instead of scanning and parsing every path. File bytes are base64 so the fallback shape can +survive JSON, Web Storage, document databases, and SQL text. The extra storage and encoding work is accepted only for this +complete-record path. Native file and object adapters do not use that representation. -`FileHandle` and `DirectoryHandle` are facades. They are not native `FileSystemHandle` objects and they do not claim to be. +The record contract also has optional byte lanes. A store can provide metadata-only stat, direct ranges, streaming reads, direct +materialized writes for selected modes, or streaming writes for selected modes. The generic record adapter advertises only the +lanes the store declares. Deno KV uses these lanes so a partitioned file is not reconstructed into one base64 record for stat, +listing, range reads, streaming reads, materialized append/update, or streamed replacement. Its append/update lane constructs a +new immutable generation one part at a time. This still performs provider I/O for untouched bytes, but it keeps JavaScript memory +bounded by the configured part/concurrency policy. Simpler record stores keep the small complete-record contract. -The facades preserve the useful OPFS programming shape: +An object store has a different strength: large objects, byte ranges, prefix listing, whole-object replacement, conditional +requests, and provider-side copy. `ObjectStoreType` preserves those concepts before `createObjectAdapter()` translates them into +filesystem operations. + +Files map directly to object keys. Empty directories need a marker object because a pure prefix does not exist until at least one +child exists: ```text -root.getFileHandle() -root.getDirectoryHandle() -file.getFile() -file.createWritable() -file.createSyncAccessHandle() -directory.removeEntry() -directory.resolve() -entries()/keys()/values() +/photos -> photos/ marker +/photos/a.jpg -> photos/a.jpg +/photos/2026/b.jpg -> photos/2026/b.jpg ``` -They also expose a package-specific canonical `path` property because the adapter architecture needs a stable logical address. +The adapter also accepts implicit directories inferred from foreign prefixes. This matters when the bucket/container is not +created exclusively by this library. -`createWritable()` stages an in-memory file image and commits only on close. Abort discards the staged image. This mirrors the commit-on-close behavior an application expects from the File System API, but it is intentionally not the recommended large-file path. Large sequential writes should use `FileSystemType.writeFile()` so a streaming-capable adapter can bypass the staged image. +An object namespace can contain both `mixed` and `mixed/child`. A real filesystem cannot. The filesystem view resolves an exact +`mixed` object as the file, because exact `stat()` already does that. Reads and writes follow the same rule. This creates one +stable interpretation for a foreign namespace instead of making `stat()` and `writeFile()` disagree. -Coordination invariant ----------------------- +Streaming stays native only when the backend really streams +----------------------------------------------------------- -There are two classes of mutation. +`writeFile()` accepts strings, Blob, ArrayBuffer, typed-array views, ReadableStream, and AsyncIterable input. -File mutation: +When the selected adapter supports native streaming for the requested write mode, the facade forwards a byte stream directly. +When it does not, the facade collects the stream below `maxBufferedWriteBytes` and then calls the materialized adapter write. +Crossing the limit cancels the producer and returns `too-large`. ```text -shared tree lock +ReadableStream | -exclusive /path/to/file lock + +-- native mode supported ------> adapter.writeStream() | -write or sync file lifetime + `-- no direct adapter stream lane + | + v + bounded collector + | | + | +-- over limit -> cancel producer -> too-large + v + Uint8Array + | + v + adapter.writeFile() ``` -Structural mutation: +This makes memory behavior visible. A simple record-backed adapter can accept streamed input through the public API while still +reporting an emulated stream route because the complete record is materialized before storage. A specialized record store can +report a partitioned stream lane when its own physical layout preserves backpressure. Deno KV does exactly that for +replacement streams when partitioning is enabled. + +Partitioning is not hidden as an optimization. It changes durable physical layout, so the adapter publishes `mode`, part size, +threshold, maximum parts, and layout identity. Deno KV exposes `never | auto | always`. Its parts are written under a new +generation and the manifest is committed last. A pre-manifest crash can leak unreachable parts but cannot publish a partial new +logical file. + +Multipart and block-upload clients use `@std/async/pool` for bounded request admission. The surrounding client still owns the +provider lifecycle: S3 waits for already-started part requests before it sends AbortMultipartUpload, while Azure documents that +uncommitted blocks have no equivalent abort operation. The pool limits concurrent work; it does not become authority for remote +commit, cleanup, or the terminal provider failure. + +HTTP retry policy is separate from body replayability. Direct clients rebuild authorization on every retry and use exponential +backoff with jitter, but a mechanically replayable request can still be semantically non-idempotent. S3 multipart initiation and +completion therefore disable automatic request retry. Uploaded parts use stable part numbers and can use the normal retry policy. +The low-level S3/Azure request APIs expose `retry: false` so a caller can make the same decision for provider-specific operations. +One-shot `ReadableStream` request bodies are never retried automatically. + +Object append and update are optimistic read-modify-write +-------------------------------------------------------- + +Object stores do not expose a portable in-place byte update. Append and update therefore use the current object as the starting +image, modify that image, and replace the object. + +Without a precondition, two writers can both read version A and then publish different replacements; the later one silently +loses the earlier write. When the object client advertises conditional writes, the adapter uses the current ETag as `If-Match`. +A concurrent change then fails the replacement instead of becoming silent data loss. ```text -exclusive tree lock - | -copy / move / recursive remove / emptyDir +writer A: HEAD ETag=A -> GET A ---------> PUT if-match A -> succeeds +writer B: HEAD ETag=A -> GET A -------------------------> PUT if-match A -> fails ``` -This lets independent files make progress at the same time while ensuring that a recursive tree mutation cannot race an active library file mutation. +If a provider claims conditional writes but does not return an ETag for an existing object, the adapter refuses append/update. +That is safer than publishing an unconditional write while the capability record says optimistic protection exists. -`local` coordination shares lock state by lock name inside one JavaScript realm. New readers queue behind an already-waiting exclusive request so writers do not starve. +The provider client can disable `conditionalWrite` when a compatible protocol implementation does not support the required +precondition. The library does not choose provider behavior from a provider-name table. -`web-locks` uses the browser Web Locks API. `auto` selects Web Locks when present and local FIFO locks otherwise. `none` retains cancellation checks but does not coordinate mutations. +Copy and move preserve their real commit behavior +------------------------------------------------- -The adapter still owns any stronger backend-level locking. Library locks are application-level coordination for callers that use this library. +Native host filesystems use their copy and rename operations when available. Object stores use provider-side copy when the +client can do it. The facade removes/rejects the destination according to its own overwrite contract before invoking native copy, +so the adapter does not have to invent another overwrite policy. -Asynchronous positional file lifecycle --------------------------------------- +When no native copy exists, file bytes move through a stream when both ends support streaming or through bounded materialization +otherwise. -A long-lived asynchronous writable owns the same file mutation lock from its create/check/open sequence until terminal cleanup. +A native move can be atomic or near-atomic according to the host/provider operation. The portable fallback is explicitly: ```text -facade file lock <------ same lifetime ------> adapter writable file - | | - +-------------- close/abort --------------+ +copy source -> destination + | + +-- copy failed -> source remains + | + `-- copy succeeded -> remove source ``` -The facade deliberately does not build this contract from repeated `writeFile(update)` calls. Backends with native or staged file resources can preserve positional-write throughput, while record backends remain explicit about the fact that they materialize whole values. +That fallback is not atomic. A failure after copy and before remove can leave both entries. The API documents this instead of +claiming POSIX rename semantics on every backend. -Cancellation stops ordinary `write()`, `truncate()`, and `flush()` calls. It does not prevent `close()` or `abort()` from releasing the backend resource and the facade lock. If a backend cannot roll back writes, callers that need all-or-nothing publication should use a staging path. +Before copy or move, source and destination are checked for overlap. The library never removes an overwrite destination that is +an ancestor or descendant of the source. -Synchronous file lifecycle --------------------------- +Coordination protects cooperating callers, not the whole storage system +------------------------------------------------------------------------ + +There are two classes of mutation. -A synchronous file has two resources with one lifetime: +A file mutation acquires a shared tree lock plus an exclusive lock for that canonical file path: ```text -facade path lock <------ same lifetime ------> adapter sync file - | | - +---------------- close() ----------------+ +shared tree lock + | +exclusive /a/file lock + | +write or sync-file lifetime ``` -`ManagedSyncFile` keeps the path lock until the native resource closes. This prevents an async write through the same facade from entering while synchronous random access is active. +A structural mutation such as recursive copy, move, remove, or empty-directory work acquires the exclusive tree lock: -`writeAll()` must handle partial writes. It repeats the write until the complete input is committed or the backend reports no progress. +```text +exclusive tree lock + | +structural mutation +``` -Move semantics --------------- +Independent files can therefore make progress concurrently while a tree mutation cannot race an active library file mutation. +The in-realm lock implementation queues new readers behind an already-waiting writer so a busy read/write workload does not +starve structural work. -Adapters with a native rename/move set `nativeMove: true` and provide `move()`. +`coordination: "web-locks"` uses the browser Web Locks API. `auto` uses Web Locks when exposed and falls back to in-realm FIFO +coordination. `local` is one-realm coordination only. `none` preserves cancellation and adapter semantics but makes the caller +responsible for concurrency. -```text -Deno/Bun/Node -source -------- native rename --------> destination -``` +None of these modes becomes a distributed lock. Separate Node processes, browser profiles, hosts, or independent applications +need provider/database coordination when same-path atomicity matters across those processes. -Adapters without that primitive use: +Synchronous and asynchronous writable resources own locks for their complete lifetime +-------------------------------------------------------------------------------------- + +A synchronous file is not one short method call. It owns both the adapter file resource and the facade path lock until close: ```text -source ---- copy ----> destination - | - +---- remove source after successful copy +facade path lock <-------- same lifetime --------> adapter sync file + | | + +------------------- close() ------------------+ ``` -The second sequence is not atomic. A failure between copy and remove can leave both entries. The API and documentation state this rather than presenting every backend as a POSIX filesystem. +This prevents an asynchronous write through the same facade from entering while synchronous random access is active. +`writeAll()` loops over partial native writes until the complete input is written or the backend reports no progress. -Before either form, source and destination are checked for ancestor overlap. An overwrite never removes an ancestor or descendant containing the source. +The OPFS-shaped `createWritable()` facade stages one file image and commits it on close. Abort discards the staged image. This is +useful for compatibility with File System API write commands, including seek and truncate. It is not the recommended path for +very large sequential files because the staged image is materialized. `FileSystemType.writeFile()` can use an adapter's native +streaming path instead. -Resource ownership ------------------- +Integration direction is explicit +--------------------------------- -Injected resources are borrowed by default. +An adapter is `ecosystem -> OPFS`. A driver is `OPFS -> ecosystem`. A bridge is only a descriptor that groups those directions; +it does not add a third translation layer to each operation. ```text -caller creates resource - | - +----> adapter borrows resource - | | - | +---- filesystem closes - | +---- resource stays open - | - +---- caller still owns resource +ecosystem/client -> adapter -> FileSystemType -> driver -> ecosystem contract ``` -Ownership changes only through an explicit option: +Some ecosystems are genuinely bidirectional. unstorage has a storage contract that can be consumed as a record backend and a +driver contract that can be implemented over `FileSystemType`. RxDB, db0, and Drizzle are not symmetric: their reverse +contracts require query, conflict, dialect, schema, or change-stream semantics a filesystem does not own. `defineBridge()` +therefore requires an explicit reason for unsupported directions instead of encouraging a false adapter. + +Cancellation and disposal are different operations +-------------------------------------------------- + +An `AbortSignal` asks active work to stop before a commit when possible. Closing a filesystem ends ownership of the facade. +Closing the facade does not dispose the adapter unless `disposeAdapter: true` transferred that ownership. + +The same rule continues below the adapter: ```text -disposeAdapter -disposeStorage -disposeDatabase -disposeFileSystem +caller creates database/client/cache + | + +--> adapter borrows it + | | + | +--> filesystem closes + | `--> resource remains open + | + `--> caller still owns resource ``` -This rule matters for connection pools, shared RxDB collections, process-wide unstorage instances, and server databases. A library adapter must not quietly dispose infrastructure that another subsystem still owns. +An adapter option such as `disposeDatabase`, `disposeStore`, or another explicit ownership flag changes that lifecycle. The +option exists because connection pools, RxDB collections, unstorage instances, object clients, and caches are commonly shared by +more than one subsystem. -Error invariant ---------------- +Errors normalize the portable category without erasing the provider cause +------------------------------------------------------------------------- -Backends fail differently. Browsers use DOMException names. Node commonly reports `error.code`. Database bridges can throw provider errors. +Browsers use DOMException names. Node/Deno/Bun expose host error codes. Databases and cloud providers have their own errors. +`toFileSystemError()` maps known failures to stable package categories while retaining the original `cause`. -`toFileSystemError()` normalizes known failures to stable categories while retaining the original `cause`. The package does not erase unexpected backend failures into one generic string. +S3 and Azure clients also retain provider request identities on their own errors. Those IDs matter when a service returns an +unexpected result and the provider support logs are the only authoritative trace. -Adapter import invariant ------------------------- +Import safety follows the package graph +--------------------------------------- -The root package is import-safe for browsers. Runtime-specific code remains behind explicit subpaths. +The root package exports the portable facade, native browser OPFS convenience path, schemas, errors, handles, and capability +probes. It does not export every adapter from the root. ```text -@okikio/opfs browser-safe core + native OPFS -@okikio/opfs/adapter/node node:fs imports -@okikio/opfs/adapter/deno Deno globals -@okikio/opfs/adapter/bun Bun globals + Node compatibility APIs -@okikio/opfs/adapter/drizzle optional drizzle-orm peer +@okikio/opfs browser-safe core + native OPFS +@okikio/opfs/adapter/node node:fs imports +@okikio/opfs/adapter/deno Deno runtime APIs +@okikio/opfs/adapter/bun Bun + Node-compatible APIs +@okikio/opfs/s3 Web Fetch/Crypto S3 client +@okikio/opfs/azure Web Fetch Azure Blob client +@okikio/opfs/adapter/drizzle optional drizzle-orm peer ``` -No adapter configures logging, reads environment variables, connects to a database, or mutates global application state merely because the module was imported. +Importing a module does not read environment variables, configure logs, connect to providers, start workers, or mutate a global +adapter registry. + +Schemas are executable contracts, not duplicated type declarations +------------------------------------------------------------------- + +Project-owned structural values use Zod schemas and inferred TypeScript output types. Public schema constants end in `Schema`. +Serializable project-owned types normally end in `Type`. + +Zod 4 implements Standard Schema. The exported Zod value is therefore also the Standard Schema value. Creating a second OPFS +schema wrapper would add maintenance without adding a stronger contract. diff --git a/docs/ecosystems.md b/docs/ecosystems.md index 6bc4725..73fcfde 100644 --- a/docs/ecosystems.md +++ b/docs/ecosystems.md @@ -1,228 +1,158 @@ Ecosystem integrations ====================== -The package integrates at the highest stable storage abstraction each ecosystem already provides. This is intentional. Reimplementing every upstream driver inside `@okikio/opfs` would duplicate provider code and create a second compatibility matrix that would immediately drift. +`@okikio/opfs` integrates with another storage ecosystem at the highest stable abstraction the application already owns. It does +not duplicate every provider driver from that ecosystem. + +That rule gives two complementary directions: ```text -unstorage: Storage -> RecordStoreType -> AdapterType -RxDB: RxCollection -> RecordStoreType -> AdapterType -db0: Database -> RecordStoreType -> AdapterType -Drizzle: Database+Table -> RecordStoreType -> AdapterType +existing storage resource existing OPFS filesystem + | | + v v + adapter reverse driver + | | + v v + FileSystemType KV / unstorage API ``` -The filesystem semantics above those bridges are identical. +The forward path lets filesystem-shaped application code use another storage system. The reverse path lets another ecosystem +consume any backend already reachable through `FileSystemType`. -unstorage ---------- +Bridge descriptors make asymmetry part of the contract +------------------------------------------------ -`createUnstorageAdapter(storage)` accepts the high-level unstorage `Storage` contract. It uses: +`@okikio/opfs/bridge` groups the existing forward adapter and reverse driver for an ecosystem. It does not force every +integration to be symmetric. -```text -getItem -setItem -removeItem -getKeys -optional dispose -``` +| Bridge | ecosystem -> OPFS | OPFS -> ecosystem | Why the reverse side is absent when unsupported | +| --- | --- | --- | --- | +| `UnstorageBridge` | yes | yes | both stable shapes exist | +| `RxDbBridge` | yes | no | `RxStorage` also owns queries, conflicts, change streams, cleanup, and storage-instance semantics | +| `Db0Bridge` | yes | no | a filesystem is not a SQL query/dialect engine | +| `DrizzleBridge` | yes | no | a filesystem does not own Drizzle schema, dialect, or query-builder behavior | +| `KeyValueBridge` | no | yes | the generic reverse KV shape does not define persistence semantics needed to build an adapter | -This means the bridge is independent of the mounted driver. - -unstorage's generated built-in driver catalog includes these families: - -- Azure App Configuration, Cosmos, Key Vault, Storage Blob, and Storage Table -- Capacitor Preferences -- Cloudflare Cache, KV binding/HTTP, and R2 -- db0 -- Deno KV and Deno KV Node -- fs and fs-lite -- GitHub -- HTTP -- IndexedDB -- localStorage and sessionStorage -- LRU cache and memory -- MongoDB -- Netlify Blobs -- null and overlay -- PlanetScale -- Redis -- S3 -- UploadThing -- Upstash -- Vercel Blob and Vercel Runtime Cache - -The list is upstream inventory, not a claim that every provider has filesystem-quality write semantics. A driver can be read-only, eventually consistent, size-limited, or expensive to enumerate. Configure `{ readOnly: true }` when the selected Storage cannot safely mutate. - -The adapter stores records below a reserved key prefix, `opfs` by default. Virtual path segments are encoded reversibly before they become unstorage key segments. +`defineBridge()` validates that direction declarations agree with real constructors. An unsupported direction must include a +reason. Third-party integrations can therefore publish capability honestly without inventing a method that only works for a +small subset of the upstream contract. + +Unstorage works in both directions without a provider explosion +--------------------------------------------------------------- + +The forward adapter accepts the high-level unstorage `Storage` contract: ```ts -const adapter = createUnstorageAdapter(storage, { - prefix: "my-app-fs", - readOnly: false, - disposeStorage: false, -}); +import { createStorage } from "unstorage"; +import memoryDriver from "unstorage/drivers/memory"; +import { createFileSystem } from "@okikio/opfs"; +import { createUnstorageAdapter } from "@okikio/opfs/adapter/unstorage"; + +const storage = createStorage({ driver: memoryDriver() }); +const fileSystem = createFileSystem(createUnstorageAdapter(storage)); ``` -### Reverse unstorage direction +Unstorage remains responsible for its selected driver, mounts, provider SDKs, retry behavior, and provider-specific limits. The +OPFS adapter only uses the stable high-level operations it needs to persist records. -`createUnstorageDriver(fileSystem)` lets unstorage consume any `FileSystemType`. +This is intentionally broader than maintaining separate OPFS adapters for every unstorage driver. Current unstorage drivers span +browser storage, Cloudflare, Azure, S3, Deno KV, filesystem, Redis, databases, blobs, HTTP, and other providers. Duplicating that +catalog here would create a second compatibility matrix that would drift from upstream. + +The reverse direction is more powerful after the generic key-value driver: ```text unstorage Storage - | - v + | + v @okikio/opfs unstorage Driver - | - v + | + v +KeyValueDriverType + | + v FileSystemType - | - +--- OPFS - +--- Node/Deno/Bun - +--- RxDB - +--- db0 - +--- Drizzle - +--- custom adapter + | + +-- native OPFS + +-- Node / Deno / Bun + +-- S3 / Azure Blob + +-- localStorage / IndexedDB / Cache + +-- Deno KV / SQLite + +-- RxDB / db0 / Drizzle / unstorage + `-- custom adapter ``` -The driver maps `:` hierarchy segments to private filesystem directories with reversible percent-based encoding. Each key stores its payload in a dedicated `value` leaf file. This indirection is required because unstorage can hold both `foo` and `foo:bar`, while a normal filesystem cannot make `/foo` both a file and a directory. Literal `%` and literal `~` remain distinct. - -`disposeFileSystem` defaults to false because the injected filesystem is borrowed. - -RxDB ----- - -RxDB explicitly defines `RxStorage` as the storage-engine abstraction. The upstream storage interface creates `RxStorageInstance` objects that own bulk writes, queries, attachment access, change streams, cleanup, close, and remove semantics. - -`@okikio/opfs` does not implement `RxStorage`. Instead it uses a normal RxDB collection whose underlying storage can be any RxStorage chosen by the application. - -The package exports `RxDbRecordJsonSchema`: +`createKeyValueDriver()` owns the collision-safe filesystem mapping. `createUnstorageDriver()` only translates that contract to +unstorage's driver method names and `maxDepth` flag. Both reverse views retain `inspect()`, `plan()`, and `getMetrics()` from the +backing filesystem. An ecosystem caller can therefore reject a value above `maxFileBytes`, see when a streamed write would +buffer or partition, disable a native route on the filesystem, and observe the same counters without a second capability table. -```ts -await database.addCollections({ - files: { schema: RxDbRecordJsonSchema }, -}); - -const fileSystem = createFileSystem( - createRxDbAdapter(database.files), -); -``` - -The bridge uses collection operations that preserve RxDB's document concurrency semantics: +The extra key directory is required because a KV store can contain both `foo` and `foo:bar`: ```text -findOne(path).exec() -find({ selector: { parent } }).exec() -incrementalUpsert(record) -incrementalRemove() +foo -> /key-foo/value +foo:bar -> /key-foo/key-bar/value ``` -The `path` field is the primary key. `parent` is indexed for direct directory listing. The exported schema sets `maxLength: 4096` on both indexed path fields because RxDB requires a maximum length for indexed strings; the adapter rejects longer paths before querying or writing the collection. - -### RxStorage coverage - -As reviewed from RxDB's current storage guide on 2026-08-12, upstream documents these storage implementations and wrappers: - -Native/storage implementations: - -- Memory -- LocalStorage -- premium IndexedDB -- premium OPFS -- premium Filesystem Node - -Storage wrappers/infrastructure: - -- premium Worker -- premium SharedWorker -- Remote -- premium Sharding -- premium Memory Mapped -- premium Localstorage Meta Optimizer -- Electron IPC renderer/main integration - -Third-party or premium-backed storage families documented by RxDB: - -- premium Expo Filesystem -- premium SQLite -- Dexie.js -- MongoDB -- DenoKV -- FoundationDB +A naive `:` to `/` conversion would try to make `/foo` both a file and a directory. The private `value` leaf removes that +conflict while reversible segment encoding keeps `%`, `~`, spaces, slashes inside a key segment, and other URI-sensitive text +distinct. -Because this package sits above the collection, the adapter does not need a separate implementation for each item in that list. The selected RxStorage still owns its own requirements, licensing, runtime constraints, multi-instance behavior, replication behavior, and performance characteristics. +RxDB stays above RxStorage +-------------------------- -db0 ---- +RxDB already defines `RxStorage` as its storage-engine abstraction. An RxCollection adds document behavior, indexes, conflict +handling, and the selected RxStorage implementation. -The db0 bridge targets the high-level `Database` interface: +`createRxDbAdapter()` therefore accepts an existing collection instead of implementing another RxStorage engine. The exported +`RxDbRecordJsonSchema` uses canonical `path` as the primary key and indexes `parent` for direct-child listing. ```text -dialect -prepare(sql) -statement.get/all/run -optional dispose -``` - -The current db0 type contract reports four SQL dialects: - -```text -sqlite -libsql -postgresql -mysql +RxDB application + | + v +RxCollection + | + +--> selected RxStorage + | + v +createRxDbAdapter() + | + v +FileSystemType ``` -`createDb0Adapter()` has explicit SQL generation for all four. The behavioral test suite exercises all four dialect branches. - -As reviewed from db0's generated connector catalog on 2026-08-12, upstream connector names include: - -- better-sqlite3 -- bun-sqlite and bun alias -- Cloudflare D1 -- Cloudflare Hyperdrive MySQL -- Cloudflare Hyperdrive PostgreSQL -- libSQL core, HTTP, Node, web, and alias -- mysql2 -- node-sqlite and sqlite alias -- PGlite -- PlanetScale -- PostgreSQL -- sqlite3 - -The bridge depends on the `Database` contract and dialect, not the connector name. A connector therefore does not need bespoke OPFS code when it presents a compatible db0 Database. - -Prepared statements use db0's portable `?` placeholders. PostgreSQL-family db0 connectors own translation to native `$1`, `$2`, and later parameters, so the filesystem bridge does not duplicate connector-specific parameter rewriting. - -### db0 table - -Default table: `opfs_entries`. - -The adapter can initialize it: +This keeps the adapter compatible with the collection regardless of whether the application selected memory, IndexedDB, OPFS, +filesystem, SQLite, remote, worker, or another RxStorage family. RxDB keeps ownership of replication, multi-instance behavior, +licensing, storage wrappers, and conflicts. -```ts -const adapter = await createDb0Adapter(database, { - initialize: true, - table: "opfs_entries", -}); -``` +db0 and direct SQLite share the same SQL record model +------------------------------------------------------ -The path primary key is a SHA-256 hex digest. The original path is stored separately. This is important for MySQL because arbitrary `TEXT` is not a portable primary-key choice. +`createDb0Adapter()` targets db0's high-level `Database` contract and its reported dialect. The SQL generation currently covers +SQLite, libSQL, PostgreSQL, and MySQL branches. It depends on database behavior rather than connector names, so a new db0 +connector does not need a new OPFS adapter when it presents the same database contract. -`parent_path` is used for directory listing. Large installations should add a provider-appropriate index through their normal migration system if directory listing becomes a hot query. +`createSqliteAdapter()` is the focused direct SQLite path for applications that already own a small connected statement API. It +reuses the SQLite branch of the same SQL record contract instead of maintaining a second table layout and upsert implementation. -`disposeDatabase` defaults to false. +The default SQL table is `opfs_entries`. The path identity and parent path are stored separately so direct-child listing can use +a provider-appropriate index. The db0 path uses portable parameter placeholders and lets the db0 connector translate them where +its dialect requires a different native parameter shape. -Drizzle -------- +Drizzle keeps schema ownership with the application +--------------------------------------------------- -Drizzle is not one SQL dialect. It exposes dialect-specific schema builders and many driver entrypoints. The adapter therefore does not create or migrate a universal table. +Drizzle spans several SQL dialects and runtime drivers. A universal OPFS-owned Drizzle table would either choose one dialect or +hide dialect-specific DDL details. -The caller provides: +`createDrizzleAdapter()` therefore receives: -1. a connected Drizzle database; +1. an already-connected Drizzle database; 2. a table built for that database dialect; 3. the required logical columns. -Required table properties: +Required logical fields are: ```text path @@ -235,40 +165,102 @@ lastModified mediaType ``` -`path` must be unique or a primary key. `size` and `lastModified` must round-trip JavaScript safe integers. +`path` must be unique or primary. `size` and `lastModified` must round-trip JavaScript safe integers. The bridge uses the common +select/insert/delete builder shape and keeps Drizzle an optional peer dependency. + +Inside one `FileSystemType`, normal coordination serializes same-path mutations. Separate processes or hosts are not serialized +by an in-memory facade lock. A database-backed deployment that needs cross-process replacement atomicity must use transactions, +leases, advisory locks, or another mechanism provided by its actual database/driver. -The bridge uses Drizzle's common CRUD shape: +S3-compatible providers are configured by capability, not brand guesses +------------------------------------------------------------------------ + +The direct S3 client is meant to work with AWS S3 and compatible XML/SigV4 services, but "S3-compatible" is not a promise that +all operations, preconditions, limits, checksums, or control-plane features are identical. + +The client options deliberately separate the parameters that compatible services vary: ```text -select().from(table).where(eq(...)) -insert(table).values(...) -delete(table).where(eq(...)) +endpoint +bucket +region +addressing: path | virtual +headers +copy: boolean +conditionalWrite: boolean +partSize +copyPartSize +concurrency +credentials ``` -That choice keeps the integration usable across Drizzle database objects that expose this common surface. It also means record replacement is delete-then-insert rather than dialect-specific upsert SQL. +The safe rule is to read the selected provider's current primary documentation and enable only the capabilities it actually +implements for the operations used by the filesystem. + +A few current examples show why this matters: -### Concurrency consequence +| Provider family | Current nuance that affects this client | +| --- | --- | +| AWS S3 | baseline SigV4, multipart upload, CopyObject, UploadPartCopy, conditional completion | +| Cloudflare R2 | S3-compatible endpoint with its own supported-operation set; `auto` is a documented region value | +| DigitalOcean Spaces | implements a documented subset of the S3 API and its own published object/multipart limits | +| Google Cloud Storage XML API | S3-compatible multipart exists, but documented multipart precondition behavior differs | +| Backblaze B2 S3 API | S3-compatible surface with its own unsupported/changed AWS control-plane features | + +For a Google Cloud Storage XML multipart path that does not support the preconditions expected by optimistic object +read-modify-write, create the client with `conditionalWrite: false`. That does not make append/update magically atomic; it makes +the absence of that safety property explicit. + +Cloudflare R2 and other services can also disable `copy` if their selected endpoint/path does not provide the server-side copy +contract expected by the adapter. The filesystem then falls back to the normal streamed/materialized copy path instead of +calling a native capability that was never real. -Inside one `FileSystemType` with normal coordination, same-path mutations are serialized. Across multiple processes, hosts, or independently configured applications, delete-then-insert is not an atomic database transaction. +The S3 request escape hatch is intentional +----------------------------------------- -If cross-process atomic replacement is required, provide a database-level transaction/serialization strategy appropriate for the actual Drizzle dialect and driver. The package does not hide that requirement behind a false portability claim. +A filesystem does not need to model every S3 object or bucket feature. The direct client therefore exposes signed +`request(options)` in addition to `ObjectStoreType`. -### Drizzle driver breadth +Use the filesystem/object layer for portable file behavior. Use the lower-level request API when the application needs a +provider-specific control such as an object-lock header, tag operation, checksum policy, versioning call, or another S3 operation +whose semantics should not be flattened into a generic filesystem method. -The current Drizzle source tree contains dedicated runtime/dialect integrations such as AWS Data API, better-sqlite3, Bun SQL, Bun SQLite, Cloudflare D1, Durable SQLite, Expo SQLite, and many additional PostgreSQL/MySQL/SQLite-family drivers. The package's compatibility condition is the database object's CRUD surface and the caller's correct table schema, not a hard-coded driver-name allowlist. +The same principle applies to Azure Blob +---------------------------------------- -Choosing between the ecosystem bridges --------------------------------------- +Azure Blob has enough differences that the package implements a native Azure REST client rather than translating Azure through +an S3 compatibility layer. -Use the bridge for the abstraction your application already owns. +`createAzureClient()` supports SAS, Microsoft Entra bearer tokens, Shared Key, and custom-header credentials. Its service version is explicit. Its streamed +writes use Azure block APIs, and its large server-side copies use Put Block From URL when synchronous Copy Blob From URL is too +small. -| Existing application resource | Use | +The object adapter above Azure is still the same `createObjectAdapter()` used by S3. The provider client owns Azure-specific +HTTP mechanics; the object adapter owns the file/directory translation. + +Choose the integration that matches the resource you already own +----------------------------------------------------------------- + +| Existing application resource | Preferred integration | | --- | --- | +| browser OPFS root | `createOpfsAdapter()` or `openFileSystem()` | +| host directory | Node, Deno, or Bun adapter | | unstorage `Storage` | `createUnstorageAdapter()` | | RxDB collection | `createRxDbAdapter()` | | db0 `Database` | `createDb0Adapter()` | +| connected SQLite statement API | `createSqliteAdapter()` | | Drizzle database + table | `createDrizzleAdapter()` | -| custom document/KV layer | `createRecordAdapter()` | -| host directory | Deno/Bun/Node adapter | - -Do not wrap a db0 Database in unstorage merely to reach this package if the application already owns db0 directly. Each extra storage layer adds semantics and performance behavior that has to be understood. +| Deno KV database | `createDenoKvAdapter()` | +| localStorage/sessionStorage-like Web Storage | `createLocalStorageAdapter()` | +| IndexedDB | `createIndexedDbAdapter()` / `openIndexedDbAdapter()` | +| Cache Storage `Cache` | `createCacheAdapter()` | +| S3-compatible endpoint | `createS3Client()` + `createS3Adapter()` | +| Azure Blob container | `createAzureClient()` + `createAzureAdapter()` | +| custom KV/document layer | `createRecordAdapter()` | +| custom object storage | `createObjectAdapter()` | +| any `FileSystemType` needed as KV | `createKeyValueDriver()` | +| any `FileSystemType` needed by unstorage | `createUnstorageDriver()` | + +Adding an extra ecosystem layer only because this package already has an adapter for it usually makes the system harder to +reason about. Use the direct adapter/client/driver for the abstraction the application already owns, and use bridge descriptors +when code needs to inspect both directions as one integration. diff --git a/docs/environments.md b/docs/environments.md index c2494dd..836078e 100644 --- a/docs/environments.md +++ b/docs/environments.md @@ -1,33 +1,33 @@ Execution environments ====================== -`@okikio/opfs` separates the frontend filesystem model from backend availability. This lets the same core APIs compile in Window, WebWorker, Deno, Bun, and Node TypeScript targets while concrete adapters stay on explicit runtime subpaths. +`@okikio/opfs` keeps the filesystem frontend separate from backend availability. The same core source can compile for Window, +workers, Deno, Bun, and Node while runtime-specific adapters remain on explicit subpaths. -Browser OPFS ------------- +The package does not maintain a browser-brand or runtime-brand behavior table. It asks the current realm or adapter what it can +actually do and preserves the resulting capability/failure information. -The native OPFS adapter requires `navigator.storage.getDirectory()`. +Browser OPFS follows the storage key of the current realm +--------------------------------------------------------- -The package does not select behavior from a browser brand table. It probes the APIs the current realm exposes and preserves native failures when the browser denies storage. +Native browser OPFS requires `navigator.storage.getDirectory()`. -### Window - -Use the full async facade. +In Window, use the asynchronous facade: ```ts const fileSystem = await openFileSystem(); await fileSystem.writeFile("/state.json", "{}", { parents: true }); ``` -Do not assume `createSyncAccessHandle()` is available in Window. The sync API is capability-gated. - -### DedicatedWorker +Do not assume synchronous access from Window. `openSyncFile()` succeeds only when the actual native file handle exposes the sync +handle API and the adapter reports that capability. -The async facade works when OPFS is exposed. A DedicatedWorker is also the intended browser context for synchronous access handles in the File System standard. +DedicatedWorker is the important worker case because browsers commonly expose synchronous OPFS access there. The library still +probes the handle instead of saying "DedicatedWorker means sync": ```ts const capabilities = await probeOpfs(); -if (capabilities.syncAccessExposed) { +if (capabilities.syncAccessHandleExposed) { const file = await fileSystem.openSyncFile("/database.sqlite", { create: true, parents: true, @@ -40,138 +40,139 @@ if (capabilities.syncAccessExposed) { } ``` -### SharedWorker - -Use the async facade. Do not infer synchronous access only from the fact that code is running in a worker. The package checks the actual handle method. - -### ServiceWorker - -Use async filesystem methods and keep event lifetime explicit. +SharedWorker and ServiceWorker use the same asynchronous filesystem APIs when storage is exposed. A ServiceWorker must keep the +browser event alive itself. The filesystem cannot call `event.waitUntil()` on behalf of the application. ```ts self.addEventListener("message", (event) => { - const operation = (async () => { - const fileSystem = await openFileSystem(); - await fileSystem.writeFile("/events/latest.json", "{}", { - parents: true, - }); - })(); - - event.waitUntil(operation); + event.waitUntil(saveMessage(event.data)); }); ``` -The filesystem cannot extend a service-worker event lifetime on its own. - -Iframes -------- - -### Same-origin iframe - -Normal `openFileSystem()` opens the storage associated with the iframe's current storage key. - -### Third-party iframe - -Browser storage partitioning can give a third-party iframe storage isolated by the embedding site. Normal `openFileSystem()` deliberately does not attempt to escape that policy. - -For browsers that support the Storage Access API extension for OPFS, the separate iframe module can request unpartitioned access: - -```ts -import { - requestUnpartitionedFileSystem, - supportsUnpartitionedOpfsRequest, -} from "@okikio/opfs/iframe"; -``` - -The application must make the request from the appropriate user-activation/permission flow. The package never does it automatically. +Iframes need policy-aware tests, not a blanket promise +----------------------------------------------------- -### Opaque sandbox +A same-origin iframe normally observes storage under the same applicable storage key as its embedding context. -A sandboxed iframe without a usable origin can reject storage access. Treat `probeOpfs().rootAvailable` and the returned root error as the source of truth for that context. +A third-party iframe can be partitioned by browser privacy/storage policy. The package does not try to escape that policy +implicitly. The optional iframe API is separate because requesting unpartitioned storage, where the browser supports it, belongs +inside an explicit user-activation and permission flow. -Private browsing ----------------- - -Private/incognito modes can change storage quota, persistence, availability, or lifetime. The package does not fingerprint the browsing mode. - -The decision flow is: +An opaque sandbox can reject storage because it has no usable origin. `probeOpfs()` returns the actual root result and normalized +failure instead of inferring the outcome from the sandbox flag alone. ```text -probe actual capability - | - +-- root available -> use selected OPFS strategy - | - +-- root unavailable -> inspect normalized error - choose application fallback +iframe starts + | + v +probe actual storage API + | + +-- root opens ------> use selected strategy + | + `-- root rejected ---> preserve normalized reason -> application fallback ``` -This is more reliable than inferring behavior from a browser/private-mode label. - -`file:` documents ------------------ +Private browsing, packaged `file:` pages, enterprise browser policy, quota, and persistence can also change availability or +lifetime. The package deliberately does not fingerprint private mode or promise OPFS on `file:` URLs. Probe the realm you are +actually running in. -Current browser behavior for OPFS in `file:` documents is not fully interoperable. The WHATWG File System issue tracker contains an active request to specify this case more clearly. +Playwright tests the browser contexts directly +---------------------------------------------- -Do not promise OPFS availability for a packaged `file:` application. Probe the actual context. +The canonical browser suite uses Playwright Test projects for Chromium, Firefox, and WebKit. The test matrix exercises: -Web Locks ---------- +| Context or behavior | Chromium | Firefox | WebKit | +| --- | ---: | ---: | ---: | +| Window async OPFS | probe + execute | probe + execute | probe + execute | +| DedicatedWorker async | probe + execute | probe + execute | probe + execute | +| DedicatedWorker sync handle | probe + open | probe + open | probe + open | +| SharedWorker | probe + execute | probe + execute | probe + execute | +| ServiceWorker | black-box + instrumentation | black-box | black-box | +| same-origin iframe | probe + execute | probe + execute | probe + execute | +| cross-origin iframe | observe policy result | observe policy result | observe policy result | +| opaque sandbox | observe policy result | observe policy result | observe policy result | +| fresh context isolation | execute | execute | execute | +| persistent profile reopen | execute | execute | execute | +| abort before commit | execute | execute | execute | +| localStorage / IndexedDB / Cache adapters | execute | execute | execute | -When `coordination: "auto"`, the facade uses Web Locks if `navigator.locks` is available. This can coordinate cooperating tabs and workers that share the same origin and lock names. +Playwright's deeper ServiceWorker inspection is Chromium-specific, so only that instrumentation is browser-specific. The actual +ServiceWorker OPFS behavior stays a black-box page-to-worker message test in every browser that exposes the API. -If Web Locks are unavailable, `auto` falls back to one-realm FIFO locks. That fallback cannot coordinate another tab, worker realm, or OS process. +Deno keeps the same library contracts with runtime-specific capabilities +------------------------------------------------------------------------ -Deno ----- +Use `@okikio/opfs/adapter/deno` for a host directory. The runtime needs the filesystem permissions required by the configured +root. The adapter does not request permissions or inspect environment variables itself. -Use `@okikio/opfs/adapter/deno`. +Deno KV is a separate adapter because it is a record store, not a host filesystem: ```ts -const fileSystem = createFileSystem( - createDenoAdapter({ root: "./data" }), - { coordination: "local" }, -); -``` +import { createFileSystem } from "@okikio/opfs"; +import { createDenoKvAdapter } from "@okikio/opfs/adapter/deno-kv"; -The runtime needs filesystem permissions appropriate for the configured root. The adapter does not request broader permissions or inspect environment variables itself. +const kv = await Deno.openKv("./data.kv"); +const fileSystem = createFileSystem(createDenoKvAdapter(kv)); +``` -Native move and sync random access are available through Deno filesystem APIs. +Current Deno documentation still marks KV as unstable. Real Deno KV tests therefore use `--unstable-kv`. The production adapter +accepts a structural KV contract, so simply importing the module does not require a global `Deno` object. -Bun ---- +Deno KV also has a much smaller physical value limit than an ordinary filesystem file. The adapter exposes that limit and a +partition policy through `inspect()`. With the default `partition: "auto"`, small materialized files stay inline while large +files and unknown-size replacement streams use a manifest plus raw byte parts. Callers that need a one-record layout can set +`partition: "never"`; the adapter then stops advertising its partitioned replacement-stream lane and large values fail +explicitly instead of changing layout. -Use `@okikio/opfs/adapter/bun`. +Node and Bun use explicit host adapters +--------------------------------------- -The adapter uses Bun's file primitives where they provide a clear benefit and Bun's Node-compatible filesystem surface for directory/update/sync operations. +The Node adapter uses `node:fs` and `node:fs/promises`. It supports native streaming reads and writes, ranges, copy, rename, +synchronous random access, and flush. The configured host root is the only host directory intentionally exposed through the +virtual path namespace. -The root entrypoint never imports the Bun adapter, so browser code does not evaluate Bun-specific code accidentally. +The Bun adapter uses Bun file APIs for the direct byte path and Bun's Node-compatible filesystem APIs for directory, update, +copy/move, and synchronous host-file behavior. The same public tests import `node:test`; Bun currently supports the in-process +`node:test` API when those files are run with `bun test`. -Node ----- +The Bun benchmark keeps two raw file-copy baselines: Node-compatible `copyFile()` and `Bun.write(destination, Bun.file(source))`. +The second path lets Bun select its file-backed Blob fast path. A Bun-only provider benchmark also compares Bun's native +`S3Client` with the AWS SDK baseline, this package's direct SigV4 client, the object adapter, and the filesystem facade. These +measurements are evidence for route selection; they do not make runtime brand part of the portable API contract. -Use `@okikio/opfs/adapter/node`. +Electron can use the Node adapter in a trusted main-process layer. The OPFS-shaped API is not a reason to expose an arbitrary host +root directly to untrusted renderer content. Use an application-specific IPC/service contract or browser OPFS where that matches +the trust model. -The adapter uses `node:fs` and `node:fs/promises`. It supports streaming, range reads, rename, synchronous random access, and flush. +Object clients are runtime-neutral Web clients, but deployment policy still matters +----------------------------------------------------------------------------------- -The configured host root is the only directory intentionally exposed through the virtual namespace. +The S3 and Azure clients use Web Fetch, Web Crypto where signing is required, Web Streams, AbortSignal, and focused `@std/*` +packages for concurrency, byte assembly, stream limits, path mapping, and XML. They are +therefore usable across Deno, Bun, Node, browsers, and workers that expose those Web APIs. -Electron --------- +That does not make every deployment equally appropriate. -The Node adapter is suitable for a trusted Electron main-process filesystem layer. Do not expose arbitrary host roots directly to untrusted renderer content merely because the frontend API resembles OPFS. +A browser calling S3 directly needs CORS rules that permit the required methods and headers. More importantly, long-lived cloud +storage secrets should not be shipped to untrusted browser code. Use short-lived scoped credentials or a trusted server design. -A renderer can instead communicate with a controlled main-process service or use browser OPFS where that matches the application model. +The same applies to Azure. SAS tokens and bearer tokens should be scoped to the actual client threat model. The library accepts a +refresh function so a long-lived process does not have to freeze one credential at client creation. -Record/database backends ------------------------- +Provider endpoints can also have runtime-specific network rules. A Cloudflare Worker, browser extension, serverless host, or +corporate browser policy can allow or reject different origins. Those network policies are outside the filesystem abstraction. -RxDB, unstorage, db0, and Drizzle integrations work in any runtime where the injected upstream resource works and where the package's core Web APIs (`ReadableStream`, `Blob`, `File`, `TextEncoder`, AbortSignal) are available. +Coordination scope is part of the execution environment +------------------------------------------------------- -Their filesystem file bodies are record-backed, so stream writes are bounded-buffer operations rather than native streaming. +`web-locks` can coordinate cooperating browser realms that share the relevant Web Locks namespace. `local` only coordinates one +JavaScript realm. Separate OS processes and hosts need backend-level coordination where same-path atomicity matters. -Server coordination -------------------- +Database/object adapters can use provider transactions, ETags, versions, advisory locks, or leases where the provider exposes +them. The facade does not pretend an in-memory lock became distributed merely because the persisted bytes live on a remote +service. -`coordination: "local"` coordinates only within one JavaScript realm. It does not serialize writes across separate Node/Deno/Bun processes or separate hosts. -Database-backed applications that need cross-process same-path atomicity must use backend-level transactions, leases, advisory locks, or another coordination mechanism suitable for that provider. +Provider protocol details are kept in [S3 client protocol](./s3.md), [Azure Blob client protocol](./azure.md), and the +[provider integration test guide](./providers.md). Shared Key is intended for trusted server/Azurite contexts because it exposes +the Azure storage account key to the runtime. diff --git a/docs/providers.md b/docs/providers.md new file mode 100644 index 0000000..89c2590 --- /dev/null +++ b/docs/providers.md @@ -0,0 +1,225 @@ +Object-provider integration tests +================================= + +Purpose +------- + +The direct S3 and Azure clients own HTTP protocol behavior that a pure mock +cannot prove. This guide explains the local provider environment used to test +real signing, request routing, range transfer, multipart/block state, copy, +listing, and filesystem translation without requiring cloud credentials for +every maintainer run. + +The provider environment supplements deterministic unit tests. It does not +replace Amazon S3 or Azure specifications. + + +Local topology +-------------- + +`tests/provider/compose.yml` starts two independent object services: + +```text +Deno test process + | + +--> http://127.0.0.1:8333 + | SeaweedFS S3 endpoint + | bucket: opfs-test + | SigV4 credentials: admin / secret + | + `--> http://127.0.0.1:10000/devstoreaccount1 + Azurite Blob endpoint + Shared Key development account +``` + +The compose file pins explicit provider versions so a maintainer does not get a +silent protocol change because `latest` moved: + +```text +chrislusf/seaweedfs:4.41 +mcr.microsoft.com/azure-storage/azurite:3.36.0 +``` + +SeaweedFS is used as an actively maintained independent S3-compatible server. +Its role is interoperability testing. Amazon S3 documentation remains the +source of truth for AWS-specific request behavior. + +Azurite is Microsoft's local Azure Storage emulator. It runs only the Blob +service for this test matrix, with telemetry disabled and in-memory persistence. +`--skipApiVersionCheck` allows the direct client to exercise the configured +current REST version even when the emulator has not yet added an identical +version allow-list. This option does not make Azurite behavior identical to the +Azure cloud service. + + +Run the provider suite +---------------------- + +The canonical maintainer command is: + +```sh +mise run test-providers +``` + +The task performs this lifecycle: + +```text +docker compose up -d + | + v +poll S3 and Azure HTTP endpoints + | + v +deno ci frozen-lock dependency install + | + v +deno test tests/provider.test.ts + | + v +always docker compose down -v +``` + +Cleanup is registered before readiness polling. A failing test therefore does +not intentionally leave provider volumes or containers behind. On readiness +failure, the task prints compose logs before cleanup so the environment failure +remains diagnosable. + +The GitHub Actions `providers` job calls the same mise task. CI does not carry a +second provider-startup implementation. + + + +Provider benchmarks compare equivalent layers +-------------------------------------------- + +The provider environment also backs an explicit benchmark matrix: + +```sh +mise run bench-providers +``` + +S3 is measured as: + +```text +AWS SDK v3 baseline +Bun native S3Client baseline +@okikio/opfs direct SigV4 client +direct ObjectStore adapter +FileSystemType with metrics none +FileSystemType with metrics basic +``` + +Azure is measured as: + +```text +@azure/storage-blob baseline +@okikio/opfs direct Azure REST client +direct ObjectStore adapter +FileSystemType with metrics none +FileSystemType with metrics basic +``` + +Write cases include the same post-write metadata request when the project client returns verified object metadata. Multipart is a +separate benchmark from a small replacement because request-count and commit topology are different. This avoids presenting an +SDK/client request-plan difference as facade overhead. + +The Bun provider run is separate because its native S3 implementation is runtime-specific. The local Bun filesystem benchmark +also compares Node-compatible `copyFile` with `Bun.write(destination, Bun.file(source))`; the project does not switch the adapter +copy path until measurements justify the change. + +These loopback benchmarks are overhead diagnostics, not cloud throughput claims. They should be repeated against controlled real +provider environments before using them to choose production concurrency, retry, or part sizes. + +What the live tests prove +------------------------- + +The S3 path validates: + + - a real SigV4 HTTP request accepted by an independent S3-compatible server; + - PUT and HEAD; + - byte-range GET; + - create-only conditional replacement; + - multipart stream upload with a legal non-final part size; + - provider-side copy; + - prefix listing; + - delete cleanup; + - `ObjectStoreType -> AdapterType -> FileSystemType` translation. + +The Azure path validates: + + - Shared Key accepted by Azurite; + - explicit container creation; + - Put Blob and Get Blob Properties; + - byte-range GET; + - create-only conditional replacement; + - Put Block / Put Block List streaming upload; + - same-account server-side copy; + - prefix listing; + - delete cleanup; + - `ObjectStoreType -> AdapterType -> FileSystemType` translation. + +These are end-to-end HTTP tests. The provider receives the actual headers, +query fields, body bytes, XML, and signatures created by the library. + + +What the live tests do not prove +-------------------------------- + +A provider emulator/compatible server cannot prove all details of a cloud +service. The suite does not use it as an oracle for: + + - exact AWS canonical-request text; + - AWS-only embedded error bodies returned with HTTP 200; + - every S3 service limit; + - Azure Shared Key string-to-sign construction independent of Azurite; + - every historical Azure service-version size limit; + - cloud identity/role acquisition; + - region routing and redirects; + - provider throttling behavior; + - cloud durability or consistency guarantees; + - billing, lifecycle, retention, encryption, replication, or versioning. + +Those behaviors are covered by deterministic protocol tests where possible and +remain candidates for opt-in real-cloud suites. + + +Why the matrix keeps both test styles +------------------------------------- + +A mock can assert an exact canonical signature, but it can accidentally accept a +request no real server would parse. A container can prove the request works, +but an emulator can also be more permissive than the cloud service. + +The two test styles therefore protect different failure classes: + +| Test style | Strong at | Weak at | +| ---------- | --------- | ------- | +| Deterministic request test | Exact canonical text, headers, limits, branch selection | Real HTTP parser/auth integration | +| Local provider container | Real socket/HTTP/auth/protocol interoperability | Complete cloud parity | +| Optional real cloud | Actual provider behavior | Cost, credentials, availability, reproducibility | + +A client change that affects signing, multipart/block state, copy, conditional +writes, or provider errors should add or update the deterministic test **and** +the provider test when the behavior is supported by the local implementation. + + +Future provider breadth +----------------------- + +S3 compatibility should eventually be tested against more than one independent +implementation when that adds a materially different contract. Useful future +candidates include Cloudflare R2, Backblaze B2 S3, DigitalOcean Spaces, and a +real Amazon S3 bucket through opt-in CI. These should not all become mandatory +Docker services merely to increase a provider count. + +The selection criterion is behavioral diversity: + + - different addressing requirements; + - missing copy or conditional-write behavior; + - different multipart error behavior; + - non-AWS region/signing expectations; + - list/pagination differences that affect the portable contract. + +Azure should similarly gain an opt-in cloud test for the newest service version. +Azurite remains valuable because it gives every contributor a deterministic +Shared Key integration without cloud credentials. diff --git a/docs/releasing.md b/docs/releasing.md new file mode 100644 index 0000000..fa75192 --- /dev/null +++ b/docs/releasing.md @@ -0,0 +1,111 @@ +Release process +=============== + +`@okikio/opfs` publishes one Git release to JSR and npm. Git history decides the version. `deno.json` owns the authored package graph and public exports. semantic-release does not commit generated version files back to `main`. + +Conventional commits +-------------------- + +Commits merged to `main` use Conventional Commits 1.0.0. Release-relevant examples are: + +```text +fix: correct a public behavior -> patch +feat: add a compatible capability -> minor +feat!: replace a public contract -> major + +BREAKING CHANGE: describe the consumer impact -> major +``` + +`build`, `chore`, `ci`, `docs`, `refactor`, `style`, and `test` are valid types. They do not cause a release unless the analyzed commit carries a breaking change. Pull requests should normally squash to one conventional commit so `main` remains a clear release input. + +The checked-in `0.0.1` is development metadata, not the first public release decision. semantic-release does not support selecting `0.0.1` as the initial stable version. With no earlier release tag, the normal first release on `main` is `1.0.0`. If the project is not ready for a stable `1.0.0`, configure a prerelease branch before enabling the release workflow. Registry commands receive the semantic-release version through `--set-version`, so they do not publish the development placeholder. + +Dependency graph +---------------- + +Dependency changes are not complete until both package-manager views are reproducible. After changing `deno.json` or +`package.json`, update the Deno lockfile intentionally with: + +```sh +deno install --frozen=false +pnpm install +``` + +Review and commit both `deno.lock` and `pnpm-lock.yaml`. CI uses `deno ci`, which deliberately rejects a missing or stale Deno +lockfile instead of resolving an unseen dependency graph during the build. + + +Release flow +------------ + +```text +push to main + | + v +CI + | frozen dependencies, format, lint, docs, types, + | runtime tests, browser matrix, stress, package dry-runs + v +Release workflow + | + +-- verify main still equals the tested CI SHA + +-- semantic-release reads commits since opfs@ + +-- semantic-release creates the Git tag and GitHub Release + | + `-- only when a new tag appeared: + workflow_dispatch Publish Registries + | + +----> JSR + `----> npm +``` + +The registry workflow is dispatched explicitly instead of relying on the GitHub `release` event. A release created with the repository `GITHUB_TOKEN` does not normally trigger another workflow from that release event. GitHub does allow `workflow_dispatch` events created with `GITHUB_TOKEN`, so the release workflow can start the separate publisher without a long-lived release PAT. + +A rerun on a commit that already has an `opfs@` tag does not dispatch another automatic publication. Manual registry retries remain available through `Publish Registries` and require the existing immutable tag. + +semantic-release owns only version analysis, tag creation, release notes, and the GitHub Release. It does not publish either registry and it does not create a release commit. + +npm packaging +------------- + +JSR consumes the authored TypeScript package directly. npm receives generated JavaScript and declaration files. + +`.mise/tasks/npm` runs `deno pack --no-deno-shim --set-version`. Deno derives the npm graph and exports from the same `deno.json` used by JSR. Drizzle needs one npm-only correction: it is an optional integration, so the task removes `drizzle-orm` from generated normal dependencies and writes it as an optional peer before the final `npm pack`. + +The generated tarball is then installed into a clean consumer. The verifier checks exports and declarations, Node/Deno/Bun imports when those runtimes are present, browser bundling, and the optional Drizzle subpath after the peer is installed. + +Registry setup +-------------- + +Before the first release: + +1. Create `@okikio/opfs` on JSR and link it to `okikio/opfs` for GitHub OIDC publication. +2. Bootstrap the first npm publication with a granular/automation token in `NPM_TOKEN` when trusted publisher settings are not yet available for the package. +3. After npm has the package, configure trusted publishing for `.github/workflows/publish.yml`. +4. Remove `NPM_TOKEN` when the bootstrap fallback is no longer wanted. + +The publisher uses GitHub OIDC for JSR and for normal npm trusted publication. npm is upgraded to a current npm 11 release in the publish job before trusted publishing. + +Partial failure +--------------- + +If one registry succeeds and the other fails, do not create another version. Run `Publish Registries` manually with the same existing `opfs@` tag and select only the failed registry. + +Local release checks +-------------------- + +A Deno-capable machine can run: + +```sh +mise install +mise run check +mise run test +mise run bench +deno task release:check +``` + +The GitHub release and registry workflows use mise for their Node, Deno, and Bun toolchain too. Release-only policy remains in +GitHub Actions: OIDC permissions, immutable tag resolution, semantic-release, registry authentication, and the final publish +commands are deployment concerns rather than reusable repository tasks. + +`deno task release:check` performs package dry-runs. It does not upload a release. diff --git a/docs/s3.md b/docs/s3.md new file mode 100644 index 0000000..673a9ae --- /dev/null +++ b/docs/s3.md @@ -0,0 +1,512 @@ +S3 client protocol guide +======================== + +Purpose +------- + +This document defines the S3 protocol contract implemented by +`@okikio/opfs/s3`. It is written for maintainers who need to change signing, +upload, copy, listing, conditional-write, or compatibility behavior without +silently changing the public filesystem guarantees. + +The client is intentionally not a replacement for the complete AWS SDK. It +implements the S3 REST operations needed by the object-store adapter and keeps a +low-level signed `request()` method for S3-compatible features that do not belong +in the portable filesystem API. + +The implementation is direct: + +```text +ObjectStoreType / S3ClientType + | + v +request construction + | + v +AWS Signature Version 4 + | + v +Web Fetch + | + +--> Amazon S3 + `--> S3-compatible endpoint +``` + +The protocol code uses Web Crypto for SHA-256 and HMAC-SHA256. `@std/encoding` +owns hexadecimal encoding, `@std/async/pool` owns bounded multipart concurrency, +`@std/xml` owns XML parsing/serialization, and `@std/bytes` supports the shared +chunk layer. No AWS SDK package is imported. + +This document distinguishes three evidence classes: + + - **Implemented** means the current source contains the behavior. + - **Protocol** means the behavior is required or described by current AWS S3 + documentation. + - **Provider-dependent** means an S3-compatible service can intentionally + differ and the client exposes configuration for that difference. + +The current implementation was reviewed against AWS S3 REST documentation on +August 14, 2026. The authoritative source remains AWS documentation when this +file and the service specification disagree. + + +The client owns a focused S3 contract +------------------------------------- + +`S3ClientType` extends the package `ObjectStoreType`, so it provides the object +operations the filesystem adapter needs: + +| Object operation | S3 REST operation | Current behavior | +| ---------------- | ----------------- | ---------------- | +| `head()` | `HeadObject` | Exact-key metadata or `null` for `404` | +| `get()` | `GetObject` | Complete object or one `Range` | +| `put(Uint8Array)` | `PutObject` | One materialized request up to 5 GB | +| `put(stream)` | Multipart upload | Bounded concurrent parts and explicit completion | +| `delete()` | `DeleteObject` | Missing object is already removed | +| `list()` | `ListObjectsV2` | Prefix, delimiter, limit, continuation token | +| `copy()` small | `CopyObject` | Server-side copy through the 5 GB single-copy limit | +| `copy()` large | Multipart `UploadPartCopy` | Server-side ranged copy without JS body transfer | + +The S3-specific surface also exposes: + +```text +request() +createUpload() +uploadPart() +completeUpload() +abortUpload() +``` + +Those operations are public because an S3 consumer can need storage class, +checksums, encryption, object lock, tagging, or another provider extension that +is not a filesystem concern. The low-level request method signs the request but +does not interpret every S3 feature on the caller's behalf. + + +Addressing and canonical request construction +--------------------------------------------- + +The client supports path-style and virtual-hosted-style addressing. + +Path style: + +```text +https://endpoint.example/bucket/path/to/object +``` + +Virtual-hosted style: + +```text +https://bucket.endpoint.example/path/to/object +``` + +`addressing: "path"` is the compatibility-oriented default because many local +and third-party S3 implementations expose one HTTP endpoint without wildcard +DNS for buckets. + +The client preserves slash separators in object keys while percent-encoding each +path segment. Signature Version 4 is sensitive to exact path and query +serialization. The implementation therefore does not use locale-sensitive +sorting and does not normalize the canonical URI after object-key construction. + +For the canonical query string, the client: + +1. expands repeated query values; +2. URI-encodes each name and value; +3. sorts the encoded names and then encoded values by code-unit order; +4. joins the pairs with `&`. + +For canonical headers, the client: + +1. lowercases header names; +2. normalizes internal whitespace; +3. sorts the signed header names; +4. includes the request authority as `host` even though browser Fetch controls + the actual Host or HTTP/2 `:authority` field; +5. includes required `x-amz-*` headers. + +The signing flow is: + +```text +method + + canonical URI + + canonical query + + canonical headers + + signed-header names + + payload SHA-256 + | + v +canonical request SHA-256 + | + v +AWS4-HMAC-SHA256 string to sign + | + v +kDate -> kRegion -> kService(s3) -> kSigning + | + v +Authorization signature +``` + +Credential sources can be static or refreshable. Refreshable credentials are +resolved immediately before each signed request. Session credentials add +`x-amz-security-token` before canonical signing. + +The client uses Web Crypto directly because browser-compatible SHA-256 and HMAC +are already platform APIs. `@std/crypto` does not implement S3 Signature +Version 4, so adding it would not remove protocol code or improve ownership. + +### Payload hashes + +Replayable Web request bodies are SHA-256 hashed before signing. The client +hashes strings, `ArrayBuffer`, `ArrayBufferView`, `Blob`, and +`URLSearchParams` values. Requests without bodies use the standard SHA-256 of +an empty payload. + +A low-level caller can supply `payloadHash`. This exists for S3 modes such as +`UNSIGNED-PAYLOAD` where the selected provider accepts that contract. The +client does not silently choose an unsigned payload when the exact request +bytes can be determined without consuming or re-encoding caller state. + +A streamed low-level body cannot be consumed once merely to calculate a hash and +then consumed again by Fetch. `FormData` has a related problem because Fetch +owns its multipart boundary and exact wire encoding. Those two low-level body +forms therefore use `UNSIGNED-PAYLOAD` unless the caller supplies an explicit +`payloadHash`. Callers that need AWS streaming-signature chunk framing must +implement that S3-specific streaming mode above `request()`. The normal +high-level streamed `put()` avoids this ambiguity by using multipart parts, +each of which is materialized before its individual signed request. + + +Object and multipart limits are part of planning +------------------------------------------------ + +`S3_LIMITS` records the protocol limits that affect this implementation: + +| Limit | Value used by the client | Why it matters | +| ----- | ------------------------ | -------------- | +| Maximum object | 53,687,091,200,000 bytes | Exact 10,000 x 5 GiB multipart ceiling (48.8 TiB) | +| Single `PutObject` | 5,000,000,000 bytes | Larger replacement must use multipart upload | +| Single `CopyObject` | 5,000,000,000 bytes | Larger copy must use multipart copy | +| Minimum multipart part | 5 MiB | Every non-final upload part must meet the S3 minimum | +| Maximum multipart part | 5 GiB | Client part-size configuration cannot exceed it | +| Maximum multipart parts | 10,000 | Known-size streams must choose a large enough part size | + +The distinction between GB and GiB is intentional. AWS documents the +single-request PUT/copy threshold in decimal GB, while multipart part limits use +binary-sized MiB/GiB values. AWS product documentation often calls the maximum +object size 50 TB, while the current object guide and multipart arithmetic make +the exact ceiling 10,000 x 5 GiB = 53,687,091,200,000 bytes (48.8 TiB, about +53.7 TB). The client uses the exact multipart-derived value so it does not +reject objects that S3 can legally assemble. + +`partSize` defaults to 8 MiB. `copyPartSize` defaults to 1 GiB. Both are +validated when the client is created. + +If `ObjectPutOptionsType.size` is supplied, the upload planner calculates the +minimum part size required to stay at or below 10,000 parts and uses the larger +of that value and the configured `partSize`. + +```text +known body size + | + v +ceil(size / 10,000) + | + +--> <= configured partSize -> configured partSize + | + `--> larger -> required part size + | + `--> reject if > 5 GiB +``` + +When a stream size is unknown, the configured part size remains authoritative. +The chunk iterator fails before part 10,001 instead of creating an invalid +multipart request. A caller with a very large stream should supply `size` so +the planner can choose a safe part size before network work begins. + +The final upload byte count is compared with the declared `size`. A mismatch is +a caller/data-source error and the multipart upload is aborted. + +Failure and cancellation do not transfer cleanup authority to the caller. Once +`CreateMultipartUpload` succeeds, the high-level streamed `put()` owns that +upload until completion or abort. If a part fails or the caller cancels, the +client stops consuming the source, waits for admitted part requests to settle, +and then sends `AbortMultipartUpload` with a separate cleanup signal. +`abortTimeoutMs` bounds that best-effort cleanup and defaults to 30 seconds. +Using a separate signal matters because the caller's cancellation signal is +already aborted at exactly the time cleanup becomes necessary. + + +Multipart upload is a commit protocol +------------------------------------- + +A streamed replacement follows four protocol stages: + +```text +CreateMultipartUpload + | + v +split stream into fixed-size owned byte parts + | + +--> UploadPart 1 + +--> UploadPart 2 bounded by `concurrency` + +--> ... + `--> UploadPart N + | + v +CompleteMultipartUpload + | + +--> success -> HEAD final object + | + `--> failure -> AbortMultipartUpload after active parts settle +``` + +`@std/async/pool` owns the concurrency admission. The client does not start an +unbounded Promise for every part. + +Each successful `UploadPart` must return an ETag. The completion document +contains one ordered `PartNumber`/`ETag` record per uploaded part. Before the +client sends that XML, it verifies: + + - at least one part exists; + - no more than 10,000 parts exist; + - every part number is an integer in the legal range; + - a part number is not duplicated; + - every ETag is non-empty; + - the final order is ascending by part number. + +The XML document is built with `@std/xml/stringify`; protocol escaping is not a +hand-written string replacement. + +The completion request can carry destination `If-Match`, `If-None-Match`, and +`x-amz-mp-object-size`. Preconditions belong to the commit stage rather than +multipart initiation because commit is the point where the destination object +becomes authoritative. + +S3 has an unusual completion failure mode: the service can send HTTP `200 OK` +before final assembly finishes and then put an `` document in the +response body. `completeUpload()` therefore parses the success body and treats +an embedded `` as a failed commit. + +If any streamed part operation fails, the client waits for the pool's already +started requests to settle before it sends `AbortMultipartUpload`. This order +prevents a late `UploadPart` from racing after the abort request. + +An abort failure does not replace the original upload failure. The original +operation remains the terminal error because it is what caused cleanup. +Unfinished multipart state can still remain at the provider when both the +operation and cleanup request fail, so production buckets should also use an S3 +lifecycle rule for stale multipart uploads. + + +Server-side copy has two distinct paths +--------------------------------------- + +`copy()` first performs `HeadObject` on the source. This confirms existence, +obtains size for planning, and supplies metadata needed by multipart initiation. + +For a source at or below the single-copy limit, the client sends `CopyObject`. +For a larger source, it creates a multipart upload at the destination and sends +one `UploadPartCopy` request per byte range. + +```text +HEAD source + | + +-- size <= 5 GB --------> CopyObject + | + `-- size > 5 GB ---------> CreateMultipartUpload + | + +--> UploadPartCopy range 1 + +--> UploadPartCopy range 2 + `--> ... + | + v + CompleteMultipartUpload +``` + +Multipart copy chooses a range size large enough to keep the destination at or +below 10,000 parts and rejects a plan that would require a part above 5 GiB. +The range is inclusive because `x-amz-copy-source-range` uses inclusive byte +positions. + +Source preconditions map to S3 copy-source headers: + +| Package option | S3 header | +| -------------- | --------- | +| `sourceIfMatch` | `x-amz-copy-source-if-match` | +| `sourceIfNoneMatch` | `x-amz-copy-source-if-none-match` | +| `sourceIfModifiedSince` | `x-amz-copy-source-if-modified-since` | +| `sourceIfUnmodifiedSince` | `x-amz-copy-source-if-unmodified-since` | + +Destination `ifMatch` and `ifNoneMatch` apply directly to `CopyObject` or to the +multipart completion request. They are deliberately removed from individual +`UploadPartCopy` requests because a destination object does not become the +completed value until commit. + +`CopyObject` and `UploadPartCopy` can also encode a service error inside an HTTP +200 response. Both paths use the same success-XML inspection as multipart +completion. + + +Listing preserves provider pagination +------------------------------------- + +`list()` uses `ListObjectsV2` with these mappings: + +| Package field | S3 query field | +| ------------- | -------------- | +| `prefix` | `prefix` | +| `delimiter` | `delimiter` | +| `limit` | `max-keys` | +| `cursor` | `continuation-token` | + +`Contents` records become `ObjectEntryType`. `CommonPrefixes` become child +prefixes. `NextContinuationToken` is returned as the next opaque cursor. + +The object-store filesystem adapter is responsible for consuming pages and +interpreting directory-marker metadata. The S3 client itself does not pretend +that prefixes are native directories. + + +Conditional writes are capability claims +----------------------------------------- + +The client advertises `conditionalWrite: true` by default because Amazon S3 +honors the preconditions used by this implementation. An S3-compatible service +that ignores or only partially implements these conditions must set +`conditionalWrite: false` in `createS3Client()`. + +That flag changes filesystem behavior. The object adapter will not claim that a +read-modify-write append/update is protected from concurrent replacement when +the backend cannot enforce the ETag precondition. + +`copy: false` similarly disables server-side copy for an endpoint whose S3 API +does not implement the required copy operations correctly. + +These overrides are explicit because compatibility means "uses the S3 protocol" +not "implements every Amazon S3 behavior". + + +Failures retain provider evidence +--------------------------------- + +Non-success responses are parsed as S3 XML when possible. `S3Error` retains: + +```text +status +S3 code +requestId +hostId +original Response +``` + +The original `Response` remains available so a caller can inspect headers and +provider-specific diagnostics that the stable error fields do not model. + +A response body that is not valid S3 XML still produces a failure based on HTTP +status and available text. The client does not convert an unknown provider +response into a fake known S3 error code. + +Cancellation uses `AbortSignal` on each Fetch request. Multipart cancellation +is not a distributed transaction: cancellation can stop local admission and +abort HTTP work, while an already accepted provider request may still have +created remote multipart state. Cleanup is therefore explicit. + +Request retry is explicit and operation-aware. `request` in `S3ClientOptionsType` configures retries, exponential delay, jitter, +and an optional per-attempt timeout. The implementation uses `@std/async/retry` for 408, 429, 5xx, and transport failures. +Authorization is rebuilt for every attempt so refreshable credentials and SigV4 timestamps are current. Signed redirects are +manual and are returned to the caller instead of being followed to another authority. + +Body replayability is only one admission condition. A one-shot `ReadableStream` receives one attempt. A mechanically replayable +body can still belong to a non-idempotent protocol operation, so low-level `request()` also accepts `retry: false`. The high-level +client disables automatic retry for `CreateMultipartUpload` and `CompleteMultipartUpload` because a lost response can make the +server-side outcome ambiguous. Stable part-number PUTs, reads, deletes, lists, and ordinary replacements use the configured +policy. `request: { retries: 0 }` disables automatic retry client-wide. + +`getMetrics()` returns direct HTTP request counts, retry counts, terminal failures, response counts, and optional Fetch duration. +Set `metrics: "none"` when measuring the raw protocol path, `basic` for counters, or `timing` for counters plus durations. + + +Provider compatibility and known non-goals +------------------------------------------ + +The direct client supports custom endpoint, region, headers, path/virtual +addressing, copy capability, and conditional-write capability. This is enough +to use Amazon S3 and many S3-compatible products while keeping compatibility +choices visible. + +The current client does **not** claim complete coverage of: + + - SigV4 streaming chunk signatures; + - presigned URL creation; + - S3 Express directory-bucket session management; + - access points, Object Lambda, or Outposts host construction; + - Multi-Region Access Point SigV4A; + - SSE-C/SSE-KMS convenience APIs; + - checksum negotiation beyond the payload hash needed for signing; + - object tagging, ACLs, retention, legal hold, replication, or lifecycle APIs; + - version-ID aware filesystem paths; + - bucket creation or bucket policy management; + - adaptive throttling and provider-specific `Retry-After` scheduling beyond the shared exponential retry policy. + +A low-level signed request can still reach some provider features when the +caller knows the exact S3 REST contract. A feature should receive a typed +high-level API only after the library can document and test its semantics. + + +Validation strategy +------------------- + +S3 validation is intentionally split into protocol and provider tests. + +`tests/s3.test.ts` is deterministic. It uses a controlled Fetch implementation +to inspect exact requests and covers: + + - Signature Version 4 canonicalization; + - deterministic timestamps and credentials; + - minimum/maximum multipart part configuration; + - multipart part ordering and duplicate rejection; + - `x-amz-mp-object-size` at completion; + - HTTP 200 embedded service errors; + - source and destination copy preconditions; + - large-copy multipart planning; + - part-count behavior. + +`tests/provider.test.ts` runs against the S3-compatible server in +`tests/provider/compose.yml`. It proves a real HTTP implementation can accept +our signed requests for PUT, HEAD, range GET, conditional create, multipart +streaming upload, copy, listing, delete, and the object-store filesystem adapter. + +The provider container is SeaweedFS, not a statement that SeaweedFS defines the +S3 specification. The container proves interoperability with one independent +S3-compatible implementation. AWS-specific wire details remain covered by the +deterministic protocol tests and the AWS documentation listed below. + +Before release, the maintainer test matrix should also run an opt-in real Amazon +S3 suite with short-lived credentials when CI secret policy permits it. That +suite must use a dedicated disposable bucket/prefix and explicit cleanup. + + +Primary specification sources +----------------------------- + +The implementation and this guide should be checked against these primary AWS +sources when S3 behavior changes: + + - AWS Signature Version 4 canonical request: + https://docs.aws.amazon.com/AmazonS3/latest/API/sig-v4-header-based-auth.html + - Multipart upload limits: + https://docs.aws.amazon.com/AmazonS3/latest/userguide/qfacts.html + - CompleteMultipartUpload: + https://docs.aws.amazon.com/AmazonS3/latest/API/API_CompleteMultipartUpload.html + - CopyObject: + https://docs.aws.amazon.com/AmazonS3/latest/API/API_CopyObject.html + - UploadPartCopy: + https://docs.aws.amazon.com/AmazonS3/latest/API/API_UploadPartCopy.html + - ListObjectsV2: + https://docs.aws.amazon.com/AmazonS3/latest/API/API_ListObjectsV2.html + +Secondary S3-compatible provider documentation can explain provider-specific +configuration, but it must not override the Amazon S3 wire contract when the +client claims Amazon S3 behavior. diff --git a/docs/sources.md b/docs/sources.md index 52d61a6..ff15bef 100644 --- a/docs/sources.md +++ b/docs/sources.md @@ -1,165 +1,375 @@ Research and source register ============================ -Research date: 2026-08-12. +Research date: 2026-08-15. -Source priority ---------------- +This register records the external contracts used to design and test the implementation. Source code and provider behavior can +change, so a release review should recheck current primary sources rather than assuming this date remains current. -When sources disagree, use this order: +Use this authority order when sources disagree: -1. current standards and current upstream source contracts; -2. current `okikio/mediad` repository rules and current Kaiju Platform/Crawl architecture guides; +1. current standards and current upstream source/contracts; +2. current repository implementation rules and project architecture guides; 3. current package implementation and tests; -4. older experiments and secondary articles. +4. older experiments and secondary performance reports. -The old `okikio/testing-opfs` experiment was reviewed for intent only. It is not an implementation base. +The package intentionally distinguishes implemented behavior from provider claims and proposals. A compatibility note in this +file is not evidence that an adapter passed a live integration test against that provider. -Browser File System / OPFS --------------------------- +Browser File System and OPFS +---------------------------- -Primary standards and interoperability sources: +Primary sources: - WHATWG File System Standard: - WHATWG File System issues: -- Web Platform Tests File System suite: -- WPT interoperability issue supplied for review: -- MDN Origin Private File System overview: -- web.dev OPFS article: +- Web Platform Tests File System results: +- WPT interoperability tracking: +- MDN Origin Private File System overview: + +- web.dev OPFS overview: -Important design facts traced into code/tests: +The implementation follows these observed design facts: -- normal file/directory handle operations are asynchronous; -- synchronous access handle exposure is context/capability-specific and can hold native file locks; -- writable streams and sync files have explicit close/abort lifecycle; -- current portable OPFS does not provide the same universal native rename contract as a host filesystem; -- error names, locks, storage policy, private browsing, iframe partitioning, and `file:` documents contain interoperability details that must not be hidden by browser-name guessing. +- asynchronous file/directory handles are the portable frontend; +- sync access is context/capability-specific and owns a native file lock for its lifetime; +- writable streams and sync handles have explicit close/abort lifecycle; +- browser storage policy can change availability, partitioning, persistence, quota, and error shape; +- third-party/opaque iframe behavior must be tested from the actual context; +- portable OPFS does not imply the same native rename model as Node/Deno host filesystems. -Additional OPFS material supplied by the user and reviewed for behavior/performance context: +Secondary OPFS/performance context reviewed earlier in the project: -- - - - -These secondary/performance sources informed test cases and tradeoffs. They do not override the standard or current upstream contracts. +These sources informed performance questions. They do not override the File System Standard or real browser tests. -Deno standard filesystem ------------------------- +Playwright +---------- + +Primary documentation: + +- Browsers: +- Test projects: +- Browser contexts/isolation: +- BrowserType persistent contexts: +- Frames: +- Service workers: +- Test configuration and webServer: + +The canonical browser test architecture uses Playwright Test for Chromium, Firefox, and WebKit. Chromium gets deeper +ServiceWorker instrumentation because Playwright documents that inspection surface as Chromium-specific; observable +ServiceWorker behavior remains a black-box test in the other browsers. + +Deno, Node, and Bun +------------------- + +Primary runtime sources: + +- Deno testing: +- Deno Node compatibility: +- Deno API reference: +- Deno KV API: +- Bun Node compatibility: +- Bun benchmarking guidance: +- Node documentation: + +Current Deno documentation treats `node:test` as a first-class test API and currently marks Deno KV unstable. The real Deno KV +suite therefore uses `--unstable-kv` without making that flag part of unrelated source imports. Current Deno KV documentation +states a 2 KiB serialized key limit, a 64 KiB serialized value limit, 1,000 mutations per atomic operation, and an 800 KiB total +atomic-operation limit. The Deno KV adapter exposes these constraints and uses a configurable manifest/part layout instead of +pretending one logical file must fit in one 64 KiB value. + +Current Bun compatibility documentation says its in-process `node:test` API works when files run under `bun test`, while some +advanced Node test-runner/reporting features remain incomplete. The repository uses the common `describe`/`it`/hooks subset and +keeps the test API itself as `node:test`. + +Bun's current File I/O documentation says `Bun.write(destination, Bun.file(source))` selects fast platform system calls for +file-to-file copies. The current Bun Rust source also keeps file-backed Blob state distinct so file-to-file paths can avoid a +naive user-space read/write loop. The benchmark therefore compares Bun's direct copy shape with Node-compatible `copyFile` +before changing the adapter implementation. + +Bun's S3 documentation exposes `S3Client`, `S3File`, `write`, `stat`, `stream`, and multipart `writer()` APIs. A Bun-only provider +benchmark now uses that native implementation as a second S3 baseline beside AWS SDK v3. The project does not treat Bun main +branch implementation work as proof about a released runtime; the mise pin remains the current released version selected by the +repository until a deliberate toolchain update. + +Deno standard libraries and Standard Schema +------------------------------------------- + +Primary sources reviewed from the current `denoland/std` repository and JSR +packages: + + - `@std/async`: + - `@std/bytes`: + - `@std/encoding`: + - `@std/expect`: + - `@std/fs`: + - `@std/http`: + - `@std/path`: + - `@std/streams`: + - `@std/xml`: + - `@std/crypto`: + - Standard Schema: + - Zod 4: + +The review was operation-led. A standard package replaces project code only +when its contract matches the filesystem or provider requirement without hiding +a stronger invariant. + +`@std/async/pool` owns bounded multipart and block concurrency. The stable +`pooledMap()` contract limits active requests and lets already-started requests +settle after one item fails. S3 cleanup waits for that settlement before it +sends `AbortMultipartUpload`, so a late part cannot arrive after the cleanup +request. + +`@std/async/retry` owns the direct clients' exponential backoff, jitter, AbortSignal, and retriable-error loop. The protocol +layer still classifies whether a request may enter that loop. One-shot streams are not replayed, and S3 multipart initiation and +completion disable automatic retry because a lost response can make the remote lifecycle outcome ambiguous. The low-level +request APIs also expose `retry: false` for provider-specific operations. + +`@std/bytes/concat` owns byte-array concatenation used by bounded chunk +assembly. The package does not maintain another concatenation implementation. + +`@std/streams` owns bounded materialization through +`LimitedBytesTransformStream` and final stream collection through `toBytes()`. +The current `FixedChunkStream` API is still marked unstable, so fixed-size +provider chunks remain in the package's small streaming adapter until that +standard contract is suitable for a public dependency. + +`@std/encoding` owns Base64 and hexadecimal encoding through their direct +subpaths. S3 uses hexadecimal SHA-256 output, Azure Shared Key and block IDs +use Base64, and record stores use Base64 for portable byte persistence. + +`@std/path` owns host path normalization and resolution for the Deno, Bun, and +Node adapters. The OPFS virtual path model remains project-owned because it +rejects and normalizes a different namespace than an operating-system path. + +`@std/fs` was reviewed for copy, move, walk, ensure, and host filesystem +operations. Those are intentionally not used inside the primitive Deno/Bun/Node +adapters. The public filesystem facade already owns recursive copy/move/walk, +overwrite, cancellation, and adapter-neutral semantics. Calling `@std/fs` from +one host adapter would duplicate that layer and introduce host-only symlink and +filesystem assumptions. `@std/path`, by contrast, directly replaces custom host +path manipulation without changing facade semantics. + +`@std/http/etag` was reviewed for conditional request support. The clients keep +provider ETags opaque instead of generating or evaluating them locally. S3 +multipart ETags and Azure ETags are provider tokens, not hashes that this +library should reinterpret. The package therefore forwards `If-Match` and +`If-None-Match` values to the provider rather than applying `@std/http/etag` in +the client. The unstable HTTP message-signature utilities also do not implement +AWS Signature Version 4 or Azure Shared Key. + +`@std/xml` owns provider control-document parsing and serialization. S3 list, +error, multipart, and copy responses and Azure list/error/block-list documents +use the standard XML tree instead of regular expressions or hand-written XML +escaping. Storage payloads themselves do not pass through XML parsing. + +`@std/crypto` was reviewed but is not used for provider signing. Web Crypto +already exposes browser-compatible SHA-256 and HMAC-SHA256, while the standard +crypto package does not implement AWS Signature Version 4 or Azure Shared Key +canonicalization. Adding it would introduce a wrapper without removing the +protocol code that actually carries the risk. + +`@std/expect` remains the assertion API on top of `node:test`. Zod 4 implements +Standard Schema, so the repository exports the Zod schemas directly instead of +maintaining a second validation wrapper for Standard Schema consumers. + + +S3 and Signature Version 4 +-------------------------- + +AWS primary references: + +- Signature Version 4 request authentication: + +- Signature Version 4 canonical request: + +- ListObjectsV2: +- CreateMultipartUpload: +- UploadPart: +- CompleteMultipartUpload: +- AbortMultipartUpload: +- CopyObject: +- UploadPartCopy: +- S3 multipart limits: -- Deno standard library repository: -- `@std/fs`: -- current `fs/mod.ts`, `fs/walk.ts`, `fs/copy.ts`, and `fs/move.ts` source were reviewed. +Implementation details derived from these contracts include: -Useful patterns retained: +- canonical signing includes `host` even though browser Fetch does not let application code set the Host header directly; +- multipart upload parts are bounded and the destination publishes on CompleteMultipartUpload; +- conditional `If-Match`/`If-None-Match` behavior belongs to multipart completion for the commit path used here; +- CompleteMultipartUpload can return an HTTP 200 response whose XML body later reports an error; +- CopyObject can also report an embedded error in an HTTP 200 response; +- CopyObject has a 5 GB source limit, so larger provider-side copies use UploadPartCopy; +- multipart uploads permit at most 10,000 parts and have defined part-size limits. -- lazy tree walking; -- explicit overwrite behavior; -- source/destination overlap checks; -- bounded, understandable helper APIs. +The package uses Web Crypto and Web Fetch rather than the AWS SDK so the direct client remains small, runtime-neutral, and +explicit about the S3 protocol surface it actually implements. -Native-host assumptions deliberately not copied into OPFS/record adapters: +S3-compatible providers +----------------------- -- symbolic links; -- host permission bits; -- OS path identity; -- portable timestamp mutation; -- universal native rename. +Provider-specific primary sources reviewed for compatibility differences: + +Cloudflare R2: + +- S3 API compatibility: +- release notes: + +DigitalOcean Spaces: + +- S3 compatibility: +- limits: +- direct API/SigV4: + +Google Cloud Storage XML API: + +- interoperability/migration: +- XML multipart uploads: + +Backblaze B2 S3-compatible API: + +- S3-compatible API: + +These providers illustrate why capability overrides exist. Endpoint, region, addressing, copy support, multipart preconditions, +checksum behavior, and unsupported control-plane operations can differ even when basic object requests use the S3 protocol. + +Azure Blob Storage +------------------ + +Microsoft primary references: + +- Azure Blob REST API: +- Shared Key authorization: +- Put Blob: +- Put Block: +- Put Block List: +- Copy Blob From URL: +- Put Block From URL: +- List Blobs: +- Versioning for Azure Storage services: + +The implementation keeps the service version explicit because accepted block sizes and Shared Key canonicalization depend on +the service version. Shared Key support starts at the augmented Blob format introduced in `2009-09-19`; zero-length +`Content-Length` signing changes after `2014-02-14`, and empty `x-ms-*` header canonicalization changes at `2016-05-31`. +Current copy behavior uses synchronous Copy Blob From URL for the smaller path and Put Block From URL ranges for large +provider-side copies. + +Unstorage +--------- + +Primary sources: + +- repository: +- current Driver/Storage contracts: +- current storage implementation: +- custom drivers: +- built-in driver catalog: + +The forward bridge targets `Storage`, not individual unstorage drivers. The reverse driver implements the stable Driver subset +needed for values, raw bytes, metadata, keys, clear, and disposal. `maxDepth` is advertised because the reverse driver applies +the depth filter itself. RxDB ---- +Primary sources: + - RxStorage guide: - RxStorage interface: - RxCollection implementation: -- RxDocument type contract: -The integration point is `RxCollection`, while RxDB retains responsibility for the chosen RxStorage implementation, wrappers, replication, multi-instance behavior, conflicts, and licensing. +The bridge accepts an RxCollection. RxDB retains responsibility for the selected RxStorage, replication, conflicts, +multi-instance behavior, wrappers, and licensing. -unstorage ---------- +db0 and Drizzle +--------------- -- repository: -- `src/types.ts` for `Storage` and `Driver` contracts -- generated `src/_drivers.ts` for current built-in driver inventory +Primary sources: -The forward bridge targets `Storage`. The reverse bridge implements the stable Driver subset used by unstorage. +- db0: +- db0 repository: +- Drizzle ORM: +- Drizzle repository: -db0 ---- +The db0 bridge targets the Database/dialect contract rather than connector names. Direct SQLite reuses that same record schema. +Drizzle keeps table/DDL ownership with the application because its schema builders and database behavior are dialect-specific. -- site: -- repository: -- `src/types.ts` for Database/Statement/dialect contracts -- generated `src/_connectors.ts` for current connector inventory +Upstream issue and pull-request review +-------------------------------------- -The adapter targets `Database` and its reported SQL dialect, not a connector name. +Current upstream issue/PR review was used to find failure modes that happy-path API docs do not reveal. The implementation does +not copy another library's behavior blindly; the issues are evidence for tests and invariants. -Drizzle -------- +Bun S3/Rust work reviewed included fixes for retry coverage, exponential backoff, timeouts, manual redirect handling, option +propagation, multipart abort on writer error, long SigV4 inputs, in-place multipart part assembly, XML parsing, proxy handling, +and worker-termination lifetime safety. The repeated lessons are: signed redirects must not be followed automatically, remote +cleanup has its own lifecycle, part concurrency needs a memory budget, and retry policy must not be inferred from body type alone. -- repository: -- current package metadata and `drizzle-orm/src` driver/dialect tree -- SQLite core database/query builder source for the common CRUD shape +AWS SDK v3 issues reviewed included very large upload memory growth, unknown-size multipart completion hangs, empty-stream lockups, +stream chunk-integrity regressions, conditional-header gaps in `lib-storage`, browser decompression/checksum mismatches, socket +exhaustion, and S3-compatible provider deserialization/endpoint regressions. The project benchmark keeps the AWS SDK as a +baseline while retaining a smaller direct protocol client with independently testable semantics. -Drizzle schema and DDL remain caller-owned because they are dialect-specific. The integration uses a caller-supplied table and common select/insert/delete builders. +Azure SDK issues reviewed included paused-stream abort hangs, invalid upload buffer arguments producing zero-byte blobs, large +buffer/block-size constraints, historical stream/file data corruption, copy polling request noise, and concurrency/default-size +questions. These reinforce explicit size/concurrency limits, bounded block admission, real abort tests, and provider request-count +benchmarks. -Mediad conventions ------------------- +Unstorage issues reviewed included non-atomic filesystem writes, S3 pagination/prefix bugs, XML entity decoding, file/prefix +collisions, SQL disposal, binary Redis storage, and Cloudflare Cache method binding. RxDB issues reviewed included OPFS/Expo file +truncation after crashes or rapid writes, large-replication corruption, and concurrency/benchmark questions. db0 issues reviewed +included connector/dialect exposure, caller-owned connections, deprecated sqlite3, and Drizzle result-shape mismatches. These are +why the OPFS project keeps ownership, collision semantics, partial-result failure, and backend capability differences explicit. -Current private repository reviewed through the connected GitHub source: - -- `okikio/mediad/AGENTS.md` -- root workspace/package/TypeScript configuration -- `docs/` organization -- `packages/media/*` organization -- `packages/media/storage` source - -Rules applied here include: - -- one-word capability-oriented folders where practical; -- precise verbs; -- `Schema` suffix for Zod schema constants; -- `Type` suffix for project-owned data types; -- same core TypeScript source across runtimes; -- explicit runtime subpaths; -- caller-owned injected resources by default; -- TSDoc that explains examples, impact, ownership, limits, failure behavior, and necessary background. - -Kaiju Platform and Crawl conventions ------------------------------------- - -The connected Library sources reviewed include: - -- Kaiju Platform Programming Model -- library-first architecture guidebook -- Kaiju naming and folder structure guide -- Kaiju code formatting guide -- Kaiju readable Markdown / technical writing handbook -- Kaiju Platform package/service architecture handoff -- Kaiju Crawl architecture and capability alignment handoff - -The project guidance used here includes: - -- library-first composition; -- explicit resource ownership and disposal; -- import-safe capability packages; -- exact runtime-resource names instead of vague terms; -- focused public subpaths; -- lazy iterators and bounded active memory; -- `.agents/` for temporary Node validation when Deno/JSR are not available; -- authored documentation must explain how exports compose into real developer workflows. - -Uploaded skill pack -------------------- +Deno KV issue review also covered historical reports about large prefix-list cost and selector/transaction limits: + +- +- +- + +The Deno KV physical key layout therefore indexes a logical entry by `(namespace, "entry", parentPath, name)`. Listing one +directory uses `(namespace, "entry", parentPath)` as the provider prefix, so descendants of a child directory are not part of +that prefix result. Physical body parts use the complete canonical path as one tuple component rather than expanding each path +segment into the provider prefix. This keeps exact lookup and direct-child enumeration aligned with the filesystem contract. + +Recent Drizzle issue review included SQLite/libSQL transaction-lifetime failures and migration/data-loss cases: + +- +- +- +- +- + +These are not all adapter-runtime bugs, but they reinforce a deliberate contract here: the generic Drizzle bridge does not +claim universal cross-process atomic replacement or own application migrations. The caller keeps dialect/driver/table lifecycle +and can provide a stronger database-specific transaction strategy when that concrete driver proves the required semantics. -`skills(20260806-212711).zip` was reviewed before this refactor. Relevant software delivery references included: +Project architecture and writing sources +---------------------------------------- -- documentation requirements; -- comment/TSDoc requirements; -- TypeScript requirements; -- library architecture and packaging; -- Deno software packaging; -- storage/database design and Drizzle guidance. +The implementation was reviewed against the attached/current project guides covering: + +- library-first capability composition; +- resource ownership and cancellation; +- short concrete naming and focused folders; +- compact code formatting; +- smooth narrative TSDoc/comments with invariants, examples, and lifecycle explanation; +- `node:test` plus `@std/expect` as the repository test API; +- runtime-neutral TypeScript and explicit runtime subpaths; +- verification against real runtimes and extracted release artifacts. + +Older OPFS experiments were treated as intent/history only. The current repository, current project rules, and current upstream +contracts are the implementation authority for this pass. + + +Client protocol handoffs +------------------------ -The skill pack is a process/reference input. It is not copied into the package artifact. +The detailed implementation contracts live in [s3.md](./s3.md) and [azure.md](./azure.md). The Docker-backed interoperability +matrix is documented in [providers.md](./providers.md). These files separate protocol requirements from emulator evidence and +record the unsupported surface explicitly. diff --git a/docs/validation.md b/docs/validation.md new file mode 100644 index 0000000..51b5f64 --- /dev/null +++ b/docs/validation.md @@ -0,0 +1,306 @@ +Validation strategy +=================== + +The test architecture separates portable filesystem semantics from the runtimes and providers that supply concrete storage. +This is deliberate. A fast memory test should not be the evidence for browser OPFS interoperability, and a browser test should +not be the only evidence for a deterministic path or copy invariant. + +The canonical layers are: + +```text +node:test + @std/expect + portable schemas, paths, facade behavior, record/object translations, + ecosystem bridges, S3/Azure protocol behavior with deterministic fakes + +real server runtimes + Deno host filesystem + Deno KV + Node host filesystem + node:sqlite + Bun host filesystem + +Playwright Test + Chromium / Firefox / WebKit + Window / Worker / ServiceWorker / iframe / persistence / browser storage + +Mitata + raw backend baseline -> adapter primitive -> filesystem facade +``` + +A test states the contract it protects. Avoid tests that only mirror the current implementation line by line. + +Portable tests protect filesystem semantics +------------------------------------------- + +The portable suite uses `node:test` with `describe` and `it`, plus `@std/expect` for expectations. Deno runs these same source +files directly. Node runs the same source. Bun runs the same `node:test` API through `bun test`. + +The suite covers: + +- schema acceptance/rejection and Standard Schema exposure; +- canonical path normalization and root-escape rejection; +- file and directory handle semantics; +- replace, append, update, truncate, and byte-range behavior; +- staged writable close versus abort; +- stream cancellation after the operation becomes terminal; +- bounded stream materialization for simple record adapters; +- Deno KV partitioned large-file stat/list/range/stream behavior, bounded append/update patching, and manifest-last visibility; +- optimization-disabled differential paths and matching preflight plans; +- filesystem route/peak-buffer metrics; +- copy/move overwrite and source/destination overlap protection; +- file mutation versus structural mutation coordination; +- queued cancellation recovery; +- sync-file lock lifetime and partial-write looping; +- adapter disposal ownership; +- record-store semantics; +- generic object-store directories, ranges, streaming replacement, optimistic read-modify-write, and native copy; +- foreign object layouts where an exact file key and a child prefix coexist; +- unstorage forward and reverse integration; +- the generic reverse key-value driver and collision-safe keys; +- RxDB, db0 dialect, Drizzle, and direct SQLite translation; +- S3 Signature Version 4, XML list/error parsing, multipart commit preconditions, HTTP-200 embedded failures, multipart + server-side copy, retry/backoff, timeout, manual redirects, one-shot body admission, and non-idempotent multipart lifecycle retry guards; +- Azure list/error parsing, large server-side range copy, bearer/SAS source authorization, provider request identities, + retry/backoff, and one-shot/explicit no-retry behavior. + +Focused commands: + +```sh +deno task test:portable +deno task test:node +deno task test:bun +``` + +The deterministic stress run shuffles and repeats the portable suite so hidden test order does not become a dependency: + +```sh +deno task test:stress +``` + +Runtime suites prove runtime adapters against the real API +---------------------------------------------------------- + +`tests/deno.test.ts` exercises the real Deno host filesystem adapter. `tests/node.test.ts` exercises real Node host filesystem +operations and runs the SQL bridge against Node's built-in SQLite engine. `tests/bun.test.ts` exercises the Bun adapter against +Bun's actual runtime. + +Deno KV has a separate real integration test because current Deno requires the unstable KV flag: + +```sh +deno task test:deno-kv +``` + +The adapter module is still type-checked with the server set. The unstable flag belongs to the real Deno KV execution, not to +unrelated package imports. + +The normal server-runtime matrix is: + +```sh +deno task test:deno +deno task test:deno-kv +deno task test:node +deno task test:bun +``` + +The pinned mise task runs these after installing the frozen dependency graph: + +```sh +mise run test +``` + +GitHub Actions does not recreate this runtime setup with separate Node, Deno, and Bun setup actions. `jdx/mise-action` installs +the pinned mise release and only the tools required by the current job. The job then calls the same focused mise task a +maintainer can run locally, such as `mise run test-deno`, `mise run test-node`, or `mise run test-bun`. This keeps tool versions +and test commands in the repository instead of duplicating them in workflow YAML. + +Playwright owns browser installation and browser lifecycle +--------------------------------------------------------- + +The browser suite lives under `tests/browser/`. There is no custom browser-launch loop or custom test-result protocol. +Playwright owns browser installation, contexts, server lifecycle, traces, retries, and test attribution. + +Install the compatible browser builds, then run the matrix: + +```sh +deno task test:browser:install +deno task test:browser +``` + +or: + +```sh +mise run test-browser +``` + +The same semantic tests run in Chromium, Firefox, and WebKit. Tests probe runtime capability and then assert the actual result. +They do not encode statements such as "Firefox has no sync OPFS" or "WebKit always rejects this iframe" into the test logic. +Those are exactly the assumptions an interoperability suite is supposed to detect when browser behavior changes. + +The browser cases include: + +```text +Window + OPFS probe + async write/read + abort before commit + +DedicatedWorker + async OPFS + synchronous handle probe and open attempt + +SharedWorker + async OPFS through the actual SharedWorker realm + +ServiceWorker + black-box registration + postMessage in all browsers + deeper serviceWorkers() instrumentation in Chromium only + +iframes + same-origin + cross-origin + opaque sandbox + +storage lifecycle + fresh BrowserContext isolation + persistent profile close/reopen + +browser record adapters + localStorage + IndexedDB + Cache Storage +``` + +The iframe and ServiceWorker tests report unsupported runtime APIs as capability skips. A supported realm whose OPFS root is +rejected is not silently skipped; the test asserts that a normalized root failure is present. + +Benchmarks measure overhead against the direct backend +------------------------------------------------------ + +A benchmark without a raw baseline cannot tell whether the adapter is fast or merely whether one code path is faster than +another OPFS code path. The benchmark layout therefore keeps three layers visible: + +```text +raw backend + | + v +adapter primitive + | + v +FileSystemType + | + +-- coordination: none + `-- coordination: local +``` + +`bench/memory.bench.ts` measures raw `Map`, direct `RecordStoreType`, direct memory adapter, and facade overhead. This exposes the +cost of record serialization separately from the higher-level filesystem contract. + +`bench/node.bench.ts`, `bench/deno.bench.ts`, and `bench/bun.bench.ts` compare raw host filesystem reads/writes and copy with +the direct adapter and facade. Bun measures both Node-compatible `copyFile` and `Bun.write(destination, Bun.file(source))` so a +future adapter change has a runtime baseline instead of an assumption. `bench/deno-kv.bench.ts` does the same for real local Deno +KV, and `bench/sqlite.bench.ts` compares a raw SQLite BLOB row with the direct record adapter and facade. Metrics are disabled for +facade baseline measurements. + +```sh +deno task bench:memory +deno task bench:deno +deno task bench:deno-kv +deno task bench:node +deno task bench:sqlite +deno task bench:bun +``` + +`mise run bench` runs the server/memory set with the pinned runtimes. + +The browser benchmark keeps the raw browser API, direct adapter, and facade visible in each real browser. Native OPFS uses 25 +replace/read iterations with a 64 KiB payload. localStorage, IndexedDB, and Cache Storage use 20 iterations with a 16 KiB +payload. Each sample records the raw, adapter, and facade durations plus the adapter/raw, facade/raw, and facade/adapter ratios +as a Playwright attachment. + +```sh +deno task bench:browser +# or +mise run bench-browser +``` + +Microbenchmarks are evidence about overhead in the measured operation. They are not universal provider throughput numbers. +Object-store latency, geographical distance, TLS, provider multipart behavior, and connection reuse can dominate the small +client/facade cost. + +`mise run bench-providers` uses the pinned SeaweedFS/Azurite fixture to compare official SDK/native runtime baselines with the +direct protocol clients, object adapters, and filesystem facade. S3 includes AWS SDK v3 and a Bun-native `S3Client` run; Azure +uses `@azure/storage-blob`. The small write baseline includes the same follow-up stat/properties request as the direct project +client, and multipart/block cases are separate. `metrics: "none"` versus `metrics: "basic"` makes instrumentation overhead +visible instead of hiding it. Real-cloud performance still requires an opt-in controlled provider benchmark. + +Type, lint, format, and documentation gates stay separate +--------------------------------------------------------- + +`deno task check` type-checks the code in environment-focused groups so unrelated ambient globals do not accidentally make an +invalid target look valid: + +```text +check:core + root/core + provider-neutral adapters/clients + reverse drivers + +check:browser + Window/browser storage adapters + Playwright specs/config + +check:workers + DedicatedWorker / SharedWorker / ServiceWorker fixtures with WebWorker libs + +check:server + Deno / Deno KV / Node / Bun / SQLite adapters and server benchmarks + +check:tests + portable + runtime test source + +check:deno-kv + Deno KV test source with the unstable KV flag +``` + +The normal quality gates are: + +```sh +deno ci +deno task check +deno task lint +deno task doc +deno task fmt:check +``` + +`deno ci` is important because the committed manifests and lockfile must describe one dependency graph. A changed dependency is +not ready for release until the real lockfile has been regenerated and the frozen install succeeds. + +Release validation checks the artifact, not only the source tree +--------------------------------------------------------------- + +Before publication, the repository runs the complete source-level checks plus registry dry-runs. npm packaging uses Deno's +package output and then adjusts Drizzle from a normal generated dependency to the optional peer relationship authored by this +project. + +The artifact gate should verify: + +1. the public export map contains every intended subpath and no internal-only file; +2. the generated npm package has JavaScript/declarations that import in Node, Deno, and Bun; +3. browser-safe imports bundle without pulling server-only adapters into the root graph; +4. optional Drizzle remains optional until its subpath is imported; +5. package files exclude tests, benchmarks, coverage, temporary output, and repository-only tooling; +6. the extracted artifact passes the same checks that are meaningful after packaging. + +The release command is: + +```sh +deno task release:check +``` + +Browser tests and browser benchmarks remain explicit matrix jobs because downloading three browser engines is a large operation +and should not be hidden inside every local unit-test invocation. + +Docker-backed provider tests +---------------------------- + +`mise run test-providers` starts the pinned SeaweedFS S3 endpoint and Azurite Blob emulator from `tests/provider/compose.yml`, +runs `tests/provider.test.ts`, and always removes the containers and volumes. The suite proves real HTTP/signing/interoperability +for the direct clients. It does not replace deterministic request-shape tests or real-cloud conformance. See +[providers.md](./providers.md) for the exact coverage and limitations. `mise run bench-providers` reuses the same containers for +the official-client/direct-client/adapter/facade benchmark matrix.