diff --git a/AGENTS.md b/AGENTS.md index 3adf43d..5900947 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -1,19 +1,55 @@ # Repository implementation rules +Read `README.md` and `docs/design.md` before changing architecture or public APIs. + - Treat `@okikio/opfs` as a library programming model, not as an application runtime. - Keep the root entrypoint import-safe in Window, Worker, Deno, Bun, and Node contexts. -- Put concrete storage integrations under `src/adapter/` and expose them through explicit public subpaths. -- Put reverse ecosystem interfaces under `src/driver/`. -- Prefer one-word file and folder names. Use more words only when the precise concept requires them. +- Use the storage path literally: `client -> driver -> adapter -> FileSystemType -> bridge`. +- A client owns a wire protocol when one exists. It does not own OPFS semantics. +- A driver owns backend-native storage mechanics, requirements, limits, optimization policy, physical metrics, and + resources that it explicitly acquires. +- An adapter translates one driver into the small canonical OPFS primitive contract. It does not reimplement provider + mechanics. +- A bridge starts from `FileSystemType` and implements a real ecosystem contract. Do not call direction metadata or a + nominal wrapper a bridge. +- Put direction/support metadata under `src/integration/`. Definitions remain import-safe and require no global + registry. +- Do not preserve obsolete layer names or compatibility entrypoints. Update every current consumer, test, export, + document, and benchmark when a public contract changes. +- Prefer one-word file and folder names. Use two or three words only when the precise concept requires them. - All Zod schema constants end in `Schema`. - Project-owned data types normally end in `Type`. -- Prefer direct schema and type exports. Use namespace imports only when they improve short operation call sites. -- Prefer `get`, `create`, `open`, `save`, `inspect`, `plan`, `convert`, `read`, `write`, `close`, `remove`, `copy`, and `move` over vague verbs. -- Avoid `generate`, `execute`, `handle`, `process`, `manager`, `helper`, `common`, `shared`, and `misc` unless an external protocol requires the word. -- Keep Deno, Bun, Node, browsers, and Workers on the same core TypeScript source. Runtime-specific adapters can use runtime-specific APIs behind explicit subpaths. -- Prefer Web APIs and existing standard-library capabilities before custom infrastructure. -- Adapters never configure logging, read environment variables, or acquire unrelated global resources at import time. -- The caller owns injected database, collection, storage, and filesystem resources unless an adapter option explicitly transfers ownership. -- TSDoc and comments use plain technical English. They teach options, examples, impact, reasoning, ownership, limits, failure behavior, and necessary background. -- Document public schemas, types, properties, functions, classes, and adapter contracts. Document internal symbols when their invariant, lifecycle, or failure behavior is not obvious. +- Prefer direct schema and type exports. Use namespace imports only when they improve short coherent operation call + sites. +- Prefer `get`, `create`, `open`, `save`, `inspect`, `plan`, `read`, `write`, `close`, `remove`, `copy`, and `move` over + vague verbs. +- Avoid `generate`, `execute`, `handle`, `process`, `manager`, `helper`, `common`, `shared`, and `misc` unless an + external protocol requires the word. +- Keep Deno, Bun, Node, browsers, and Workers on the same core TypeScript source. Runtime-specific drivers stay behind + explicit subpaths. +- Prefer Web APIs and `@std/*` before custom infrastructure when they provide the required semantics. +- Importing a module must not connect to a provider, read credentials, configure logging, start workers, or perform + unrelated global work. +- The caller owns injected databases, collections, clients, stores, mounts, and filesystems unless an explicit option + transfers ownership. +- Every behavior-changing optimization must be independently disableable and visible through inspection. +- Keep provider hard limits, implementation safety limits, user policy, and dynamic probe results distinct. Do not + collapse them into one unexplained number. +- `plan()` is deterministic preflight. It must not perform provider I/O. A probe is a separate explicit operation. +- Partitioning belongs to the driver that owns the physical layout. The driver must describe visibility, part limits, + cleanup, and whether the layout changes observable behavior. +- Use `node:test` with `describe` and `it`; use `@std/expect` for expectations. Playwright owns real browser environment + tests. +- Use mise as the repository command authority. GitHub Actions owns triggers, permissions, matrices, outputs, and + secrets, then calls `mise run ...`. +- Testcontainers owns disposable provider fixtures. Do not restore fixed host ports, hand-written readiness polling, or + Compose lifecycle scripts for S3/Azure tests. +- Benchmarks compare the native/provider baseline, project client when present, driver, adapter, facade with metrics + disabled, and measured facade. Compare filesystem clients such as Mountpoint and BlobFuse only on operations their + documented semantics support. +- TSDoc and comments use plain technical English. Teach options, examples, impact, reasoning, ownership, limits, + cancellation, performance, and necessary background. +- Document important non-exported symbols when they own an invariant, lifecycle rule, serialization layout, or failure + rule. - Comments explain why a rule exists or what must remain true. Do not restate obvious syntax. +- Keep agent-only validation under `.agents/`. It must never become a production import or published artifact. diff --git a/README.md b/README.md index 55cac1e..075d33a 100644 --- a/README.md +++ b/README.md @@ -1,323 +1,384 @@ -@okikio/opfs -============ +# @okikio/opfs -`@okikio/opfs` is an OPFS-shaped filesystem programming model that can sit on top of browser OPFS, host filesystems, -object stores, key-value stores, browser storage, document databases, and SQL databases. +`@okikio/opfs` is an OPFS-shaped storage programming model for browser OPFS, host filesystems, object stores, key-value +stores, browser storage, document databases, and SQL databases. -The public filesystem owns the semantics that application code should not have to rebuild: canonical virtual paths, -OPFS-shaped handles, recursive copy and move, cancellation, staged writable files, bounded stream fallbacks, coordination, -normalized failures, and resource ownership. An adapter translates those operations into one concrete backend. +The package does not pretend those systems are identical. It separates protocol behavior, backend storage mechanics, +OPFS translation, portable filesystem semantics, and reverse ecosystem projections so applications can inspect the real +route and its limits before work starts. ```text -application - | - +-- path API -------------------+ - | readFile / writeFile | - | copy / move / walk | - | v - +-- OPFS-shaped handles --> FileSystemType - | - canonical adapter operations - | - +---------------------+---------------------+ - | | | - v v v - native files record stores object stores - OPFS / Node / Deno KV / DB rows S3 / Azure Blob - / Bun | - +-- localStorage - +-- IndexedDB - +-- Cache Storage - +-- Deno KV - +-- unstorage - +-- RxDB - +-- db0 / SQLite - `-- Drizzle +ecosystem / native API + | + v + client protocol client when a wire protocol exists + | + v + driver backend-native storage contract + | requirements / limits / optimizations / physical metrics + v + adapter driver -> canonical OPFS primitives + | + v + FileSystemType paths / handles / locks / fallbacks / logical metrics + | + v + bridge FileSystemType -> real ecosystem contract ``` -The reverse direction is useful too. `@okikio/opfs/driver/kv` exposes any `FileSystemType` as a small hierarchical key-value -store, and `@okikio/opfs/driver/unstorage` adapts that view to unstorage. This means an application can give unstorage an OPFS, -Node, Deno, Bun, S3, Azure Blob, IndexedDB, Deno KV, SQLite, or another custom OPFS backend without a second provider matrix. +A client is optional. Node, Deno, Bun, browser OPFS, IndexedDB, and SQLite can start at the driver layer. S3 and Azure +Blob have explicit protocol clients because their wire contracts are independently useful. -Install and start with the backend you actually own ----------------------------------------------------- +The filesystem facade owns the behavior application code should not rebuild for every backend: canonical virtual paths, +OPFS-shaped handles, recursive work, cancellation, staged writable files, bounded stream fallbacks, coordination, +normalized failures, resource ownership, and deterministic preflight planning. -Deno and JSR can import the package directly: +## Start with the storage you own -```ts -import { openFileSystem } from "jsr:@okikio/opfs"; - -const fileSystem = await openFileSystem(); -try { - await fileSystem.writeFile("/state/app.json", "{}", { parents: true }); -} finally { - await fileSystem.close(); -} -``` +Browser OPFS has a root convenience function: -npm-compatible runtimes use the same TypeScript API: +```ts +import { openFileSystem } from "@okikio/opfs"; -```sh -npm install @okikio/opfs +await using fileSystem = await openFileSystem(); +await fileSystem.writeFile("/state/app.json", "{}", { parents: true }); ``` -Server code selects a concrete adapter instead of importing a different filesystem API: +Server code can compose each layer explicitly: ```ts import { createFileSystem } from "@okikio/opfs"; import { createNodeAdapter } from "@okikio/opfs/adapter/node"; -const fileSystem = createFileSystem( +await using fileSystem = createFileSystem( createNodeAdapter({ root: "./data" }), { coordination: "local" }, ); ``` -The root entrypoint is intentionally import-safe in browsers, workers, Deno, Bun, and Node. Runtime-specific dependencies stay -on explicit subpaths. Importing `@okikio/opfs` does not import `node:fs`, inspect environment variables, connect to databases, -or configure application logging. - -The first-party backend set is deliberately broad, but the layers stay small: - -| Subpath | Backend or role | Important behavior | -| --- | --- | --- | -| `adapter/opfs` | native browser OPFS | native handles, streams, sync access when exposed | -| `adapter/node` | `node:fs` | streams, ranges, copy/rename, sync random access | -| `adapter/deno` | Deno filesystem | streams, ranges, copy/rename, sync random access | -| `adapter/bun` | Bun + Node-compatible fs | Bun read/write fast paths plus host filesystem operations | -| `adapter/memory` | in-memory records | deterministic tests and temporary state | -| `adapter/record` | `RecordStoreType` | common translation for value/document/SQL stores | -| `adapter/object` | `ObjectStoreType` | common translation for object stores without hiding object semantics | -| `adapter/s3` | `S3ClientType` | direct S3/S3-compatible storage, no AWS SDK | -| `adapter/azure` | `AzureClientType` | direct Azure Blob REST storage, no Azure SDK | -| `adapter/localstorage` | Web Storage | synchronous string store translated through records | -| `adapter/indexeddb` | IndexedDB | indexed record persistence with caller-controlled database ownership | -| `adapter/cache` | Cache Storage | record persistence in an injected Cache | -| `adapter/deno-kv` | Deno KV | record persistence over a caller-owned KV database | -| `adapter/sqlite` | connected SQLite | focused SQLite view over the same SQL record contract as db0 | -| `adapter/unstorage` | unstorage `Storage` | forward bridge above the selected unstorage driver | -| `adapter/rxdb` | RxDB collection | forward bridge above the selected RxStorage | -| `adapter/db0` | db0 `Database` | SQL bridge across db0 dialects/connectors | -| `adapter/drizzle` | Drizzle database + table | caller-owned schema and common CRUD bridge | -| `driver/kv` | reverse key-value view | collision-safe key hierarchy over any filesystem | -| `driver/unstorage` | reverse unstorage driver | lets unstorage consume any `FileSystemType` | - -Drizzle is an optional peer dependency because the integration is only loaded through its explicit subpath. - -Object storage is not flattened into a fake local disk -------------------------------------------------------- - -S3 and Azure Blob can both back the filesystem facade, but they remain object stores underneath. That distinction affects -performance and correctness. - -A complete replacement can stream to multipart/block upload. Append and update cannot normally mutate object bytes in place, -so the object adapter performs a read-modify-write operation. When the provider exposes conditional writes, the previous ETag -is used as an optimistic precondition so a concurrent writer fails rather than being silently overwritten. - -Native copy is also a separate capability. The filesystem asks the adapter to copy before it opens a source stream, so -provider-side copy stays inside S3/Azure instead of becoming an accidental download and re-upload. +`createNodeAdapter()` is a convenience composition. The explicit form is useful when an application wants to inspect or +use the backend driver before it creates a filesystem: + +```ts +import { createFileSystem } from "@okikio/opfs"; +import { createFileAdapter } from "@okikio/opfs/adapter/file"; +import { createNodeDriver } from "@okikio/opfs/driver/node"; + +const driver = createNodeDriver({ root: "./data" }); +console.log(driver.inspect()); +console.log(driver.plan({ operation: "write", path: "/large.bin", size: 1_000_000 })); + +const adapter = createFileAdapter(driver); +await using fileSystem = createFileSystem(adapter); +``` + +The root entrypoint remains import-safe in Window, Worker, Deno, Bun, and Node contexts. Runtime-specific code stays on +explicit subpaths. Importing the root package does not connect to a provider, read credentials, start workers, or +configure application logging. + +## Layer inventory + +The first-party integration set is broad, but each layer has one job. + +| Family | Client | Driver | Adapter | Reverse bridge | +| --------------- | --------------------- | --------------------- | ---------------------- | --------------------------------------- | +| browser OPFS | n/a | `driver/opfs` | `adapter/opfs` | filesystem itself | +| Node filesystem | n/a | `driver/node` | `adapter/node` | ecosystem-specific | +| Deno filesystem | n/a | `driver/deno` | `adapter/deno` | ecosystem-specific | +| Bun filesystem | n/a | `driver/bun` | `adapter/bun` | ecosystem-specific | +| memory | n/a | `driver/memory` | `adapter/memory` | generic KV bridge possible | +| Deno KV | n/a | `driver/deno-kv` | `adapter/deno-kv` | generic KV bridge possible | +| localStorage | n/a | `driver/localstorage` | `adapter/localstorage` | generic KV bridge possible | +| IndexedDB | n/a | `driver/indexeddb` | `adapter/indexeddb` | generic KV bridge possible | +| Cache Storage | n/a | `driver/cache` | `adapter/cache` | cache-specific bridge not yet provided | +| SQLite rows | engine-owned | `driver/sqlite` | `adapter/sqlite` | see database direction below | +| db0 | connector-owned | `driver/db0` | `adapter/db0` | no fake SQL projection | +| Drizzle | dialect/driver-owned | `driver/drizzle` | `adapter/drizzle` | no fake SQL projection | +| RxDB | RxStorage-owned | `driver/rxdb` | `adapter/rxdb` | full RxStorage bridge not yet provided | +| unstorage | upstream driver-owned | `driver/unstorage` | `adapter/unstorage` | `bridge/unstorage` | +| S3 | `s3` | `driver/s3` | `adapter/s3` | object-specific bridge not yet provided | +| Azure Blob | `azure` | `driver/azure` | `adapter/azure` | object-specific bridge not yet provided | + +`adapter/file`, `adapter/record`, and `adapter/object` are reusable translators for third-party drivers. `driver/file`, +`driver/record`, and `driver/object` are the corresponding backend-native contracts. + +## Drivers are independently useful + +A driver is not an adapter with a new name. It owns backend mechanics that remain meaningful without `FileSystemType`. + +A configured driver reports: ```text -filesystem.copy() - | - +-- nativeCopy ----> provider/server-side copy - | - `-- fallback ------> source stream -> bounded transfer -> destination +name and storage family +stable backend operations/capabilities it provides +backend resource ownership: none / borrowed / owned +requirements and current availability +provider hard limits +implementation safety limits +user-selected policy limits +dynamic limits that still require probing +behavior-changing and transparent optimizations +structured preflight problems and actions +physical metrics when available +owned-resource disposal when applicable ``` -The direct S3 client implements Signature Version 4, range reads, ListObjectsV2, multipart upload, conditional completion, -CopyObject, and multipart UploadPartCopy for objects above CopyObject's 5 GB source limit. It also checks S3's unusual -success-with-error-body responses for copy and multipart completion. +Limits always include provenance. For example, Deno KV can report its serialized provider ceiling separately from the +smaller raw payload budget this library chooses to leave room for serialization overhead. A caller can therefore +distinguish a provider rule from an implementation safety choice and from its own `maxParts` policy. + +`plan()` is deterministic and performs no provider I/O: ```ts -import { createFileSystem } from "@okikio/opfs"; -import { createS3Adapter } from "@okikio/opfs/adapter/s3"; -import { createS3Client } from "@okikio/opfs/s3"; - -const client = createS3Client({ - endpoint: "https://s3.us-east-1.amazonaws.com", - bucket: "my-bucket", - region: "us-east-1", - credentials: { accessKeyId, secretAccessKey }, +const plan = driver.plan({ + operation: "write", + path: "/archive/data.bin", + size: 80 * 1024, + source: "bytes", + mode: "replace", }); -const fileSystem = createFileSystem(createS3Adapter(client)); +for (const problem of plan.problems) console.log(problem.code, problem.limit); +for (const action of plan.actions) console.log(action.kind); ``` -S3 compatibility is a protocol family, not one identical product. The client therefore accepts endpoint, region, addressing, -headers, and capability overrides. For example, an S3-compatible service that does not support multipart preconditions should -set `conditionalWrite: false` instead of pretending the safety property exists. `client.request()` remains available for S3 -features that do not belong in the portable filesystem contract. - -Azure uses its own REST model instead of being forced through an S3 abstraction. It supports SAS, Microsoft Entra bearer -tokens, Shared Key, and caller-defined authorization headers. Large server-side copies use Put Block From URL after Azure's -smaller synchronous Copy Blob From URL path is no longer sufficient. +Dynamic facts such as available quota remain unknown until a separate explicit probe supplies them. The planner never +invents a quota or silently performs network/storage I/O. -The protocol clients are documented separately because their wire contracts are larger than the filesystem adapter surface: +## Capabilities are layered instead of flattened -- [S3 client protocol](./docs/s3.md) covers SigV4, request canonicalization, multipart upload/copy, conditions, limits, errors, - compatibility controls, and known non-goals. -- [Azure Blob client protocol](./docs/azure.md) covers REST versions, SAS/bearer/Shared Key authorization, block upload/copy, - conditions, limits, errors, and Azurite behavior. -- [Provider integration tests](./docs/providers.md) explains the Testcontainers-backed SeaweedFS and Azurite matrix and what those - local providers can and cannot prove. +`driver.inspect()` describes the backend. `FileSystemType.inspect()` adds the adapter translation and effective facade +routes: -The facade makes capability, limits, routing, and cost inspectable ----------------------------------------------------------------- +```ts +const inspection = fileSystem.inspect(); + +inspection.driver; // provides, ownership, requirements, limits, optimizations +inspection.adapter; // native OPFS translation capabilities +inspection.support; // effective native/emulated/partitioned/unsupported routes +inspection.optimizations; // facade route switches +inspection.metrics; // logical filesystem counters +inspection.driverMetrics; // physical backend counters when available +``` -`AdapterCapabilitiesType` describes immediate adapter behavior. `FileSystemType.inspect()` describes the configured stack after -facade fallbacks and optimization policy are applied. This distinction lets callers ask whether a route is `native`, `emulated`, -`partitioned`, or `unsupported` without guessing from the adapter name. +`FileSystemType.plan()` combines the driver preflight with adapter and facade policy. Problems remain structured and +retain the layer that identified them. Human-readable messages are presentation data, not the machine contract. ```ts -const fileSystem = createFileSystem(adapter, { - maxBufferedWriteBytes: 32 * 1024 * 1024, - metrics: "basic", - optimizations: { - nativeCopy: false, - }, -}); - -console.log(fileSystem.inspect()); -console.log(fileSystem.plan({ +const plan = fileSystem.plan({ operation: "write", + path: "/video.bin", source: "stream", - mode: "replace", size: 512 * 1024 * 1024, -})); +}); + +if (!plan.supported) { + // Example actions: partition, change-policy, reduce-input, select-driver. + console.log(plan.problems, plan.actions); +} ``` -`inspect()` includes native capabilities, effective support, hard limits known by the adapter, partition layout, resolved -optimization controls, the facade buffer ceiling, and a detached metrics snapshot. `plan()` is deterministic and does no I/O. -When size is known it can reject a request before work begins, show expected facade materialization, or explain the physical -part count selected by a partitioned adapter. Unknown provider limits remain unknown rather than being invented. +Every optimization that can change request count, failure timing, storage layout, consistency, atomicity, or another +observable property is independently disableable and visible through inspection. -Write planning separates the resulting logical file from the bytes supplied by the current call. `size` checks logical -file/partition limits. `inputBytes` checks whether a non-native input stream fits under `maxBufferedWriteBytes`. Replace usually -needs only `size`; append/update should provide both values when they are known. +## Object storage keeps object semantics -Optimizations that select a materially different route are independently disableable: native stream read/write, direct range -read, native/server-side copy, and native move. The fallback is used only when it can preserve the portable filesystem contract. -For example, disabling provider-side copy can force bytes through this process and cannot reproduce provider-private control-plane -metadata such as every ACL, tag, lock policy, or checksum policy. Portable file bytes and `mediaType` are preserved. +S3 and Azure Blob do not become fake POSIX disks. Their clients keep protocol-specific operations while object drivers +expose portable object mechanics to the adapter. -`metrics: "none"` removes facade counter updates for baseline benchmarks. `basic` counts operations, bytes, failures, route -selection, and peak facade materialization. `timing` adds monotonic durations. The direct S3 and Azure clients expose separate HTTP -request/retry metrics so protocol overhead and facade overhead can be measured independently. +```text +S3 REST / Azure Blob REST + | + v + client + | + v + object driver + | + v + object adapter + | + v + FileSystemType +``` -Large values are a backend capability, not a promise that every value store is unlimited. `adapter/deno-kv` is the first record -backend with a physical partition layout. Small files stay inline; large files use raw binary parts and a manifest-last commit. -Metadata lookup, directory listing, byte ranges, and stream reads do not reconstruct the complete logical file. The partition -policy is `never | auto | always`, so applications that do not want a changed durable layout can disable it explicitly. -Materialized append/update writes also build a new generation part-by-part, so the existing logical file is not joined into one -large base64 record before a small patch can be applied. +A complete replacement can stream through multipart/block upload. Append and update normally require read-modify-write. +Server-side copy remains a separate capability because routing bytes through JavaScript is not equivalent to a provider +control plane copy. -`streamWriteModes` remains mode-specific. A simple record adapter can have no native stream lane, while Deno KV can advertise a -partitioned replacement stream and an object store can advertise native replacement streaming. Append and update can still be -emulated or unsupported independently. Deno KV's materialized append/update lane is direct, but streamed append/update remains -emulated because the incoming stream must first fit under the facade buffer ceiling. +The S3 client exposes two optimization switches that are useful to inspect and benchmark: -Bridges make both integration directions explicit -------------------------------------------------- +- `delayedMultipart`: buffers the first bounded part so a small unknown-length stream can use `PutObject`. Disable it + when the multipart request lifecycle itself is required. +- `signingKeyCache`: reuses the derived SigV4 signing key for unchanged credentials/date/region/service. It does not + change the signed request semantics. -Adapters remain `ecosystem -> OPFS`. Drivers remain `OPFS -> ecosystem`. A bridge groups both directions and records an explicit -reason when one direction cannot honestly exist. +Azure exposes `blockUpload` and `serverCopy` independently. Disabling block upload removes native streamed-write support +rather than pretending a stream can still be sent without staging. Disabling server copy makes the object adapter/facade +choose an honest fallback when one can preserve the requested semantics. -```ts -import { UnstorageBridge, RxDbBridge } from "@okikio/opfs/bridge"; +See [docs/s3.md](./docs/s3.md) and [docs/azure.md](./docs/azure.md) for the protocol contracts. + +## Deno KV is the reference partitioned record driver -console.log(UnstorageBridge.directions); -// { toOpfs: { supported: true }, fromOpfs: { supported: true } } +Deno KV demonstrates why one flat `maxFileBytes` number is insufficient. The driver separates: -console.log(RxDbBridge.directions.fromOpfs); -// unsupported: a filesystem is not an RxStorage query/conflict/change-stream engine +```text +provider: serialized key/value and atomic-operation ceilings +implementation: conservative inline/part payload budgets +user policy: partition mode, part size, max parts, I/O concurrency +derived: logical file capacity for the selected layout ``` -The included bridge descriptors cover unstorage, RxDB, db0, Drizzle, and the generic reverse key-value view. Reverse KV and -unstorage drivers also expose the backing filesystem's `inspect()`, `plan()`, and `getMetrics()` methods so capability, size, -partition, optimization, and instrumentation decisions remain visible after the direction changes. Third parties can -use `defineBridge()` without a global registry. An unsupported direction must include a reason, which prevents a bridge from -silently pretending that asynchronous filesystem behavior can provide an unrelated synchronous or query-oriented contract. +Large files use immutable physical generations and a manifest-last visibility point: -Use schemas directly --------------------- +```text +old manifest -> old parts + +write new part 0..N + | + v +write new manifest visibility point + | + v +remove old reachable generation +``` -Project-owned structural data is defined by Zod schemas and inferred TypeScript types. Schema constants end in `Schema`, and -project-owned serializable types normally end in `Type`. +A failed write does not publish a partial logical file. Unknown-length streamed replacement uses the partitioned lane +when partitioning is enabled. `partition: "never"` disables that behavior and makes oversized/streaming requests fail or +use a bounded facade fallback instead of changing durable layout silently. -```ts -import { PathSchema, type PathType } from "@okikio/opfs/schema"; +The preflight planner also evaluates the concrete virtual path. Deno KV limits serialized keys, so file size alone is +not enough to decide whether an operation is admissible. + +## Database direction matters -const path: PathType = PathSchema.parse("/cache/result.bin"); +There are two different database architectures and the package documents them separately. + +### Database-backed filesystem + +The existing SQLite/db0/Drizzle/RxDB drivers store logical filesystem records in an existing database abstraction: + +```text +Database / collection + | + v +record driver + | + v +record adapter + | + v +FileSystemType +``` + +For Drizzle, the caller supplies a connected database and a table. The generic driver intentionally reports replacement +as best-effort because its portable CRUD route is delete then insert. A dialect-specific driver can expose stronger +binary, transaction, upsert, or partition behavior without weakening the generic contract. + +### Database file stored on OPFS + +Using OPFS as storage for a SQLite database is the opposite direction: + +```text +application + | + Drizzle + | +SQLite engine + | +SQLite VFS + | +FileSystemType / native OPFS ``` -Zod 4 schemas implement Standard Schema, so consumers that accept Standard Schema can use these exported schemas directly. The -package does not maintain a parallel wrapper layer that could drift from the executable Zod contract. +`adapter/sqlite` does **not** implement this topology. It stores virtual filesystem records inside an already connected +SQLite database. A future SQLite VFS integration must implement the database engine's real VFS contract. It must not be +represented as an SQL bridge that only renames filesystem methods. + +See [docs/ecosystems.md](./docs/ecosystems.md) for concrete Drizzle/db0/RxDB behavior. + +## Bridges are real reverse contracts -Test the semantics where they actually run ------------------------------------------- +A bridge starts from `FileSystemType` and implements an ecosystem contract. Direction metadata lives under +`integration/` and is not itself a bridge. -Portable filesystem contracts use `node:test` and `@std/expect`. Deno runs the same portable test source, Node runs the same -source, and Bun runs the same `node:test` API through its compatibility layer. Runtime-specific suites then prove the real host -filesystem adapters. +The package currently provides: -Playwright Test owns the browser matrix. The same tests run in Chromium, Firefox, and WebKit and exercise Window OPFS, -DedicatedWorker, SharedWorker, ServiceWorker, same-origin and cross-origin iframes, opaque sandbox behavior, fresh-context -isolation, persistent-profile reopen, cancellation, and browser storage adapters. Tests probe capabilities instead of selecting -behavior from browser names. +- `bridge/kv`: a small hierarchical key-value contract over any `FileSystemType`. +- `bridge/unstorage`: an unstorage Driver-shaped implementation over any `FileSystemType`. -Mitata benchmarks compare three layers where possible: +The unstorage layout keeps `foo` and `foo:bar` distinct even though a normal filesystem path cannot be both a file and a +directory. ```text -raw backend API - | - v -adapter primitive - | - v -FileSystemType facade - | - +-- coordination: none - `-- coordination: local +unstorage + | +bridge/unstorage + | +FileSystemType + | +any configured adapter/driver stack ``` -Browser benchmarks compare raw native APIs, direct adapters, and the facade for OPFS, localStorage, IndexedDB, and Cache -Storage in Chromium, Firefox, and WebKit. Node, Deno, and Bun benchmarks compare their raw filesystem APIs with direct adapters and the facade. Bun additionally compares -Node-compatible `copyFile` with `Bun.write(destination, Bun.file(source))` rather than assuming one host copy path is faster. -Deno KV and SQLite have the same raw-to-adapter-to-facade measurements. +`integration` definitions state which directions are real. RxDB, db0, and Drizzle currently remain honest one-way +integrations because a filesystem alone is not an RxStorage query/conflict/change-stream engine or a SQL engine. +Third-party packages can use `defineIntegration()` and the driver/adapter primitives without registering global state. + +## Testing and benchmarks follow the layers + +Portable behavior uses `node:test` with `@std/expect`. The same source is checked/run in the supported server runtimes +where the runtime capability exists. Playwright Test owns actual Window, Worker, iframe, ServiceWorker, persistence, and +browser-storage coverage. Testcontainers owns disposable SeaweedFS and Azurite provider fixtures. -Provider benchmarks use the same local provider fixture but keep each layer separate: official AWS/Azure SDK baseline, direct -protocol client, direct object adapter, facade with metrics disabled, and facade with basic metrics. A Bun-native S3 run compares -Bun's Rust-backed `S3Client` against the same project layers. Multipart/block cases are separate from single-request writes so a -different request plan is never presented as abstraction overhead. +Provider benchmarks compare: -With the pinned mise toolchain: +```text +official/native baseline + | +project protocol client + | +project driver + | +project adapter + | +facade metrics:none + | +facade metrics:basic +``` + +`bench/filesystem-provider.bench.ts` adds the same staircase above already-mounted AWS Mountpoint and Azure BlobFuse +filesystems. Set `OPFS_MOUNTPOINT_S3_ROOT` and/or `OPFS_BLOBFUSE_ROOT` before `mise run bench-filesystem-clients`. The +benchmark uses only file operations that can be compared through the mounted filesystem contract. It does not count an +unsupported filesystem operation as a performance failure. + +Mise is the repository command authority: ```sh mise install mise run check mise run test -mise run bench mise run test-browser -mise run bench-browser +mise run test-providers +mise run bench mise run bench-providers +mise run bench-browser +mise run bench-filesystem-clients # requires external mounts ``` -GitHub Actions uses the same tool declarations and mise tasks. The workflow installs mise once per job, asks mise to install -only the runtimes that job needs, and then calls `mise run ...`. Runtime matrix jobs override one configured version with -`MISE__VERSION`; the Node matrix uses this to test Node 22, 24, and 26 without introducing a second tool-version -source. Third-party actions are pinned to immutable commit SHAs, and the mise binary version is pinned separately. - -The focused Deno tasks are documented in [docs/validation.md](./docs/validation.md). - -Read the rest by the question you have --------------------------------------- - -- [Public API](./docs/api.md) explains the developer-facing filesystem and handle contracts. -- [Adapters](./docs/adapters.md) explains every first-party backend and the contracts for custom storage. -- [Architecture](./docs/design.md) explains invariants, streaming, copy/move, locks, ownership, and failure behavior. -- [Ecosystems](./docs/ecosystems.md) explains unstorage, RxDB, db0, Drizzle, S3-compatible services, and reverse drivers. -- [Environments](./docs/environments.md) explains Window, workers, iframes, Deno, Bun, Node, and provider clients. -- [Validation](./docs/validation.md) defines the canonical test and benchmark matrix. -- [Sources](./docs/sources.md) records the standards and upstream contracts that the implementation follows. -- [Releasing](./docs/releasing.md) explains JSR/npm packaging and release checks. +GitHub Actions owns triggers, permissions, matrices, outputs, and secrets. It then invokes the same mise tasks. Release +and publish commands also live under `.mise/tasks/` rather than becoming a second command layer in workflow YAML. + +## Read next + +- [Public API](./docs/api.md) explains filesystem, inspection, planning, driver, adapter, and bridge entrypoints. +- [Architecture](./docs/design.md) explains layer ownership, invariants, partitioning, metrics, and lifecycle. +- [Adapters and drivers](./docs/adapters.md) explains the first-party translation/backend matrix and extension + contracts. +- [Ecosystems](./docs/ecosystems.md) explains unstorage, RxDB, db0, Drizzle, SQLite direction, and reverse bridge + constraints. +- [Environments](./docs/environments.md) explains browser realms, Deno, Bun, Node, and server coordination. +- [Providers](./docs/providers.md) explains Testcontainers and provider/filesystem baseline benchmarks. +- [Validation](./docs/validation.md) defines the release gates and this repository's test matrix. +- [Sources](./docs/sources.md) records the standards and upstream contracts used by the implementation. +- [Releasing](./docs/releasing.md) explains mise-owned release and registry publication. diff --git a/docs/adapters.md b/docs/adapters.md index 70701ce..edad4af 100644 --- a/docs/adapters.md +++ b/docs/adapters.md @@ -1,131 +1,135 @@ -Adapter guide -============= +# Drivers and adapters -An adapter translates the package's canonical virtual filesystem operations into one concrete backend. The filesystem facade -owns filesystem semantics. The adapter owns backend mechanics. +## Purpose -That distinction lets the required backend contract stay small: +A driver owns backend-native storage. An adapter translates that driver into the small canonical filesystem primitive +contract. Keeping those roles separate lets applications use a driver directly, inspect real provider limits, and +measure adapter/facade overhead independently. ```text -stat read one entry's metadata -readFile read one file or range -writeFile commit one materialized write -readDir lazily list direct children -createDir create exactly one directory -remove remove one file or empty directory +backend/native API + | + driver + | + adapter + | +FileSystemType ``` -The facade builds parent creation, recursive walking, recursive copy/remove, OPFS-shaped handles, write-command staging, -coordination, and normalized errors on top. A backend can add native operations when it can do better than the facade fallback. +Use the convenience adapters for normal application code. Use explicit drivers when you need backend planning, physical +metrics, provider-specific operations, or a custom translation. + +## First-party matrix + +| Storage | Driver | Adapter | Native family | +| ------------------------ | --------------------- | ---------------------- | ------------- | +| browser OPFS | `driver/opfs` | `adapter/opfs` | file | +| Node filesystem | `driver/node` | `adapter/node` | file | +| Deno filesystem | `driver/deno` | `adapter/deno` | file | +| Bun filesystem | `driver/bun` | `adapter/bun` | file | +| memory | `driver/memory` | `adapter/memory` | record | +| Deno KV | `driver/deno-kv` | `adapter/deno-kv` | record | +| localStorage | `driver/localstorage` | `adapter/localstorage` | record | +| IndexedDB | `driver/indexeddb` | `adapter/indexeddb` | record | +| Cache Storage | `driver/cache` | `adapter/cache` | record | +| SQLite rows | `driver/sqlite` | `adapter/sqlite` | record | +| unstorage Storage | `driver/unstorage` | `adapter/unstorage` | record | +| RxDB collection | `driver/rxdb` | `adapter/rxdb` | record | +| db0 Database | `driver/db0` | `adapter/db0` | record | +| Drizzle database + table | `driver/drizzle` | `adapter/drizzle` | record | +| S3 | `driver/s3` | `adapter/s3` | object | +| Azure Blob | `driver/azure` | `adapter/azure` | object | + +Reusable family translators: ```text -openReadStream native streaming read -writeStream native streaming for declared write modes -copy native/server-side file copy -move native rename/move -openWritableFile long-lived asynchronous positional writes -openSyncFile synchronous random access +driver/file -> adapter/file +driver/record -> adapter/record +driver/object -> adapter/object ``` -`AdapterCapabilitiesType` must describe these native paths truthfully. `streamWriteModes` is a list rather than one boolean -because replacement, append, and update can have different backend costs. `nativeCopy` is separate from `nativeMove` because -object stores often copy efficiently but cannot rename an object atomically. +## File drivers -Adapters can also expose `limits` and `partition`. Limits are hard facts known by the configured backend, such as maximum file, -value, key, part, batch, or concurrency sizes. Missing fields mean unknown, not unlimited. Partition describes a durable physical -layout used when one logical file spans multiple provider values. These fields feed `FileSystemType.inspect()` and `plan()` but -do not change the required adapter method set. +`FileDriverType` preserves real file-like operations. Required primitives are metadata, materialized read/write, +direct-child listing, one-directory creation, and single-entry removal. Optional direct operations include streams, +copy, move, positional files, and synchronous random access. -Route-changing optimizations live on the facade, not inside capability flags. `optimizations.streamRead`, `streamWrite`, -`rangeRead`, `nativeCopy`, and `nativeMove` can force the safe fallback for differential testing or application policy. An adapter -should therefore implement the best native route it can and let the caller decide whether to use it. - -Use `createFileSystem()` to put the public API over any adapter: +A third-party file driver can be created with `defineFileDriver()`: ```ts import { createFileSystem } from "@okikio/opfs"; +import { createFileAdapter } from "@okikio/opfs/adapter/file"; +import { defineFileDriver } from "@okikio/opfs/driver/file"; -const fileSystem = createFileSystem(adapter, { - coordination: "auto", - maxBufferedWriteBytes: 64 * 1024 * 1024, +const driver = defineFileDriver(backend, { + name: "my-files", + capabilities: { + read: true, + write: true, + streamRead: true, + streamWriteModes: ["replace"], + rangeRead: true, + nativeCopy: false, + nativeMove: false, + positionalWrite: false, + syncAccess: false, + }, }); + +const fileSystem = createFileSystem(createFileAdapter(driver)); ``` -The first-party adapters cover three different storage shapes -------------------------------------------------------------- - -Native filesystems expose files and directories directly. Record stores expose values keyed by logical identity. Object stores -expose whole-object replacement, ranges, prefixes, and provider-side copy. Keeping those shapes separate is what prevents one -"universal" adapter from hiding important performance and consistency behavior. - -| Public subpath | Backend | Translation layer | -| --- | --- | --- | -| `adapter/opfs` | browser OPFS | native filesystem | -| `adapter/node` | Node `fs` | native filesystem | -| `adapter/deno` | Deno filesystem | native filesystem | -| `adapter/bun` | Bun + Node-compatible fs | native filesystem | -| `adapter/memory` | in-memory map | records | -| `adapter/record` | custom value/document store | records | -| `adapter/localstorage` | Web Storage | records | -| `adapter/indexeddb` | IndexedDB | records | -| `adapter/cache` | Cache Storage | records | -| `adapter/deno-kv` | Deno KV | records | -| `adapter/sqlite` | connected SQLite | records through db0-compatible SQL | -| `adapter/unstorage` | unstorage `Storage` | records | -| `adapter/rxdb` | RxDB `RxCollection` | records | -| `adapter/db0` | db0 `Database` | records | -| `adapter/drizzle` | Drizzle database + table | records | -| `adapter/object` | custom object store | objects | -| `adapter/s3` | direct S3/S3-compatible client | objects | -| `adapter/azure` | direct Azure Blob client | objects | - -Native browser and host filesystems ------------------------------------ - -`openFileSystem()` is the convenience path for native browser OPFS: +The backend must implement every capability it advertises. The adapter does not fabricate a native method from a flag. -```ts -import { openFileSystem } from "@okikio/opfs"; +### Browser OPFS -const fileSystem = await openFileSystem(); -``` +`createOpfsDriver(root)` owns native browser handles. `createOpfsAdapter(driver)` is the explicit translation. The +convenience `openFileSystem()` acquires `navigator.storage.getDirectory()`, creates the OPFS driver and adapter, then +creates the facade. -The explicit form is useful when the caller already owns the native root: +The driver retains the native root for advanced browser interop. Sync access is advertised only when the actual file +handle exposes the required method in the current realm. -```ts -import { createFileSystem } from "@okikio/opfs"; -import { createOpfsAdapter } from "@okikio/opfs/adapter/opfs"; +### Node -const root = await navigator.storage.getDirectory(); -const fileSystem = createFileSystem(createOpfsAdapter(root)); -``` +`createNodeDriver({ root })` maps virtual `/` below one host directory. The host-path mapper rejects escape from that +root. -The OPFS adapter reports synchronous access only when an actual file handle exposes `createSyncAccessHandle()`. The package -does not infer the feature from a browser name or from "worker" alone. +Node exposes: -Node, Deno, and Bun map virtual `/` below one configured host directory: +- materialized and streaming reads; +- ranged reads; +- materialized and streaming writes; +- native file copy; +- native rename/move; +- asynchronous positional files; +- synchronous random access and flush. -```ts -import { createNodeAdapter } from "@okikio/opfs/adapter/node"; +Node built-ins resolve only when the explicit Node driver is created/imported. The root module does not import Node +runtime code. -const adapter = createNodeAdapter({ root: "./data" }); -``` +### Deno + +`createDenoDriver({ root })` uses Deno file APIs for persistence and `@std/path` only for the host-root mapper. It +supports the same major file routes as the Node driver where Deno provides the native primitive. -The host path mapper resolves the configured root once and rejects every virtual path whose resolved host path would leave that -root. Host adapters expose ranges, streams, native copy, native move, and synchronous random access when the underlying runtime -provides them. +The driver requires filesystem permissions chosen by the host application. It does not request broad permissions itself. -Bun uses Bun's file APIs where they provide a direct read/write path and uses Bun's Node-compatible filesystem surface for the -operations whose exact semantics already live there. Importing the Bun adapter does not require the `Bun` global until adapter -creation. +### Bun -Record stores start small and can add byte lanes ------------------------------------------------ +`createBunDriver({ root })` uses `Bun.file()` and `Bun.write()` where they improve the direct read/replace path, then +delegates operations that need stronger host-filesystem semantics to the Node-compatible file driver. -The required `RecordStoreType` stays intentionally small: +The Bun global is resolved lazily during driver creation. Importing the module in Node or Deno does not require Bun. + +## Record drivers + +`RecordDriverType` is the native contract for value/document/database persistence. + +Required logical operations: ```ts -interface RecordStoreType { +interface RecordBackendType { get(path): Promise; set(record): Promise; delete(path): Promise; @@ -133,284 +137,251 @@ interface RecordStoreType { } ``` -This complete-record path is enough for memory, Web Storage, RxDB, unstorage, and SQL-backed integrations. File records use -base64 because the same durable shape must round-trip through JSON-oriented stores. Base64 is a compatibility format, not a claim -that every record backend is suitable for large binaries. - -A store with a more capable physical layout can add optional lanes without implementing the filesystem facade again: +Optional byte lanes can avoid the generic base64 fallback: ```text -stat metadata without file body -readFile direct/range byte read -openReadStream backpressure-preserving logical stream -writeFile selected direct materialized modes -writeStream selected direct stream modes +stat +readFile +openReadStream +writeFile +writeStream ``` -`RecordStoreCapabilitiesType` declares `rangeRead`, `streamRead`, `writeModes`, and `streamWriteModes`. The record adapter turns -only those declared lanes into native adapter capabilities. If a lane is absent, the complete-record implementation remains the -fallback. This is the extension point for third-party KV/document stores that can do better than one large JSON-shaped record. +Record capability metadata also identifies: -`createMemoryAdapter()` uses the complete-record path for deterministic tests and temporary data. +```text +replacement atomic | best-effort +binary native binary storage available +transactions driver has provider transaction behavior relevant to its writes +``` -`createLocalStorageAdapter(storage)` accepts an injected Web Storage object. Web Storage is synchronous, quota-limited, and -string-only underneath the adapter. It does not claim a portable maximum item size or native streaming. +A custom record driver uses `defineRecordDriver()` and then `createRecordAdapter()`. -`openIndexedDbAdapter()` can open its own IndexedDB database, while `createIndexedDbAdapter(database)` can borrow an existing -one. The store is keyed by canonical `path` and indexed by `parent` so direct directory listing stays indexed. Ownership remains -with the caller unless the adapter option explicitly transfers it. +The generic record adapter does not claim native streaming. If the driver does not provide `writeStream()`, the facade +can materialize an input only under `maxBufferedWriteBytes`. -`createCacheAdapter(cache)` stores records in an injected `Cache` using synthetic request URLs. No network request is made. Cache -Storage quota, eviction, and persistence policy remain browser decisions. +### Memory -Deno KV uses an explicit partition layout ------------------------------------------ +The memory driver is deterministic and dependency-free. It is useful for tests, examples, and temporary state. It is not +durable storage. -`createDenoKvAdapter(kv)` accepts an already-open Deno KV database. Deno KV has a 2 KiB serialized key limit and a 64 KiB -serialized value limit, so treating one filesystem file as one KV value would create a small and surprising file ceiling. The -default adapter policy is `partition: "auto"`. +### Deno KV + +The Deno KV driver has a provider-aware partition layout. Important policy options are: ```text -logical entry key - [prefix, "entry", parentPath, name] +partition auto | always | never +partBytes decoded bytes per raw part +inlineBytes decoded body budget for one inline record +maxParts logical-file part ceiling +concurrency bounded physical part I/O +``` -list one parent - prefix [prefix, "entry", parentPath] - -> direct children only +`partBytes` and `inlineBytes` are intentionally smaller than Deno KV's serialized value ceiling. The provider limit +applies after serialization, so accepting the full provider number as decoded application bytes would be unsafe. -small file - entry -> normal FileRecord +The driver planner also evaluates the concrete path against a conservative serialized-key estimate before provider I/O. -large file - [prefix, "part", canonicalPath, generation, 0] - [prefix, "part", canonicalPath, generation, 1] - ... - entry -> manifest committed last -``` +`DenoKvDriverType.collect()` performs explicit maintenance for crash-left physical generations. It scans only the +private part namespace, retains the published generation, ignores recent unpublished generations for a one-hour grace +period by default, and stops after the caller's deletion budget. Ordinary reads and writes never start this scan +implicitly. -Default decoded sizes are 32 KiB inline and 48 KiB per raw binary part. `maxParts` defaults to 10,000 and part I/O concurrency -to 8. Callers can set `partition: "never" | "auto" | "always"`, `inlineBytes`, `partBytes`, `maxParts`, and `concurrency`. The -adapter exposes these as inspectable limits/partition policy. +### localStorage -Manifest-last publication is the visibility rule. A reader sees the previous complete generation until all new parts exist and -the new manifest is stored. A process crash before the manifest commit can leave unreachable new-generation parts. That is a -storage leak, not a partially visible logical file. The adapter does not currently run a global orphan scavenger because doing so -would require a separate ownership/retention policy. +The localStorage driver maps canonical records into a private key prefix. It inherits Web Storage's synchronous +underlying API, but the package presents the normal asynchronous driver contract to keep the storage stack composable. -Large-file hot paths avoid generic reconstruction: +Applications should treat browser quota as dynamic. The driver does not invent a stable quota number. -- `stat()` reads the entry/manifest only. -- `list()` uses the direct-parent key prefix, so it reads direct-child metadata only and never scans descendant entry keys or body parts. -- range reads fetch only overlapping parts. -- stream reads load one physical part at a time under consumer backpressure. -- materialized append/update builds a new generation part-by-part and never joins the previous large file into one record. -- streamed replacement writes parts with bounded concurrency and publishes the manifest last. +### IndexedDB -Append/update still copy the untouched logical bytes into a new immutable generation because Deno KV has no provider-side range -copy primitive. The copy is bounded by `partBytes` and `concurrency`; the tradeoff is provider I/O proportional to the resulting -file size rather than JavaScript memory proportional to that size. Streamed append/update is not advertised as a direct lane, -so an incoming stream must still fit under the facade `maxBufferedWriteBytes` ceiling before this bounded patch path runs. +The IndexedDB driver borrows or owns an injected database according to options. It uses an object store and a parent +index for direct-child listing. The application remains responsible for database versioning/upgrades outside the driver +unless ownership is explicitly transferred. -`partition: "never"` disables the partitioned streaming write lane. A large streamed write then follows the facade's normal -bounded materialization rule and fails `too-large` once it exceeds `maxBufferedWriteBytes`. This gives applications a deliberate -way to reject the changed durable layout. Deno KV remains an unstable Deno API, so real integration tests run with -`--unstable-kv`. +### Cache Storage -`createSqliteAdapter(database)` accepts a small connected SQLite statement interface. It deliberately reuses the same SQL record -mapping used by the SQLite branch of `createDb0Adapter()` instead of creating a second schema and upsert implementation. +The Cache driver stores records under private request URLs. Cache Storage is a record/value persistence mechanism here, +not an HTTP cache policy abstraction. The driver only interprets entries in its private namespace. -Existing ecosystem adapters stay above the abstraction the application already owns: +### unstorage -```text -unstorage Storage -> RecordStoreType -RxDB RxCollection -> RecordStoreType -db0 Database -> RecordStoreType -Drizzle DB + table -> RecordStoreType -``` +`createUnstorageDriver(storage)` consumes the high-level unstorage `Storage` object. This deliberately sits above +whichever unstorage provider driver the application selected. -Bridge descriptors group these forward adapters with reverse drivers when a real reverse contract exists. They do not fabricate -a reverse direction for RxDB, db0, or Drizzle. See [ecosystems.md](./ecosystems.md). +Use `bridge/unstorage` for the opposite direction, where an existing `FileSystemType` must satisfy unstorage's Driver +contract. -Object stores keep object-store semantics visible -------------------------------------------------- +### RxDB -`ObjectStoreType` is the common client contract for S3, Azure Blob, and custom object storage. It models the operations those -systems actually have: +`createRxDbDriver(collection)` targets `RxCollection`, not `RxStorage`. RxDB keeps responsibility for its selected +RxStorage, conflict mechanics, wrappers, replication, multi-instance behavior, and licensing. -```text -HEAD exact key -GET full object or range -PUT replacement -DELETE exact key -LIST prefix + delimiter -COPY inside provider, when supported -``` +The package exports `RxDbRecordJsonSchema` for the collection used by this integration. `path` is the primary key and +`parent` is indexed for direct-child listing. -Its metadata includes byte size, media type, last modification time, ETag, provider version identity, and user metadata. Its -capabilities state whether range read, streaming read, streaming replacement, provider-side copy, and conditional writes are -really available. +### db0 -`createObjectAdapter()` maps that model to filesystem paths. A normal file maps to one object key. An empty directory maps to a -trailing-slash marker with private metadata, and prefix listing recognizes both those markers and foreign provider prefixes. +`createDb0Driver(database)` targets the db0 `Database` contract and its reported dialect. The current SQL branches are +SQLite, libSQL, PostgreSQL, and MySQL. -```text -/photos -> photos/ -/photos/a.jpg -> photos/a.jpg -/photos/2026/b.jpg -> photos/2026/b.jpg -``` +The driver owns its filesystem-record table only when configured to initialize it. The injected database is borrowed +unless `disposeDatabase` is true. -A raw object namespace can physically contain both `mixed` and `mixed/child`. The filesystem view resolves the exact `mixed` -object as the file because `stat("/mixed")` does the same. That rule keeps read, stat, and write behavior internally consistent -when foreign object layouts do not obey filesystem restrictions. +### Drizzle -Replacement can stream when the provider supports it. Append and update cannot normally mutate object bytes in place, so the -adapter performs: +`createDrizzleDriver({ database, table })` accepts a caller-owned connected Drizzle database and a caller-owned table +shape. Drizzle is not one SQL dialect, so this generic driver does not own DDL or migrations. + +Required logical columns are: ```text -HEAD current object - | - v -GET current bytes - | - v -apply append/update in memory - | - v -conditional PUT replacement +path +parent +name +kind +data +size +lastModified +mediaType ``` -When conditional writes are enabled and an existing object does not return an ETag, the adapter fails rather than quietly -performing an unsafe read-modify-write. When the file is being created through append/update, it uses create-only semantics where -the provider exposes them. +The portable replacement route is delete then insert, so the generic driver reports best-effort replacement rather than +claiming cross-process atomicity. A dialect-specific future driver can expose stronger transaction/upsert/binary +behavior. -The S3 client implements the protocol directly ------------------------------------------------ +### SQLite rows -`createS3Client()` uses Web Fetch, Web Crypto, `@std/encoding`, and `@std/xml`. It does not depend on the AWS SDK. +`createSqliteDriver(database)` stores filesystem rows inside an already connected SQLite database. This is the +**database-backed filesystem** direction. -```ts -import { createS3Client } from "@okikio/opfs/s3"; -import { createS3Adapter } from "@okikio/opfs/adapter/s3"; - -const client = createS3Client({ - endpoint: "https://s3.us-east-1.amazonaws.com", - bucket: "example", - region: "us-east-1", - credentials: async () => await credentials.get(), - addressing: "path", - concurrency: 4, -}); +It is not a SQLite VFS and does not make SQLite store its database file on `FileSystemType`. See `ecosystems.md` for +that opposite direction. -const adapter = createS3Adapter(client, { prefix: "app" }); -``` - -Signature Version 4 includes the request authority in canonical headers. Browser Fetch forbids application code from setting -the `Host` header, so the client signs `url.host` while leaving actual Host or `:authority` transmission to Fetch. Static browser -credentials are usually a security mistake; browser deployments should use appropriately scoped short-lived credentials or a -trusted service design. +## Object drivers -A streamed replacement uses multipart upload with bounded part concurrency: +`ObjectDriverType` preserves object storage concepts required for efficient translation: ```text -ReadableStream - | - v -fixed-size chunker - | - +--> UploadPart 1 --+ - +--> UploadPart 2 --+--> CompleteMultipartUpload - +--> UploadPart N --+ - | - failed operation - | - `--> wait active parts -> AbortMultipartUpload +stat object +get bytes/range +put bytes/stream +list prefix +remove object +native copy when available ``` -The client applies `If-Match` and `If-None-Match` to multipart completion, which is the commit operation that current S3 exposes -for these preconditions. It waits for already-started parts before aborting so a late part cannot arrive after the abort request. +An object driver can also report object-specific capability details, provider limits, continuation behavior, +partition/upload policy, and physical metrics. -S3 has two failure cases that are easy to miss. `CompleteMultipartUpload` can return HTTP 200 and then stream an XML error, and -`CopyObject` can also return an embedded error in HTTP 200. The client parses and rejects both bodies. +### S3 -`CopyObject` has a 5 GB source limit. Larger copies use `UploadPartCopy` ranges into a multipart destination. The copy part size -increases when necessary to stay within S3's 10,000-part limit. Source bytes stay inside the object provider rather than crossing -JavaScript memory or network twice. +The S3 client owns REST, SigV4, request policy, multipart operations, copy, listing, and protocol errors. -The filesystem copy contract intentionally preserves file bytes, media type, and user metadata where the client can do so. It -does not claim to clone every S3 control-plane property such as ACLs, tags, object-lock state, or every checksum policy. Use the -low-level signed `client.request()` API when the S3 object itself, rather than its filesystem view, is the thing being managed. +`createS3Driver(client)` adds backend capability/limit/optimization inspection. `createS3Adapter(driver)` translates +object keys and directory prefixes into filesystem primitives. -S3-compatible does not mean behavior-identical. Configure the client from the selected provider's current contract: +The client optimizations include independently controllable delayed multipart promotion and derived signing-key caching. +See `s3.md`. -- Cloudflare R2 commonly uses the `auto` region and has provider-specific supported/unsupported S3 operations. -- DigitalOcean Spaces implements a compatible subset rather than every AWS S3 feature. -- Google Cloud Storage's XML multipart API documents different precondition behavior; disable `conditionalWrite` when the - selected path does not provide the safety contract expected by the object adapter. -- Other compatible providers should be treated the same way: verify endpoint, signing region, addressing, copy, conditional - requests, multipart limits, checksums, and error behavior before enabling a capability flag. +### Azure Blob -The Azure Blob client keeps Azure's own model --------------------------------------------- +The Azure client owns Blob REST, authentication, block upload, server-side copy, listing, and provider errors. -`createAzureClient()` also uses Web Fetch and `@std/xml`, but it does not force Azure Blob through an S3-shaped client. +`createAzureDriver(client)` adds backend inspection. `createAzureAdapter(driver)` supplies filesystem translation. Block +upload and server copy are independently disableable. See `azure.md`. + +## Adapter contract + +`AdapterType` always references the driver it translates: ```ts -import { createAzureClient } from "@okikio/opfs/azure"; -import { createAzureAdapter } from "@okikio/opfs/adapter/azure"; +interface AdapterType { + readonly name: string; + readonly driver: DriverType; + readonly capabilities: AdapterCapabilitiesType; + // filesystem primitives... +} +``` -const client = createAzureClient({ - endpoint: "https://account.blob.core.windows.net", - container: "example", - credential: { kind: "sas", token }, -}); +`AdapterCapabilitiesType` describes native adapter routes, not every public filesystem operation. -const adapter = createAzureAdapter(client, { prefix: "app" }); +```text +read +write +streamRead +streamWriteModes +rangeRead +nativeCopy +nativeMove +positionalWrite +syncAccess ``` -Credentials can be SAS, refreshable bearer tokens, or a custom header callback. The service version is explicit and defaults to -the version pinned by this package. Block-size limits are selected from that service version rather than one timeless constant. +The facade can still emulate operations. `FileSystemType.inspect().support` is the authority for the effective route +after adapter capabilities and facade optimization policy are composed. -Streamed replacements use Put Block followed by Put Block List. Azure has no abort call for uncommitted blocks, so a failure -waits for already-started requests, leaves the old committed blob untouched, and allows Azure to garbage-collect the uncommitted -blocks later. +Adapters may retain compact `limits` or `partition` summaries for translation diagnostics. Detailed backend limits and +their provenance live on the driver. -Copy Blob From URL has a smaller synchronous copy limit. Larger files use Put Block From URL ranges followed by Put Block List. -This keeps large copies server-side without hiding a size cliff behind `nativeCopy: true`. +## Cancellation -Custom adapters must preserve the same invariants -------------------------------------------------- +Every async backend method that accepts `AbortSignal` checks it before expensive work and between bounded chunks. A +failed or aborted stream write cancels the upstream producer when practical. -Use `defineAdapter()` for a backend that already exposes filesystem-like primitives: +Provider cleanup can outlive the caller signal. Protocol drivers/clients use a separate bounded cleanup signal when an +already aborted caller signal would make cleanup impossible. -```ts -import { defineAdapter } from "@okikio/opfs/adapter"; +## Ownership -export const adapter = defineAdapter({ - name: "provider", - capabilities: { - read: true, - write: true, - streamRead: false, - streamWriteModes: [], - rangeRead: false, - nativeCopy: false, - nativeMove: false, - positionalWrite: false, - syncAccess: false, - }, - async stat(path) { /* ... */ }, - async readFile(path, options) { /* ... */ }, - async writeFile(path, bytes, options) { /* ... */ }, - async *readDir(path, options) { /* direct children only */ }, - async createDir(path, options) { /* parent already exists */ }, - async remove(path, options) { /* file or empty directory */ }, -}); -``` +Injected resources are borrowed by default. + +Examples of explicit ownership transfer: -Every adapter receives canonical virtual paths. It must respect requested ranges and write modes. It must yield direct children, -not recursive descendants, from `readDir()`. It must not configure logging, inspect process environment, or open unrelated -resources during module evaluation. +```text +disposeDatabase +disposeStorage +disposeDriver +disposeAdapter +``` -Injected resources are borrowed by default. Transfer ownership only through an explicit option such as `disposeDatabase`, -`disposeStore`, or `disposeAdapter`. A filesystem closing must never surprise another subsystem by disposing infrastructure that -it still owns. +A convenience adapter that creates a driver internally transfers ownership of that newly created driver to the adapter. +A caller that creates a driver explicitly can choose whether the adapter should dispose it. A driver exposes backend +disposal only when its own construction options transferred backend ownership; disposing an adapter therefore cannot +close a resource that the driver only borrowed. + +`driver.inspect().ownership` reports that relationship as `none`, `borrowed`, or `owned`. `driver.inspect().provides` +reports the stable backend operations or capabilities available on the configured driver. Higher layers can therefore +explain backend ownership and breadth without inferring either from adapter flags. + +A configured record driver can also be read-only. In that mode `driver.capabilities.write` is false, write primitives +are not exposed, and direct `set()`/`delete()` calls fail before backend mutation. The adapter reflects the same state +instead of relying on adapter-only policy. + +## Import safety + +Runtime and provider code stays behind explicit subpaths. Importing the package root does not: + +- import Node/Bun/Deno-only modules; +- resolve credentials; +- open a database; +- connect to a network endpoint; +- configure global logging; +- start worker/process resources. + +## Extension checklist + +Before adding a driver/adapter, verify: + +1. The driver is independently meaningful without `FileSystemType`. +2. Provider requirements and known limits are structured and attributed. +3. Unknown limits stay unknown rather than being treated as unlimited. +4. Observable optimizations can be disabled. +5. The driver planner can reject known bad inputs before I/O. +6. Every advertised direct operation has a real implementation. +7. Large work has explicit byte/part/concurrency/retry bounds. +8. Resource ownership is explicit. +9. The adapter contains translation, not duplicated provider behavior. +10. Tests exercise the driver directly and through the adapter/facade. +11. Benchmarks include the backend/client baseline and each added layer. diff --git a/docs/api.md b/docs/api.md index 62a4832..2c4fd67 100644 --- a/docs/api.md +++ b/docs/api.md @@ -1,626 +1,780 @@ -Public API guide -================ +# Public API guide -This guide is organized by developer task. Exact low-level schemas and types are also available through the explicit package subpaths. +## Purpose -Open or create a filesystem ---------------------------- +The public API is organized by layer. Normal application code can use the root filesystem facade and convenience +adapters. Storage libraries and infrastructure code can use the explicit client, driver, adapter, bridge, and +integration subpaths. + +```text +client -> driver -> adapter -> FileSystemType -> bridge +``` + +No public constructor requires a global registry. + +## Filesystem root ### `openFileSystem(options?)` -Opens the browser's native Origin Private File System and returns `FileSystemType`. +Opens the current browser realm's native Origin Private File System and returns `FileSystemType`. ```ts import { openFileSystem } from "@okikio/opfs"; -const fileSystem = await openFileSystem(); +await using fileSystem = await openFileSystem(); +await fileSystem.writeFile("/state.json", "{}", { parents: true }); ``` -Use this only when native browser OPFS is the chosen backend. Server runtimes should create a runtime adapter and pass it to `createFileSystem()`. +This convenience path composes native OPFS root -> OPFS driver -> OPFS adapter -> filesystem facade. ### `createFileSystem(adapter, options?)` Creates the adapter-independent facade. ```ts -const fileSystem = createFileSystem(adapter, { - coordination: "auto", - lockPrefix: "my-app:filesystem", - maxBufferedWriteBytes: 64 * 1024 * 1024, - metrics: "basic", - optimizations: { nativeCopy: true }, - disposeAdapter: false, -}); +import { createFileSystem } from "@okikio/opfs"; +import { createNodeAdapter } from "@okikio/opfs/adapter/node"; + +const fileSystem = createFileSystem( + createNodeAdapter({ root: "./data" }), + { + coordination: "local", + maxBufferedWriteBytes: 64 * 1024 * 1024, + metrics: "basic", + disposeAdapter: true, + }, +); ``` -`FileSystemOptionsType`: +Important `FileSystemOptionsType` fields: -- `coordination`: `auto`, `web-locks`, `local`, or `none`. -- `lockPrefix`: stable lock namespace used for cooperating filesystem facades. -- `maxBufferedWriteBytes`: maximum stream/file size the facade may materialize for a fallback route. -- `metrics`: `none`, `basic`, or `timing`. `none` is intended for baseline overhead measurements. -- `optimizations`: partial override for `streamRead`, `streamWrite`, `rangeRead`, `nativeCopy`, and `nativeMove`. Every route defaults to enabled. -- `disposeAdapter`: transfers adapter disposal ownership to the facade when true. +- `coordination`: `auto | web-locks | local | none`; +- `lockPrefix`: stable namespace for cooperating facade locks; +- `maxBufferedWriteBytes`: maximum facade-owned materialization for fallback routes; +- `metrics`: `none | basic | timing`; +- `optimizations`: independent facade route switches; +- `disposeAdapter`: transfer adapter ownership to the facade. -`coordination` is runtime-validated by `CoordinationModeSchema`. +## Filesystem methods +`FileSystemType` exposes the portable API: -Inspect and plan before I/O ---------------------------- +```text +getDirectoryHandle +getFileHandle +getFile +stat +exists +mkdir +ensureDir +ensureFile +readFile +readText +openReadStream +writeFile +readDir +walk +copy +move +remove +emptyDir +openWritableFile +openSyncFile +inspect +plan +getMetrics +close +root +``` + +All path methods use the canonical virtual namespace. Public input is normalized before adapter/driver calls. -### `inspect()` +## Read APIs -Returns a synchronous `InspectionType` for the configured filesystem stack. It contains: +### `readFile(path, options?)` + +Returns `Uint8Array`. + +Options include: ```text -adapter -native capabilities -effective support -known hard limits -optional partition layout -resolved optimization policy -maxBufferedWriteBytes -metrics mode and current metrics snapshot +at zero-based offset +length maximum bytes after at +signal cancellation ``` -Effective support uses `native | emulated | partitioned | unsupported`. Native capability flags never include facade emulation. -This makes a caller able to distinguish a fast provider range read from a full-read-and-slice fallback, or a Deno KV partitioned -stream from a complete-record materialization. +A driver/adapter with native range support can avoid materializing the complete file. Otherwise the facade can emulate +the range and reports that route through `inspect()`/`plan()`. -### `plan(input)` +### `readText(path, options?)` -Creates a deterministic `PlanType` without touching the backend. Supported operations are `read`, `write`, `copy`, and `move`. -For writes, `size` is the resulting logical file size and is used for `maxFileBytes` and partition-count checks. `inputBytes` is -the number of bytes supplied by the current write and is used for facade stream-buffer admission. For replace, omitted -`inputBytes` falls back to `size` because those values are normally equal. Append/update callers should provide both when known. +Reads file bytes and decodes text. UTF-8 is the default encoding. -```ts -const plan = fileSystem.plan({ - operation: "write", - source: "stream", - mode: "replace", - size: 200 * 1024 * 1024, - inputBytes: 200 * 1024 * 1024, -}); +### `openReadStream(path, options?)` -if (!plan.supported) throw new Error(plan.reasons.join(" ")); -``` +Returns `ReadableStream`. -An unknown limit remains unknown. The planner does not invent provider guarantees. For an unknown-size emulated stream, it warns -that `maxBufferedWriteBytes` remains the runtime admission limit. +When a native stream exists and `optimizations.streamRead` is enabled, the facade forwards it. Otherwise a readable +backend can be adapted to a materialized stream. `inspect().support.streamRead` identifies the effective route. -### `getMetrics()` +## Write APIs -Returns a detached `MetricsType`. `basic` tracks counts, failures, bytes, native/emulated/partitioned route counts, current -facade-owned buffered bytes, and peak buffered bytes. `timing` additionally records total and maximum durations. `none` keeps the -hot-path collector inactive. +### `writeFile(path, data, options?)` -Optimization controls can force a fallback for differential tests or application policy. A disabled optimization is never -relabelled native. If the fallback cannot satisfy the request within `maxBufferedWriteBytes`, runtime execution and `plan()` both -fail with the same `too-large`/unsupported condition when size is known. +Accepted input: -Path API --------- +```text +string +Blob +ArrayBuffer +ArrayBufferView +ReadableStream +AsyncIterable +``` -### `getDirectoryHandle(path, options?)` +Write modes: -Returns a package `DirectoryHandleType` for one directory. +```text +replace +append +update +``` -Options: +Important options: -- `create`: create exactly that directory when absent. -- `recursive`: create missing ancestors as well. -- `signal`: abort before commit. +```text +at +truncate +parents +mediaType +signal +``` -The virtual root `/` always exists. +A streaming source uses a driver-native stream route only when the selected write mode supports it and the route is +enabled. Otherwise the facade materializes the stream under `maxBufferedWriteBytes`. Crossing that limit cancels the +producer when possible and fails with `too-large`. -### `getFileHandle(path, options?)` +## Directory and tree APIs -Returns a package `FileHandleType`. +### `readDir(path, options?)` -Options: +Returns a lazy direct-child iterator. -- `create`: create the file when absent. -- `parents`: create missing parent directories. -- `signal`: abort before commit. +### `walk(path, options?)` -A read-only lookup never creates a file or directory. +Returns a lazy recursive iterator. Options control depth, root inclusion, file inclusion, directory inclusion, and +cancellation. The facade does not collect the complete tree first. -### `getFile(path, options?)` +### `copy(source, destination, options?)` -Returns a `File` snapshot. Changes written later are not reflected in the already-returned File object. +Copies one file or directory tree. The facade uses native/server-side copy when the adapter provides it and the route is +enabled. Otherwise it composes read/write work with bounded concurrency. -### `stat(path, options?)` +### `move(source, destination, options?)` -Returns: +Uses native move when available. The fallback is copy then remove and is not atomic. `plan()` returns a structured +warning for that route. -```ts -type StatType = FileStatType | DirectoryStatType; -``` +### `remove(path, options?)` -File stat includes canonical path, name, size, last-modified milliseconds, and media type. Directory stat includes canonical path, name, and last-modified when the adapter provides it. +Removes one entry or recursively removes descendants when requested. The virtual root cannot be removed. -### `exists(path, options?)` +### `emptyDir(path?, options?)` -Returns an advisory boolean. `kind` can restrict the answer to `file` or `directory`. +Removes children while keeping the directory. Root is the default. -Do not use `exists()` as a substitute for operation error handling. Another context can mutate the backend after the check. +## OPFS-shaped handles -### `mkdir(path, options?)` +Every facade exposes `root: DirectoryHandleType`. -Creates one directory. `recursive: true` creates missing ancestors. +`DirectoryHandleType`: -### `ensureDir(path, options?)` +```text +kind +name +path +getDirectoryHandle +getFileHandle +removeEntry +resolve +entries +keys +values +isSameEntry +Symbol.asyncIterator +``` -Ensures a directory and its parents exist. A file at the same path produces `type-mismatch`. +`FileHandleType`: -### `ensureFile(path, options?)` +```text +kind +name +path +getFile +createWritable +createSyncAccessHandle +isSameEntry +``` -Ensures an empty file exists and creates its parent directories. +These are package facades, not native browser handle instances. `path` is a package-specific canonical virtual path. -Read APIs ---------- +## Writable files -### `readFile(path, options?)` +`createWritable()` returns `WritableFileStreamType` with OPFS-style commands: -Returns `Uint8Array`. +```ts +await writable.write(bytes); +await writable.write({ type: "write", position: 10, data: bytes }); +await writable.write({ type: "seek", position: 20 }); +await writable.write({ type: "truncate", size: 100 }); +await writable.close(); +``` -Options: +The staged image commits on close and is discarded on abort. Large sequential writes should prefer `writeFile()` because +that path can select a native streaming driver route. -- `at`: zero-based byte offset. -- `length`: maximum bytes after `at`. -- `signal`: cancellation. +`openWritableFile()` exposes a direct asynchronous positional resource only when the adapter reports that capability. -### `readText(path, options?)` +## Synchronous files -Reads bytes and decodes them. `encoding` defaults to UTF-8. +`openSyncFile(path, options?)` returns `SyncFileType` only when the selected adapter exposes native synchronous random +access. -### `openReadStream(path, options?)` +Operations: -Returns `ReadableStream`. +```text +read +write +writeAll +getSize +truncate +flush +close +``` -If the adapter provides native stream reads, the facade forwards them. Otherwise it creates a stream from the adapter's materialized read result. Cancellation remains connected after the stream opens. +The facade keeps the path lock for the complete resource lifetime. `writeAll()` handles partial native writes. -Write API ---------- +## Inspection -### `writeFile(path, data, options?)` +### `fileSystem.inspect()` -Accepted `WriteDataType` values: +Returns `InspectionType`: -```text -string -Blob -ArrayBuffer -ArrayBufferView -ReadableStream -AsyncIterable +```ts +const inspection = fileSystem.inspect(); + +inspection.driver; +inspection.adapter; +inspection.support; +inspection.optimizations; +inspection.maxBufferedWriteBytes; +inspection.metricsMode; +inspection.metrics; +inspection.driverMetrics; ``` -Options: +`inspection.driver` contains: -- `mode`: `replace` (default), `append`, or `update`. -- `at`: starting byte offset for update mode. -- `truncate`: truncate at the final cursor. -- `parents`: create missing parents. -- `mediaType`: metadata for record/native adapters that can preserve it. -- `signal`: cancellation. +- `provides`: stable operation/capability identifiers exposed by this configured driver; +- `ownership`: `none`, `borrowed`, or `owned` for the long-lived backend resource; +- backend-native requirements and their current availability; +- provenance-aware provider, implementation, user, and probe limits; +- independently controllable driver optimization state. -The mode is runtime-validated by `WriteModeSchema`. +`provides` is intentionally an open string vocabulary. A third-party driver can add a provider-specific operation +without waiting for a core enum revision. The stable core driver families still expose typed operational methods +separately. -A non-streaming adapter buffers stream input up to `maxBufferedWriteBytes`. Crossing the limit cancels the producer and throws `too-large`. +`inspection.adapter` contains the translation layer's native route flags and compact translation summaries. -Long-lived positional output ----------------------------- +`inspection.support` contains effective routes after facade fallback and optimization policy: -### `openWritableFile(path, options?)` +```text +native +emulated +partitioned +unsupported +``` -Returns `WritableFileType` only when `adapter.capabilities.positionalWrite` is true. The facade does not emulate this operation with repeated `writeFile(..., { mode: "update" })` calls because a record-backed adapter can otherwise rematerialize the complete file for every chunk. +`inspection.metrics` is logical facade work. `inspection.driverMetrics`, when present, is physical backend/protocol +work. -Options: +## Planning -- `create`: create an empty file when it is absent. -- `parents`: create missing parent directories when `create` is true. -- `signal`: cancel ordinary work before commit. Cleanup through `close()` or `abort()` still releases the owned resource after cancellation. +### `fileSystem.plan(input)` -The returned resource owns the file mutation lock for its complete lifetime. The create/check/open sequence occurs under that same lock, so another mutation cannot enter between file creation and adapter open. +Preflights a concrete request without touching storage. + +Supported high-level operations are currently: + +```text +read +write +copy +move +``` + +Example: ```ts -const file = await fileSystem.openWritableFile("/media/output.mp4", { - create: true, - parents: true, +const plan = fileSystem.plan({ + operation: "write", + path: "/archive.bin", + source: "stream", + size: 800 * 1024 * 1024, + inputBytes: 800 * 1024 * 1024, + mode: "replace", }); -try { - await file.write(header, { at: 0 }); - await file.write(chunk, { at: chunkOffset }); - await file.flush(); - await file.close(); -} catch (error) { - await file.abort(error); - throw error; +if (!plan.supported) { + console.log(plan.problems); + console.log(plan.actions); } ``` -`write()` is positional, `truncate()` changes byte length, and `flush()` requests backend durability without closing. `close()` and `abort()` are idempotent terminal operations. A browser OPFS writable can discard its staged image on abort. Host filesystems generally cannot roll back bytes already written, so an application that needs publish-on-success semantics should write a staging path and move it after close. +`PlanType` includes: -Record/database adapters report `positionalWrite: false`. They remain valid for ordinary materialized writes and bounded stream buffering, but they are not presented as a large-file positional output path. +```text +operation +supported +support +driver +bufferBytes +partBytes +parts +problems[] +actions[] +``` -Directory iteration -------------------- +Problems have a stable code, layer, severity, message, and optional referenced limit. Actions have a stable kind and +optional code/detail. -### `readDir(path, options?)` +## Driver API -Lazy direct-child iterator. +`@okikio/opfs/driver` exports the generic driver definition model: -```ts -for await (const entry of fileSystem.readDir("/projects")) { - console.log(entry.kind, entry.name, entry.path); -} -``` +- `ProblemLayerSchema` / `ProblemLayerType`; +- `ProblemSeveritySchema` / `ProblemSeverityType`; +- `ActionKindSchema` / `ActionKindType`; +- `ProblemSchema` / `ProblemType`; +- `ActionSchema` / `ActionType`; +- `DriverOperationSchema` / `DriverOperationType`; +- `DriverPlanInputSchema` / `DriverPlanInputType`; +- `DriverPlanSchema` / `DriverPlanType`; +- `DriverInspectionSchema` / `DriverInspectionType`; +- `DriverType`; +- `DefineDriverOptionsType`; +- `defineDriver()`. -### `walk(path, options?)` +`defineDriver()` validates definition metadata. Concrete storage should normally use one of the family contracts below. -Lazy recursive iterator. +## File-driver API -Options: +`@okikio/opfs/driver/file` exports: -- `maxDepth`: maximum depth below the requested path. -- `includeRoot`: include the requested path itself before descendants. -- `includeFiles`: yield files. Defaults to true. -- `includeDirectories`: yield directories. Defaults to true. -- `signal`: cancel traversal between yielded entries. +- `FileDriverCapabilitiesSchema` / `FileDriverCapabilitiesType`; +- `FileDriverSignalOptionsType`; +- `FileDriverReadOptionsType`; +- `FileDriverWriteOptionsType`; +- `FileDriverCopyOptionsType`; +- `FileDriverMoveOptionsType`; +- `FileDriverDirectoryEntryType`; +- `FileDriverFileStatType`; +- `FileDriverDirectoryStatType`; +- `FileDriverStatType`; +- `FileDriverWritableFileType`; +- `FileDriverSyncFileType`; +- `FileDriverType`; +- `FileBackendType`; +- `DefineFileDriverOptionsType`; +- `defineFileDriver()`. -The iterator does not eagerly collect the entire tree. +The associated adapter is `createFileAdapter(driver)` from `@okikio/opfs/adapter/file`. -Structural operations ---------------------- +## Record-driver API -### `copy(source, destination, options?)` +`@okikio/opfs/driver/record` exports: -Copies one file or tree. Directory file bodies use bounded `concurrency`, default 4. +- `RecordListType`; +- `RecordReplacementSchema` / `RecordReplacementType`; +- `RecordDriverCapabilitiesSchema` / `RecordDriverCapabilitiesType`; +- `RecordBackendType`; +- `RecordDriverType`; +- `DefineRecordDriverOptionsType`; +- `defineRecordDriver()`. -`overwrite: true` replaces the destination tree instead of merging stale entries into it. +The associated adapter is `createRecordAdapter(driver)` from `@okikio/opfs/adapter/record`. -Source and destination cannot be the same path or ancestors of each other. +## Object-driver API -### `move(source, destination, options?)` +`@okikio/opfs/driver/object` exports: -Uses adapter-native move when `nativeMove` is true. Otherwise calls copy then remove. The fallback is not atomic. +- `ObjectDriverCapabilitiesSchema` / `ObjectDriverCapabilitiesType`; +- `ObjectStatType`; +- `ObjectEntryType`; +- `ObjectListType`; +- `ObjectGetOptionsType`; +- `ObjectPutOptionsType`; +- `ObjectCopyOptionsType`; +- `ObjectListOptionsType`; +- `ObjectBackendType`; +- `ObjectDriverType`; +- `DefineObjectDriverOptionsType`; +- `defineObjectDriver()`. -### `remove(path, options?)` +The associated adapter is `createObjectAdapter(driver)` from `@okikio/opfs/adapter/object`. -Removes one file or empty directory. `recursive: true` removes descendants first. +## Adapter API -The virtual root cannot be removed. +`@okikio/opfs/adapter` exports: -### `emptyDir(path?, options?)` +- `AdapterType`; +- `FileSystemOptionsType`; +- `defineAdapter()`. -Removes children while retaining the directory. `path` defaults to `/`. Child removals use bounded concurrency. +Every `AdapterType` has a `driver` reference. Adapter capability flags describe direct translation routes. They do not +replace `driver.inspect()` or `FileSystemType.inspect()`. -Synchronous file API --------------------- +## First-party file constructors -### `openSyncFile(path, options?)` +### Browser OPFS -Returns `SyncFileType` only when `adapter.capabilities.syncAccess` is true. +```ts +import { createOpfsDriver } from "@okikio/opfs/driver/opfs"; +import { createOpfsAdapter } from "@okikio/opfs/adapter/opfs"; -Options: +const root = await navigator.storage.getDirectory(); +const driver = createOpfsDriver(root); +const adapter = createOpfsAdapter(root); // convenience path creates its own driver +``` -- `create` -- `parents` -- `signal` +`createOpfsDriver(root)` returns `OpfsDriverType`. `createOpfsAdapter(root)` returns `OpfsAdapterType` and retains +`nativeRoot`. -The resource owns its native file and the facade path lock for the complete lifetime. +### Node ```ts -const file = await fileSystem.openSyncFile("/db.sqlite", { - create: true, - parents: true, -}); +createNodeDriver({ root: "./data" }); +createNodeAdapter({ root: "./data" }); +``` -try { - file.writeAll(bytes, { at: 0 }); - file.flush(); -} finally { - file.close(); -} +Types: `NodeDriverOptionsType`, `NodeAdapterOptionsType`. + +### Deno + +```ts +createDenoDriver({ root: "./data" }); +createDenoAdapter({ root: "./data" }); ``` -`SyncFileType` operations: +Types: `DenoDriverOptionsType`, `DenoAdapterOptionsType`. -```text -read -write -writeAll -getSize -truncate -flush -close +### Bun + +```ts +createBunDriver({ root: "./data" }); +createBunAdapter({ root: "./data" }); ``` -`writeAll()` loops over partial native writes. +Types: `BunDriverOptionsType`, `BunAdapterOptionsType`. -OPFS-shaped handle API ----------------------- +## First-party record constructors -Every `FileSystemType` has `root: DirectoryHandleType`. +### Memory -### Directory handle +```ts +createMemoryDriver(); +createMemoryAdapter(); +``` -```text -kind -name -path -getDirectoryHandle() -getFileHandle() -removeEntry() -resolve() -entries() -keys() -values() -isSameEntry() -[Symbol.asyncIterator]() +### Deno KV + +```ts +const driver = createDenoKvDriver(database, { + partition: "auto", + partBytes: 48 * 1024, + inlineBytes: 32 * 1024, + maxParts: 10_000, + concurrency: 8, +}); ``` -### File handle +Exports include: ```text -kind -name -path -getFile() -createWritable() -createSyncAccessHandle() -isSameEntry() +DENO_KV_MAX_KEY_BYTES +DENO_KV_MAX_VALUE_BYTES +DENO_KV_MAX_ATOMIC_BYTES +DENO_KV_SAFE_PART_BYTES +DENO_KV_SAFE_INLINE_BYTES +DENO_KV_DEFAULT_PART_BYTES +DENO_KV_DEFAULT_INLINE_BYTES +DENO_KV_DEFAULT_MAX_PARTS +DENO_KV_DEFAULT_CONCURRENCY +DENO_KV_DEFAULT_COLLECT_AGE_MS +DENO_KV_DEFAULT_COLLECT_DELETES +DenoKvEntryType +DenoKvType +DenoKvDriverOptionsType +DenoKvCollectOptionsType +DenoKvCollectResultType +DenoKvDriverType ``` -These are package facades, not native browser handle instances. Their `path` property is package-specific. +The specialized `DenoKvDriverType` also exposes `collect(options?)` for bounded, age-gated reclamation of unreachable +physical parts. The adapter path exports the same provider constants and maintenance types plus `createDenoKvAdapter()` +and `DenoKvAdapterOptionsType`. -### `createWritable()` +### localStorage -Returns `WritableFileStreamType`. The staged image commits on close and discards on abort. +`driver/localstorage` exports `LocalStorageType`, `LocalStorageDriverOptionsType`, and `createLocalStorageDriver()`. +`adapter/localstorage` exports `createLocalStorageAdapter()`. -Supported write commands: +### IndexedDB -```ts -await writable.write(data); -await writable.write({ type: "write", position: 10, data }); -await writable.write({ type: "seek", position: 20 }); -await writable.write({ type: "truncate", size: 100 }); -``` +`driver/indexeddb` exports `IndexedDbDriverOptionsType`, `IndexedDbOpenOptionsType`, and `createIndexedDbDriver()`. +`adapter/indexeddb` exports `createIndexedDbAdapter()`. -Blob also has a `type` property, so the implementation identifies a command only when `type` is exactly `write`, `seek`, or `truncate`. +### Cache Storage -Adapter and storage API ------------------------ +`driver/cache` exports `CacheDriverOptionsType` and `createCacheDriver()`. `adapter/cache` exports +`createCacheAdapter()`. -`@okikio/opfs/adapter` exports the backend contract consumed by `createFileSystem()`. The required primitive set stays small: +### unstorage -```text -stat -readFile -writeFile -readDir -createDir -remove -``` +`driver/unstorage` exports `UnstorageStorageType`, `UnstorageDriverOptionsType`, and `createUnstorageDriver()`. +`adapter/unstorage` exports `createUnstorageAdapter()`. + +### RxDB + +`driver/rxdb` exports the structural collection/document/query contracts, `RxDbRecordJsonSchema`, and +`createRxDbDriver()`. `adapter/rxdb` exports `createRxDbAdapter()`. + +### db0 -Optional methods expose stronger native paths: +`driver/db0` exports `Db0PrimitiveType`, statement/database contracts, `Db0DriverOptionsType`, and `createDb0Driver()`. +The constructor is asynchronous because table initialization can perform database I/O. + +`adapter/db0` exports `createDb0Adapter()`. + +### Drizzle + +`driver/drizzle` exports: ```text -openReadStream -> capabilities.streamRead -writeStream -> mode is present in capabilities.streamWriteModes -copy -> capabilities.nativeCopy -move -> capabilities.nativeMove -openWritableFile -> capabilities.positionalWrite -openSyncFile -> capabilities.syncAccess +DrizzleTableType +DrizzleRowType +DrizzleDriverOptionsType +createDrizzleDriver ``` -The exported contract includes `AdapterCopyOptionsType` as well as the signal, read, write, move, stat, directory-entry, -writable-file, sync-file, adapter, and filesystem option types. `defineAdapter()` validates the adapter name and capability -record. It does not add a global registry or change the adapter. +`adapter/drizzle` exports `createDrizzleAdapter()`. -The capability record describes what the backend performs natively. For example, an object adapter can expose -`streamWriteModes: ["replace"]` because replacement can stream to multipart/block upload while append and update still need a -read-modify-write cycle. The facade can emulate operations, but it does not relabel an emulation as native support. +### SQLite -### Record storage +`driver/sqlite` exports: -`@okikio/opfs/adapter/record` is the common translation point for backends that naturally store values, documents, or SQL rows. -It exports `RecordStoreType`, `RecordAdapterOptionsType`, and `createRecordAdapter()`. +```text +SqliteStatementType +SqliteDatabaseType +SqliteDriverOptionsType +createSqliteDriver +``` -A record store always supports the complete logical `get/set/delete/list` contract. Simple stores can stop there. More capable -value stores can additionally expose `stat`, direct range `readFile`, `openReadStream`, selected direct materialized write modes, -and selected `writeStream` modes through `RecordStoreCapabilitiesType`. `createRecordAdapter()` translates only the declared -lanes into native adapter capabilities. +`adapter/sqlite` exports `createSqliteAdapter()`. -The portable complete-record fallback uses the versioned `RecordType` union and base64 file bytes so JSON, Web Storage, RxDB, -unstorage, and SQL text columns share one representation. A specialized store such as Deno KV can keep large body parts as raw -binary values and use metadata-only/range/stream/direct-write lanes so the generic base64 representation is not on its -large-file hot path. Deno KV declares materialized replace/append/update as direct store modes; append/update rebuild the next -immutable generation part-by-part instead of reconstructing the previous complete logical file. Its stream lane remains -replace-only, so streamed append/update can still require facade input buffering. +## Object clients and drivers -### Object storage +### S3 client -`@okikio/opfs/adapter/object` exports the provider-neutral object-storage layer: +`@okikio/opfs/s3` exports: -- `ObjectCapabilitiesSchema` / `ObjectCapabilitiesType` -- `ObjectStatType` and `ObjectEntryType` -- object GET, PUT, COPY, and LIST option types -- `ObjectStoreType` -- `ObjectAdapterOptionsType` -- `createObjectAdapter()` +- `S3AddressingSchema` / `S3AddressingType`; +- `S3CredentialsSchema` / `S3CredentialsType`; +- `S3CredentialSourceType`; +- `S3_LIMITS`; +- `S3ClientOptionsType`; +- `S3RequestOptionsType`; +- `S3CompleteOptionsType`; +- `S3UploadType`; +- `S3PartType`; +- `S3Error`; +- `S3ClientType`; +- `createS3Client()`. -`ObjectStoreType` preserves object concepts such as ETags, provider version IDs, user metadata, prefix listing, conditional -writes, and server-side copy. `createObjectAdapter()` then maps that model into files and directories. Empty directories use -trailing-slash marker objects, while ordinary prefix listing also recognizes directories created outside this library. +Important client optimization options: -The direct S3 API lives at `@okikio/opfs/s3`: +```text +delayedMultipart +signingKeyCache +``` + +The driver paths are: ```ts -import { createS3Client } from "@okikio/opfs/s3"; -import { createS3Adapter } from "@okikio/opfs/adapter/s3"; - -const client = createS3Client({ - endpoint, - bucket, - region, - credentials, -}); -const adapter = createS3Adapter(client); +createS3Driver(options); +createS3DriverFromClient(client); ``` -`S3ClientType` extends `ObjectStoreType` and also exposes signed `request()`, `createUpload()`, `uploadPart()`, -`completeUpload()`, and `abortUpload()`. The lower-level request method is the deliberate escape hatch for provider-specific -S3 features that do not belong in the filesystem API. - -The Azure counterpart lives at `@okikio/opfs/azure` and `@okikio/opfs/adapter/azure`: +The convenience adapter is: ```ts -import { createAzureClient } from "@okikio/opfs/azure"; -import { createAzureAdapter } from "@okikio/opfs/adapter/azure"; - -const client = createAzureClient({ endpoint, container, credential }); -const adapter = createAzureAdapter(client); +createS3Adapter(client, options?); ``` -`AzureClientType` retains the REST request escape hatch, provider request IDs, range access, block upload, and server-side copy. +For explicit layering, use `createObjectAdapter(createS3DriverFromClient(client))`. -The object-store interfaces are intentionally smaller than either provider protocol. See [S3 client protocol](./s3.md) and -[Azure Blob client protocol](./azure.md) for signing, version gates, multipart/block lifecycles, limits, provider failures, and -known unsupported operations. +### Azure Blob client -### Reverse key-value APIs +`@okikio/opfs/azure` exports: -`@okikio/opfs/driver/kv` exports `createKeyValueDriver()`. It maps colon-delimited keys onto private directories so both `foo` -and `foo:bar` can exist at the same time: +- `AZURE_STORAGE_VERSION`; +- `AzureStorageVersionSchema` / `AzureStorageVersionType`; +- `AzureCredentialType`; +- `AzureClientOptionsType`; +- `AzureRequestOptionsType`; +- `AzureClientType`; +- `AzureError`; +- `AZURE_LIMITS`; +- `createAzureClient()`. + +Important optimization options: ```text -foo -> /key-foo/value -foo:bar -> /key-foo/key-bar/value +blockUpload +serverCopy ``` -The driver supports string/raw get and set, existence, metadata, hierarchical key enumeration, clear, and explicit filesystem -ownership transfer. It also exposes `inspect()`, `plan()`, and `getMetrics()`. Those methods delegate to the backing -`FileSystemType`, so a reverse ecosystem consumer sees the same effective routes, limits, partition policy, buffer ceiling, and -observed metrics instead of receiving a second approximation of storage capability. +The driver paths are: -`@okikio/opfs/driver/unstorage` is a thin translation over this generic driver. It supplies unstorage method names and -`maxDepth` behavior while reusing the same collision-safe filesystem mapping. +```ts +createAzureDriver(options); +createAzureDriverFromClient(client); +``` -### Schemas +The convenience adapter is `createAzureAdapter(client, options?)`. -`@okikio/opfs/schema` exports the executable project data contracts: +## Bridge API -- `PathSchema` / `PathType` -- `AdapterNameSchema` / `AdapterNameType` -- `EntryKindSchema` / `EntryKindType` -- `OpfsContextSchema` / `OpfsContextType` -- `CoordinationModeSchema` / `CoordinationModeType` -- `WriteModeSchema` / `WriteModeType` -- `AdapterCapabilitiesSchema` / `AdapterCapabilitiesType` -- `ErrorCodeSchema` / `ErrorCodeType` -- `RecordVersionSchema` / `RecordVersionType` -- `DirectoryRecordSchema` / `DirectoryRecordType` -- `FileRecordSchema` / `FileRecordType` -- `RecordSchema` / `RecordType` -- `Db0DialectSchema` / `Db0DialectType` -- `SqlIdentifierSchema` / `SqlIdentifierType` +`@okikio/opfs/bridge/kv` exports: -The package exports Zod 4 schemas directly. Zod 4 implements Standard Schema, so a consumer that accepts that interface can use -the same schema value without a second OPFS-owned wrapper contract. +- `KeyValueMetaType`; +- `KeyValueBridgeOptionsType`; +- `KeyValueBridgeType`; +- `createKeyValueBridge()`. -### Bridge descriptors +The bridge includes `inspect()`, `plan()`, and `getMetrics()` so consumers can reason about the storage stack beneath +the KV projection. -`@okikio/opfs/bridge` groups the adapter and reverse-driver directions for an ecosystem. It does not replace either primitive. +`@okikio/opfs/bridge/unstorage` exports: -- `UnstorageBridge`: both directions. -- `RxDbBridge`: collection to OPFS only. -- `Db0Bridge`: database to OPFS only. -- `DrizzleBridge`: database/table to OPFS only. -- `KeyValueBridge`: OPFS to generic asynchronous key-value only. +- `UnstorageBridgeMetaType`; +- `UnstorageBridgeTransactionOptionsType`; +- `UnstorageBridgeType`; +- `UnstorageBridgeOptionsType`; +- `createUnstorageBridge()`. -`defineBridge()` validates that `directions.toOpfs/fromOpfs` agree with the constructors. Every unsupported direction must state a -reason. This is intentional for ecosystems where the reverse shape would require query, conflict, transaction, synchronous, or -other semantics a filesystem does not own. +## Integration API -Path utility API ----------------- +`@okikio/opfs/integration/definition` exports: -`@okikio/opfs/path` exposes the canonical virtual-path model used by adapters: +- `IntegrationDirectionSchema` / `IntegrationDirectionType`; +- `IntegrationDirectionsSchema` / `IntegrationDirectionsType`; +- `IntegrationType`; +- `defineIntegration()`. -- `ROOT_PATH`: the canonical `/` root. -- `normalizePath(path)`: resolves `.`, `..`, duplicate separators, and relative input while rejecting root escape, backslashes, and NUL. -- `splitPath(path)`: returns canonical path segments without `/`. -- `joinPath(...parts)`: joins inputs and returns a canonical `PathType`. -- `dirname(path)`: returns the canonical parent path. -- `basename(path)`: returns the final name. Root returns an empty string. -- `isAncestorPath(ancestor, path)`: tests strict ancestry after normalization. -- `validateName(name)`: validates one direct File System API child name. -- `PathType`: validated canonical virtual path type. +`@okikio/opfs/integration` exports current first-party direction definitions: -Use the high-level filesystem methods for normal application work. These helpers are primarily for adapters, drivers, and code that persists canonical paths. +```text +UnstorageIntegration +RxDbIntegration +Db0Integration +DrizzleIntegration +DrizzleIntegrationSourceType +``` -Error API ---------- +Direction metadata never makes an unsupported reverse contract executable. -The root module exports `FileSystemError`, `getErrorName()`, `getErrorMessage()`, and `toFileSystemError()`. +## Schema and path API -`FileSystemError` carries: +`@okikio/opfs/schema` owns project serializable schemas and their inferred types. Important groups include: ```text -code stable ErrorCodeType -operation filesystem operation that failed -path canonical path when one exists -cause original runtime/provider failure when retained +PathSchema / PathType +WriteModeSchema / WriteModeType +CoordinationModeSchema / CoordinationModeType +AdapterCapabilitiesSchema / AdapterCapabilitiesType +SupportModeSchema / SupportModeType +MetricsModeSchema / MetricsModeType +PartitionModeSchema / PartitionModeType +RecordSchema / RecordType +DriverKindSchema / DriverKindType +LimitKindSchema / LimitKindType +LimitSourceSchema / LimitSourceType +LimitUnitSchema / LimitUnitType +LimitSchema / LimitType +RequirementStateSchema / RequirementStateType +RequirementSchema / RequirementType +DriverOptimizationSchema / DriverOptimizationType ``` -`toFileSystemError()` maps known DOMException names and server error codes such as `ENOENT`, `EEXIST`, and quota/permission failures into the stable package categories. Unknown provider failures remain `unknown` and retain the original cause. +`@okikio/opfs/path` exposes canonical virtual-path helpers: -`getErrorName()` and `getErrorMessage()` are safe extraction helpers for diagnostics where the caught value is `unknown`. - -Browser capability APIs ------------------------ +```text +ROOT_PATH +normalizePath +splitPath +joinPath +dirname +basename +isAncestorPath +validateName +PathType +``` -### `probeOpfs()` +## Error API -Returns a non-throwing `OpfsCapabilitiesType` report with: +The root module exports: -- execution context; -- root availability and normalized root error; -- embedded/same-origin-top facts when observable; -- Web Locks availability; -- sync access exposure; -- storage estimate when available; -- persistence status when available. +```text +FileSystemError +getErrorName +getErrorMessage +toFileSystemError +``` -It does not report `isIncognito` or `isPrivate`. +`FileSystemError` carries stable code, operation, optional canonical path, and original cause. -### `getOpfsContext()` +## Browser capability API -Classifies the current browser execution context as window, dedicated worker, shared worker, service worker, generic worker, or unknown. +The root exports `probeOpfs()` and `getOpfsContext()`. -### iframe subpath +`probeOpfs()` returns a non-throwing current-realm report. It probes actual APIs/policy rather than maintaining a +browser-brand table. `@okikio/opfs/iframe` exports: -- `supportsUnpartitionedOpfsRequest()` -- `requestUnpartitionedFileSystem()` +```text +supportsUnpartitionedOpfsRequest +requestUnpartitionedFileSystem +``` -The request is explicit because browser permission/user-activation requirements must remain under application control. +The application remains responsible for the permission/user-activation flow. -Lifecycle ---------- +## Metrics API -`FileSystemType` implements `AsyncDisposable`. +`@okikio/opfs/metrics` exports logical facade metrics plus `DriverMetricsType`. -```ts -await fileSystem.close(); -``` +Logical facade metrics count operations, failures, logical bytes, native/emulated/partitioned routes, buffering, and +optional timing. + +Driver metrics are physical and provider-specific enough to include request/retry/part/physical-byte/cleanup information +without forcing that data into logical filesystem counters. + +## Lifecycle -or with supported explicit resource management syntax: +`FileSystemType` implements async disposal. Closing is idempotent. The adapter is disposed only when ownership was +transferred. Each adapter/driver/bridge has its own explicit ownership option for injected resources. ```ts await using fileSystem = createFileSystem(adapter, { @@ -628,4 +782,5 @@ await using fileSystem = createFileSystem(adapter, { }); ``` -Closing is idempotent. The adapter is closed only when ownership was explicitly transferred. +Cancellation and disposal remain distinct. A caller can cancel one operation without implicitly disposing a shared +storage resource. diff --git a/docs/azure.md b/docs/azure.md index 465cab5..3bf6d92 100644 --- a/docs/azure.md +++ b/docs/azure.md @@ -1,23 +1,22 @@ -Azure Blob client protocol guide -================================ +# Azure Blob client protocol guide -Purpose -------- +## Purpose -This document defines the Azure Blob Storage REST contract implemented by -`@okikio/opfs/azure`. It is intended for maintainers changing authentication, -service-version behavior, block upload, copy, conditional replacement, listing, -or Azurite interoperability. +This document defines the Azure Blob Storage REST contract implemented by `@okikio/opfs/azure`. It is intended for +maintainers changing authentication, service-version behavior, block upload, copy, conditional replacement, listing, or +Azurite interoperability. -The client uses the Blob REST API directly rather than wrapping the Azure SDK. -That keeps the dependency graph small, but it also means this repository owns -the protocol work it chooses to implement. The source and tests must therefore +The client uses the Blob REST API directly rather than wrapping the Azure SDK. That keeps the dependency graph small, +but it also means this repository owns the protocol work it chooses to implement. The source and tests must therefore make the exact REST contract explicit. ```text -ObjectStoreType / AzureClientType - | - v +AzureClientType + | + +--> direct protocol use + | + `--> Azure object driver -> object adapter -> FileSystemType + Blob REST request construction | +--> SAS query authorization @@ -32,27 +31,22 @@ Web Fetch `--> Azurite ``` -Web Crypto owns HMAC-SHA256 for Shared Key. `@std/encoding` owns Base64, -`@std/async/pool` owns bounded block concurrency, and `@std/xml` owns list and -block-list documents. +Web Crypto owns HMAC-SHA256 for Shared Key. `@std/encoding` owns Base64, `@std/async/pool` owns bounded block +concurrency, and `@std/xml` owns list and block-list documents. This guide uses these evidence classes: - - **Implemented** means current source contains the behavior. - - **Protocol** means current Microsoft REST documentation defines the behavior. - - **Emulator** means Azurite reproduces enough of the contract for local - integration tests but is not treated as complete Azure parity. +- **Implemented** means current source contains the behavior. +- **Protocol** means current Microsoft REST documentation defines the behavior. +- **Emulator** means Azurite reproduces enough of the contract for local integration tests but is not treated as + complete Azure parity. -The current implementation was reviewed against Microsoft Learn and current -Azurite documentation on August 14, 2026. +The current implementation was reviewed against Microsoft Learn and current Azurite documentation on August 14, 2026. +## The client targets one container -The client targets one container --------------------------------- - -`createAzureClient()` binds one endpoint and one container. Blob keys supplied -to `head()`, `get()`, `put()`, `delete()`, `copy()`, and `list()` are relative to -that container. +`createAzureClient()` binds one endpoint and one container. Blob keys supplied to `head()`, `get()`, `put()`, +`delete()`, `copy()`, and `list()` are relative to that container. A normal cloud endpoint looks like: @@ -78,36 +72,59 @@ which becomes: http://127.0.0.1:10000/devstoreaccount1/container/path/to/blob ``` -This difference matters to Shared Key canonicalization. Microsoft documents -that the emulator account segment appears once in the URL path and is prefixed -again by the signing account name. The implementation derives that duplicated +This difference matters to Shared Key canonicalization. Microsoft documents that the emulator account segment appears +once in the URL path and is prefixed again by the signing account name. The implementation derives that duplicated canonical-resource form from the URL rather than hard-coding an Azurite branch. -`AZURE_STORAGE_VERSION` defaults to `2026-04-06`. A caller can select another -service version when it needs the size or authentication behavior of an older -REST contract. +`AZURE_STORAGE_VERSION` defaults to `2026-04-06`. A caller can select another service version when it needs the size or +authentication behavior of an older REST contract. + +The driver remains a separate public layer: + +```ts +import { createAzureClient } from "@okikio/opfs/azure"; +import { createAzureDriverFromClient } from "@okikio/opfs/driver/azure"; +import { createObjectAdapter } from "@okikio/opfs/adapter/object"; + +const client = createAzureClient(options); +const driver = createAzureDriverFromClient(client); +const adapter = createObjectAdapter(driver); +``` + +The driver adds provider requirements, limits, generic optimization metadata, deterministic planning, and physical +request metrics. The adapter translates Blob keys/prefixes into the filesystem primitive contract. + +## Optimizations are independently controllable + +`blockUpload` : Defaults to true. Large/streamed complete replacements can stage blocks with bounded concurrency and +commit a block list. When disabled, the client does not advertise native stream-write/multipart behavior. Materialized +values above the single Put Blob ceiling fail rather than silently changing physical upload strategy. +`serverCopy` : Defaults to true when the selected authorization strategy can support the source and destination +semantics used by the client. When disabled, the client does not advertise provider-side copy and the object +adapter/facade can select an honest fallback. -Authorization is explicit -------------------------- +Both switches are exposed by the client and through `driver.inspect().optimizations`. A future cache, prefetch, +block-selection, or copy optimization that changes observable behavior must be inspectable and independently +disableable. + +## Authorization is explicit `AzureCredentialType` supports four strategies: -| Kind | Wire mechanism | Intended use | -| ---- | -------------- | ------------ | -| `sas` | SAS fields remain in the request query | Browser/server delegated access | -| `bearer` | `Authorization: Bearer ...` | Microsoft Entra access token | -| `shared-key` | Canonical Shared Key HMAC-SHA256 | Trusted server and Azurite | -| `headers` | Caller returns authorization headers | Provider/host integration not otherwise modeled | +| Kind | Wire mechanism | Intended use | +| ------------ | -------------------------------------- | ----------------------------------------------- | +| `sas` | SAS fields remain in the request query | Browser/server delegated access | +| `bearer` | `Authorization: Bearer ...` | Microsoft Entra access token | +| `shared-key` | Canonical Shared Key HMAC-SHA256 | Trusted server and Azurite | +| `headers` | Caller returns authorization headers | Provider/host integration not otherwise modeled | -The client does not read environment variables. Credentials are supplied by -the caller, and bearer tokens can be refresh functions resolved immediately -before the request. +The client does not read environment variables. Credentials are supplied by the caller, and bearer tokens can be refresh +functions resolved immediately before the request. -Shared Key credentials contain an account name and Base64 account key. The -account key is a root-level storage credential. It should not be embedded in an -untrusted browser bundle. Browser applications normally use a scoped SAS or a -Microsoft Entra flow with suitable permissions. +Shared Key credentials contain an account name and Base64 account key. The account key is a root-level storage +credential. It should not be embedded in an untrusted browser bundle. Browser applications normally use a scoped SAS or +a Microsoft Entra flow with suitable permissions. ### Shared Key string to sign @@ -130,30 +147,26 @@ CanonicalizedHeaders CanonicalizedResource ``` -The signer supports the augmented Blob Shared Key format from service version -`2009-09-19` onward. Earlier versions are rejected rather than being signed -with modern rules that only look plausible. +The signer supports the augmented Blob Shared Key format from service version `2009-09-19` onward. Earlier versions are +rejected rather than being signed with modern rules that only look plausible. Two canonicalization rules change with the selected service version: - - `2014-02-14` and earlier sign a zero byte `Content-Length` as the literal - `0`. Later versions contribute an empty line for the same header. - - Versions before `2016-05-31` omit empty `x-ms-*` headers from - `CanonicalizedHeaders`. Version `2016-05-31` and later retain them as - `name:\n`. +- `2014-02-14` and earlier sign a zero byte `Content-Length` as the literal `0`. Later versions contribute an empty line + for the same header. +- Versions before `2016-05-31` omit empty `x-ms-*` headers from `CanonicalizedHeaders`. Version `2016-05-31` and later + retain them as `name:\n`. Canonical `x-ms-*` headers that participate in the selected version are: -1. converted to lowercase names; -2. normalized by collapsing linear whitespace outside quoted strings while - preserving whitespace inside quoted strings; -3. sorted by code-unit order; -4. emitted as `name:value\n`. +1. converted to lowercase names; +2. normalized by collapsing linear whitespace outside quoted strings while preserving whitespace inside quoted strings; +3. sorted by code-unit order; +4. emitted as `name:value\n`. -The quoted-string rule is significant for metadata and other extension headers. -For example, `alpha beta` canonicalizes to `alpha beta`, while the two spaces -inside `alpha "beta gamma"` remain two spaces. Collapsing the quoted value -would sign different bytes from the value Azure receives. +The quoted-string rule is significant for metadata and other extension headers. For example, `alpha beta` +canonicalizes to `alpha beta`, while the two spaces inside `alpha "beta gamma"` remain two spaces. Collapsing the +quoted value would sign different bytes from the value Azure receives. The canonical resource starts with: @@ -161,8 +174,7 @@ The canonical resource starts with: /account-name/request-path ``` -then appends lowercase query names in sorted order. Repeated values are sorted -and joined with commas. +then appends lowercase query names in sorted order. Repeated values are sorted and joined with commas. The signature is: @@ -176,39 +188,33 @@ and the HTTP header is: Authorization: SharedKey account-name:signature ``` -The deterministic unit suite signs Azurite requests with the documented -`devstoreaccount1` key and a fixed timestamp. It freezes exact signatures on -both sides of the `2014-02-14` zero-length change, verifies the `2016-05-31` -empty-header change, and rejects Shared Key versions older than `2009-09-19`. -This makes the tests independent from the implementation clock and catches -service-version canonicalization drift. +The deterministic unit suite signs Azurite requests with the documented `devstoreaccount1` key and a fixed timestamp. It +freezes exact signatures on both sides of the `2014-02-14` zero-length change, verifies the `2016-05-31` empty-header +change, and rejects Shared Key versions older than `2009-09-19`. This makes the tests independent from the +implementation clock and catches service-version canonicalization drift. -A low-level streamed body using Shared Key must provide `content-length` because -the signer cannot know the stream length without consuming it. High-level -`put()` avoids this problem by splitting the stream into known-size block +A low-level streamed body using Shared Key must provide `content-length` because the signer cannot know the stream +length without consuming it. High-level `put()` avoids this problem by splitting the stream into known-size block requests. +## The service version controls write limits -The service version controls write limits ------------------------------------------ - -Azure Blob limits changed across REST service versions. The client resolves -the relevant limit from the selected version instead of assuming the newest -size everywhere. +Azure Blob limits changed across REST service versions. The client resolves the relevant limit from the selected version +instead of assuming the newest size everywhere. `AZURE_LIMITS` records the values used by planning: -| Operation/era | Client limit | -| ------------- | ------------ | -| Maximum committed blocks | 50,000 | -| Maximum uncommitted blocks | 100,000 | -| `Copy Blob From URL` synchronous copy | 256 MiB | -| Old `Put Block` | 4 MiB | -| 2016-05-31 through 2019-era `Put Block` | 100 MiB | -| Current `Put Block` | 4,000 MiB | -| Old `Put Blob` | 64 MiB | -| 2016-05-31 through 2019-era `Put Blob` | 256 MiB | -| Current `Put Blob` | 5,000 MiB | +| Operation/era | Client limit | +| --------------------------------------- | ------------ | +| Maximum committed blocks | 50,000 | +| Maximum uncommitted blocks | 100,000 | +| `Copy Blob From URL` synchronous copy | 256 MiB | +| Old `Put Block` | 4 MiB | +| 2016-05-31 through 2019-era `Put Block` | 100 MiB | +| Current `Put Block` | 4,000 MiB | +| Old `Put Blob` | 64 MiB | +| 2016-05-31 through 2019-era `Put Blob` | 256 MiB | +| Current `Put Blob` | 5,000 MiB | The implementation selects: @@ -224,27 +230,20 @@ Put Blob limit older -> 64 MiB ``` -`Put Block From URL` uses the version table published on the current REST page: -4,000 MiB from version `2020-04-08` onward and 100 MiB before that point. -Microsoft's same page currently contains a contradictory prose sentence that -still says the operation is limited to 100 MiB. The version table is also -consistent with the service's modern block-blob capacity model, so the client -follows the table. This contradiction is recorded as an upstream documentation -risk rather than hidden. The Docker suite uses small ranges and therefore does -not prove the 4,000 MiB ceiling; an opt-in real Azure test must protect that -limit before it is treated as independently verified. - -`blockSize` defaults to 8 MiB. The constructor rejects a configured block size -above the selected service-version limit. A known body can require a larger -block size to remain within 50,000 committed blocks; the planner chooses the -larger legal size and rejects an impossible request before starting the commit. +`Put Block From URL` uses the version table published on the current REST page: 4,000 MiB from version `2020-04-08` +onward and 100 MiB before that point. Microsoft's same page currently contains a contradictory prose sentence that still +says the operation is limited to 100 MiB. The version table is also consistent with the service's modern block-blob +capacity model, so the client follows the table. This contradiction is recorded as an upstream documentation risk rather +than hidden. The Docker suite uses small ranges and therefore does not prove the 4,000 MiB ceiling; an opt-in real Azure +test must protect that limit before it is treated as independently verified. +`blockSize` defaults to 8 MiB. The constructor rejects a configured block size above the selected service-version limit. +A known body can require a larger block size to remain within 50,000 committed blocks; the planner chooses the larger +legal size and rejects an impossible request before starting the commit. -High-level upload has two paths ------------------------------- +## High-level upload has two paths -A materialized `Uint8Array` at or below the selected `Put Blob` limit uses one -`Put Blob` request with: +A materialized `Uint8Array` at or below the selected `Put Blob` limit uses one `Put Blob` request with: ```text x-ms-blob-type: BlockBlob @@ -270,29 +269,24 @@ Put Block List Get Blob Properties ``` -Block IDs are deterministic Base64 values derived from zero-padded sequential -numbers. Every request in one upload therefore has a stable order and the -commit document can list exactly the intended blocks. +Block IDs are deterministic Base64 values derived from zero-padded sequential numbers. Every request in one upload +therefore has a stable order and the commit document can list exactly the intended blocks. -`Put Block` requests intentionally do not receive destination `If-Match` or -`If-None-Match`. Uncommitted blocks are not yet the authoritative destination -blob. The precondition and final metadata belong on `Put Block List`, which is -the operation that commits the new block blob. +`Put Block` requests intentionally do not receive destination `If-Match` or `If-None-Match`. Uncommitted blocks are not +yet the authoritative destination blob. The precondition and final metadata belong on `Put Block List`, which is the +operation that commits the new block blob. -The XML block list is generated through `@std/xml/stringify`, so block IDs are -serialized by a real XML implementation rather than hand-escaped text. +The XML block list is generated through `@std/xml/stringify`, so block IDs are serialized by a real XML implementation +rather than hand-escaped text. -When `ObjectPutOptionsType.size` is supplied, the final streamed byte count must -match. A mismatch rejects the operation before final commit. +When `ObjectPutOptionsType.size` is supplied, the final streamed byte count must match. A mismatch rejects the operation +before final commit. -Unlike S3 multipart uploads, Azure uncommitted blocks do not have a separate -abort REST operation. Failed uploads can leave uncommitted blocks until Azure -cleans them up according to service policy. Documentation and tests therefore -must not describe stream cancellation as an atomic remote rollback. +Unlike S3 multipart uploads, Azure uncommitted blocks do not have a separate abort REST operation. Failed uploads can +leave uncommitted blocks until Azure cleans them up according to service policy. Documentation and tests therefore must +not describe stream cancellation as an atomic remote rollback. - -Range reads use the Blob range contract ---------------------------------------- +## Range reads use the Blob range contract `get()` maps package range fields to: @@ -300,23 +294,19 @@ Range reads use the Blob range contract x-ms-range: bytes=start-end ``` -`at` is the zero-based first byte. `length` controls the inclusive final byte. -When no length is supplied, the range remains open-ended. - -The client reports `rangeRead: true` because Azure Blob Storage can satisfy the -range at the provider rather than materializing the complete object in the -library first. +`at` is the zero-based first byte. `length` controls the inclusive final byte. When no length is supplied, the range +remains open-ended. +The client reports `rangeRead: true` because Azure Blob Storage can satisfy the range at the provider rather than +materializing the complete object in the library first. -Server-side copy preserves the Azure model ------------------------------------------- +## Server-side copy preserves the Azure model -Azure has more than one URL-based copy primitive. The client selects between -them rather than presenting one fictitious universal copy call. +Azure has more than one URL-based copy primitive. The client selects between them rather than presenting one fictitious +universal copy call. -For a source up to 256 MiB, `copy()` uses synchronous `Copy Blob From URL`. -For a larger source, it performs ranged `Put Block From URL` operations and then -commits them with `Put Block List`. +For a source up to 256 MiB, `copy()` uses synchronous `Copy Blob From URL`. For a larger source, it performs ranged +`Put Block From URL` operations and then commits them with `Put Block List`. ```text Get Blob Properties(source) @@ -331,55 +321,45 @@ Get Blob Properties(source) Put Block List ``` -The large-copy block size is increased when required to remain at or below -50,000 committed blocks. It is also constrained by the selected REST version's -`Put Block From URL` range limit. +The large-copy block size is increased when required to remain at or below 50,000 committed blocks. It is also +constrained by the selected REST version's `Put Block From URL` range limit. -The copy feature is advertised only when the selected service version supports -URL-copy operations and the configured credential type lets the client derive a -source authorization strategy. +The copy feature is advertised only when the selected service version supports URL-copy operations and the configured +credential type lets the client derive a source authorization strategy. For same-client copies: - - SAS includes its authorization on the generated source URL; - - Shared Key can authorize the destination and a same-account source; - - bearer credentials can use `x-ms-copy-source-authorization` from service - version `2020-10-02` onward; - - custom header authorization is not assumed to work for the source, so the - portable `copy` capability is disabled. +- SAS includes its authorization on the generated source URL; +- Shared Key can authorize the destination and a same-account source; +- bearer credentials can use `x-ms-copy-source-authorization` from service version `2020-10-02` onward; +- custom header authorization is not assumed to work for the source, so the portable `copy` capability is disabled. -Cross-account copy has additional source-authorization requirements. A Shared -Key for the destination account cannot sign a different account's source. Use -a source SAS or a suitable bearer/source authorization design rather than +Cross-account copy has additional source-authorization requirements. A Shared Key for the destination account cannot +sign a different account's source. Use a source SAS or a suitable bearer/source authorization design rather than assuming one account key grants cross-account access. -Source conditions map to the `x-ms-source-*` condition family where the selected -operation supports them. Destination `If-Match` / `If-None-Match` apply to the -single synchronous copy or to the final block-list commit for multipart copy. +Source conditions map to the `x-ms-source-*` condition family where the selected operation supports them. Destination +`If-Match` / `If-None-Match` apply to the single synchronous copy or to the final block-list commit for multipart copy. - -Listing is container pagination, not a directory API ------------------------------------------------------ +## Listing is container pagination, not a directory API `list()` requests the container with `restype=container&comp=list` and maps: | Package field | Azure query field | | ------------- | ----------------- | -| `prefix` | `prefix` | -| `delimiter` | `delimiter` | -| `limit` | `maxresults` | -| `cursor` | `marker` | - -`Blob` entries become `ObjectEntryType`. `BlobPrefix` entries become child -prefixes. `NextMarker` is returned as the next cursor. +| `prefix` | `prefix` | +| `delimiter` | `delimiter` | +| `limit` | `maxresults` | +| `cursor` | `marker` | -The filesystem adapter interprets directory markers and provider prefixes. The -Azure client itself retains object/blob terminology because Azure has no native -filesystem directory in the Blob service contract used here. +`Blob` entries become `ObjectEntryType`. `BlobPrefix` entries become child prefixes. `NextMarker` is returned as the +next cursor. +The Azure object driver and filesystem adapter interpret directory markers and provider prefixes. The Azure client +itself retains object/blob terminology because Azure has no native filesystem directory in the Blob service contract +used here. -Errors retain Azure request evidence ------------------------------------- +## Errors retain Azure request evidence `AzureError` keeps: @@ -390,100 +370,81 @@ x-ms-request-id when present original Response ``` -Azure can return XML or provider-specific text. The error parser uses structured -XML when available and retains the response even when a field is missing. - -The client has a configurable transport retry policy built on `@std/async/retry`. Client options control retry count, exponential -delay, jitter, and an optional per-attempt timeout. The policy retries 408, 429, 5xx, and transport failures for replayable -requests. Authorization is rebuilt on every attempt, which matters for refreshable bearer/custom credentials and Shared Key -dates. Redirects are manual so authorization is not silently carried to another authority. +Azure can return XML or provider-specific text. The error parser uses structured XML when available and retains the +response even when a field is missing. -A one-shot `ReadableStream` receives one attempt. The low-level `request()` API also accepts `retry: false` because replayability -does not prove that a provider-specific operation is safe to repeat. `request: { retries: 0 }` disables automatic retry for the -client. Provider-specific `Retry-After` interpretation is not yet modeled. +The client has a configurable transport retry policy built on `@std/async/retry`. Client options control retry count, +exponential delay, jitter, and an optional per-attempt timeout. The policy retries 408, 429, 5xx, and transport failures +for replayable requests. Authorization is rebuilt on every attempt, which matters for refreshable bearer/custom +credentials and Shared Key dates. Redirects are manual so authorization is not silently carried to another authority. -`getMetrics()` returns request, retry, terminal-failure, response, and optional Fetch-duration counters. `metrics: "none"` is the -baseline benchmark setting; `basic` counts; `timing` adds monotonic duration. +A one-shot `ReadableStream` receives one attempt. The low-level `request()` API also accepts `retry: false` because +replayability does not prove that a provider-specific operation is safe to repeat. `request: { retries: 0 }` disables +automatic retry for the client. Provider-specific `Retry-After` interpretation is not yet modeled. -`AbortSignal` reaches every Fetch operation. Cancellation ends local admission and HTTP work where Fetch can abort it. It does -not guarantee that Azure failed to accept a request before the signal reached the network stack. +`getMetrics()` returns request, retry, terminal-failure, response, and optional Fetch-duration counters. +`metrics: "none"` is the baseline benchmark setting; `basic` counts; `timing` adds monotonic duration. +`AbortSignal` reaches every Fetch operation. Cancellation ends local admission and HTTP work where Fetch can abort it. +It does not guarantee that Azure failed to accept a request before the signal reached the network stack. -Azurite is an integration target, not the specification -------------------------------------------------------- +## Azurite is an integration target, not the specification -`tests/provider/fixture.ts` uses the official `@testcontainers/azurite` module -with the pinned Azurite image and the well-known development account. -Testcontainers chooses a free mapped host port, so the concrete endpoint changes -per run while the account identity remains `devstoreaccount1`. +`tests/provider/fixture.ts` uses the official `@testcontainers/azurite` module with the pinned Azurite image and the +well-known development account. Testcontainers chooses a free mapped host port, so the concrete endpoint changes per run +while the account identity remains `devstoreaccount1`. -The test uses Shared Key, creates a disposable logical Azure container through the client's -signed low-level request, then exercises PUT, HEAD, range GET, conditional -create, block upload, server-side copy, list, delete, and the object-store -filesystem adapter. +The test uses Shared Key, creates a disposable logical Azure container through the client's signed low-level request, +then exercises PUT, HEAD, range GET, conditional create, block upload, server-side copy, list, delete, the Azure driver, +and the object filesystem adapter. -Azurite is intentionally treated as an emulator. Its documentation states that -it provides best-effort Azure Storage compatibility and can differ from the -cloud service. A green Azurite suite therefore proves real HTTP/authentication +Azurite is intentionally treated as an emulator. Its documentation states that it provides best-effort Azure Storage +compatibility and can differ from the cloud service. A green Azurite suite therefore proves real HTTP/authentication interoperability, not complete conformance with every Azure version or feature. -The deterministic unit suite remains responsible for exact Shared Key string -construction, version-dependent limits, condition placement, and source bearer -version gates. - -Before release, an opt-in real Azure Blob test should run against a disposable -container with short-lived CI credentials when organizational secret policy -permits it. - +The deterministic unit suite remains responsible for exact Shared Key string construction, version-dependent limits, +condition placement, and source bearer version gates. -Known non-goals ---------------- +Before release, an opt-in real Azure Blob test should run against a disposable container with short-lived CI credentials +when organizational secret policy permits it. -The current direct client does not claim full Azure Storage coverage. Important -features outside this focused contract include: +## Known non-goals - - hierarchical namespace/Data Lake Gen2 filesystem semantics; - - append blobs and page blobs; - - leases as a first-class high-level API; - - snapshots/version-ID aware filesystem paths; - - customer-provided encryption-key convenience APIs; - - immutability policies and legal holds; - - blob index tags; - - asynchronous `Copy Blob` polling workflows; - - batch operations; - - account/container administration beyond low-level requests; - - adaptive provider throttling beyond the configured exponential retry policy, including provider-specific `Retry-After` scheduling; - - Microsoft Entra token acquisition itself. +The current direct client does not claim full Azure Storage coverage. Important features outside this focused contract +include: -`request()` can reach an unmodeled REST operation when a caller supplies the -correct method, query, headers, and body. A feature should become a typed public -operation only when its ownership, failure behavior, version gates, and tests -are explicit. +- hierarchical namespace/Data Lake Gen2 filesystem semantics; +- append blobs and page blobs; +- leases as a first-class high-level API; +- snapshots/version-ID aware filesystem paths; +- customer-provided encryption-key convenience APIs; +- immutability policies and legal holds; +- blob index tags; +- asynchronous `Copy Blob` polling workflows; +- batch operations; +- account/container administration beyond low-level requests; +- adaptive provider throttling beyond the configured exponential retry policy, including provider-specific `Retry-After` + scheduling; +- Microsoft Entra token acquisition itself. +`request()` can reach an unmodeled REST operation when a caller supplies the correct method, query, headers, and body. A +feature should become a typed public operation only when its ownership, failure behavior, version gates, and tests are +explicit. -Primary specification sources ------------------------------ +## Primary specification sources Review these Microsoft sources before changing protocol behavior: - - Shared Key authorization: - https://learn.microsoft.com/rest/api/storageservices/authorize-with-shared-key - - Versioning for Azure Storage services: - https://learn.microsoft.com/rest/api/storageservices/versioning-for-the-azure-storage-services - - Put Blob: - https://learn.microsoft.com/rest/api/storageservices/put-blob - - Put Block: - https://learn.microsoft.com/rest/api/storageservices/put-block - - Put Block List: - https://learn.microsoft.com/rest/api/storageservices/put-block-list - - Put Block From URL: - https://learn.microsoft.com/rest/api/storageservices/put-block-from-url - - Copy Blob From URL: - https://learn.microsoft.com/rest/api/storageservices/copy-blob-from-url - - List Blobs: - https://learn.microsoft.com/rest/api/storageservices/list-blobs - - Azurite: - https://learn.microsoft.com/azure/storage/common/storage-use-azurite - -The REST documentation is authoritative for Azure. Azurite source and behavior -are integration evidence for the emulator only. +- Shared Key authorization: https://learn.microsoft.com/rest/api/storageservices/authorize-with-shared-key +- Versioning for Azure Storage services: + https://learn.microsoft.com/rest/api/storageservices/versioning-for-the-azure-storage-services +- Put Blob: https://learn.microsoft.com/rest/api/storageservices/put-blob +- Put Block: https://learn.microsoft.com/rest/api/storageservices/put-block +- Put Block List: https://learn.microsoft.com/rest/api/storageservices/put-block-list +- Put Block From URL: https://learn.microsoft.com/rest/api/storageservices/put-block-from-url +- Copy Blob From URL: https://learn.microsoft.com/rest/api/storageservices/copy-blob-from-url +- List Blobs: https://learn.microsoft.com/rest/api/storageservices/list-blobs +- Azurite: https://learn.microsoft.com/azure/storage/common/storage-use-azurite + +The REST documentation is authoritative for Azure. Azurite source and behavior are integration evidence for the emulator +only. diff --git a/docs/design.md b/docs/design.md index 4cee4b1..0d47c9a 100644 --- a/docs/design.md +++ b/docs/design.md @@ -1,397 +1,747 @@ -Architecture and invariants -=========================== +# Architecture and invariants -The architecture starts from one rule: +## Purpose -> The filesystem facade owns filesystem semantics. An adapter owns the mechanics of one backend. +`@okikio/opfs` is a storage programming model with an OPFS-shaped filesystem frontend. The package supports storage +systems that have very different native contracts. Browser OPFS exposes file and directory handles. Node, Deno, and Bun +expose host file APIs. Deno KV and IndexedDB expose values and transactions. S3 and Azure Blob expose object protocols. +Drizzle, db0, RxDB, and unstorage sit above their own storage engines. -That rule matters because OPFS, Node files, Deno KV, SQLite, S3, and Azure Blob do not have the same native operations. A useful -portable library must make common application behavior consistent without hiding those differences from performance-sensitive -or correctness-sensitive code. +The architecture keeps those differences visible while giving applications one portable filesystem API where that API +can be implemented correctly. -The complete data path is: +The defining path is: ```text - application - | - +---------------+---------------+ - | | - v v - path API OPFS-shaped handles - readFile / writeFile / walk DirectoryHandle / FileHandle - | | - +---------------+---------------+ - | - v - FileSystemType - | - normalize paths / normalize failures / cancellation - parent creation / recursive operations / staging - file locks / tree locks / sync-file lock lifetime - stream selection / bounded materialization / ownership - | - v - AdapterType - +----------------+----------------+ - | | | - v v v - native filesystem record storage object storage - OPFS/Node/Deno KV/doc/SQL S3/Azure Blob - | | | - | RecordStoreType ObjectStoreType - | | | - v v v - native bytes versioned row object key +native API / ecosystem + | + v + client optional protocol client + | + v + driver backend-native persistence + | + v + adapter driver -> canonical filesystem primitives + | + v + FileSystemType portable OPFS-shaped behavior + | + v + bridge FileSystemType -> real ecosystem contract ``` -The frontend therefore has one filesystem contract, but the adapter capability record still tells the truth about how that -backend gets the work done. +`integration` definitions are separate metadata that describe which of the two directions exist. They are not executable +bridges. -The adapter contract stays deliberately small ---------------------------------------------- +## Layer ownership -Every adapter implements six primitives: `stat`, `readFile`, `writeFile`, `readDir`, `createDir`, and `remove`. Everything else -is an optional acceleration or stronger native lifecycle. +### Client -This avoids a common adapter failure mode where every backend reimplements recursive copy, walk, parent creation, handles, -locking, and error normalization separately. If those policies live in every adapter, semantics drift as soon as one backend gets -a bug fix that the others do not. +A client owns a wire protocol when the protocol is useful independently of the filesystem abstraction. -Optional operations exist only when the backend can perform them natively: +Current examples: ```text -streamRead -> openReadStream -write mode -> writeStream when mode is in streamWriteModes -nativeCopy -> copy -nativeMove -> move -positionalWrite -> openWritableFile -syncAccess -> openSyncFile +src/s3.ts S3 REST + SigV4 + multipart + request policy +src/azure.ts Azure Blob REST + authentication + block upload + request policy ``` -`streamWriteModes` is intentionally a list. A local file can stream replace, append, and update. An object store can usually -stream a complete replacement but cannot append bytes to an existing object in place. One `streamWrite: true` flag would hide -that difference and make callers reason from a capability that was too broad. +A client can expose protocol operations that do not belong in a filesystem. For example, an S3 client can retain ETags, +conditional requests, upload IDs, provider request IDs, presigned requests, object metadata, and multipart controls. + +Node, Deno, Bun, OPFS, IndexedDB, and localStorage do not need a package-owned protocol client. They begin at the driver +layer. -The adapter can additionally publish hard `limits` and a durable `partition` description. Those are facts about the configured -backend, not policy guesses. `FileSystemType.inspect()` combines them with resolved optimization controls and effective facade -support. `plan()` uses the same information before I/O, so runtime execution and preflight selection share one route model. +### Driver -Route-changing optimizations are facade policy: +A driver owns one configured backend's storage mechanics. It must remain useful without `FileSystemType`. + +A driver owns: + +- backend-native operations; +- required resources and current availability facts; +- provider hard limits; +- implementation safety limits; +- caller-selected policy limits; +- dynamic limits that still require a probe; +- driver-specific optimization switches; +- deterministic preflight planning; +- physical backend metrics when available; +- disposal of resources whose ownership was explicitly transferred. + +A driver does not own recursive filesystem semantics merely because the adapter above it needs them. + +The three reusable driver families are: ```text -streamRead -streamWrite -rangeRead -nativeCopy -nativeMove +FileDriverType + OPFS / Node / Deno / Bun + +RecordDriverType + memory / Deno KV / localStorage / IndexedDB / Cache + SQLite rows / db0 / Drizzle / RxDB / unstorage + +ObjectDriverType + S3 / Azure Blob / custom object storage ``` -Each defaults to enabled and can be disabled independently. This is deliberately different from capability detection. The adapter -should expose the strongest implementation it has; the caller can force the safe fallback for differential tests, observability, -provider workarounds, or policy. A disabled route is never relabelled native. +These families preserve stronger native concepts. They are not a forced lowest-common-denominator interface. + +### Adapter + +An adapter is deliberately smaller. It translates one driver into the filesystem primitive set consumed by +`FileSystemType`. -`nativeCopy` is also separate. A provider-side S3 or Azure copy can move terabytes without transferring the source through this -process, even though the same provider has no filesystem rename. The facade checks native copy before opening a source stream. -That ordering is a performance invariant, not an implementation detail: +Required adapter primitives: ```text -correct selection -filesystem.copy() - | - +-- native copy available -> adapter.copy() - | - `-- no native copy -------> open source stream -> transfer +stat +readFile +writeFile +readDir +createDir +remove +``` + +Optional direct routes: -incorrect selection -open source stream -> discover native copy -> source GET was already wasted +```text +openReadStream +writeStream +copy +move +openWritableFile +openSyncFile ``` -Paths are virtual identities, not host paths -------------------------------------------- +The adapter reports whether those direct routes are native. It does not say that a portable operation is unavailable +merely because the facade can emulate it. -Every adapter receives a canonical `PathType`: +This distinction is important: ```text -/ -/a -/a/b.txt +adapter.nativeMove = false +FileSystemType.move() can still exist as copy + remove ``` -The adapter seam rejects or never receives forms such as: +The first value describes the translation layer. The second describes the effective public route. + +### FileSystemType + +`FileSystemType` owns portable filesystem behavior: + +- canonical virtual paths; +- parent creation; +- OPFS-shaped file and directory handles; +- recursive walk, copy, move, remove, and empty-directory operations; +- staged writable-file semantics; +- synchronization and lock lifetime; +- bounded stream materialization when a backend cannot stream; +- normalized filesystem failures; +- logical metrics; +- adapter/facade optimization switches; +- composition of driver preflight with adapter and facade policy. + +The facade must not claim that an emulated route has the atomicity, consistency, or memory behavior of a native route. + +### Bridge + +A bridge starts from an existing `FileSystemType` and implements another ecosystem's real contract. + +Current bridges are: ```text -a/b -/a/ -/a//b -/a/./b -/a/../b -/a\b +bridge/kv hierarchical asynchronous key/value view +bridge/unstorage unstorage Driver-shaped view ``` -Public methods accept more convenient input and call `normalizePath()` first. This lets callers write ordinary path-like input -without making every adapter repeat normalization rules. +A bridge is valid only when the filesystem can satisfy the ecosystem contract. A filesystem cannot become a SQL engine +by renaming methods. A real RxDB reverse bridge would need to implement the complete `RxStorage` semantics, including +conflict, query, checkpoint, change-stream, cleanup, and lifecycle behavior. + +### Integration definition + +`integration/definition` stores import-safe direction metadata: + +```text +toOpfs ecosystem/native resource -> driver/adapter -> FileSystemType +fromOpfs FileSystemType -> ecosystem bridge +``` -The host filesystem adapters map this virtual namespace under one configured host root. A virtual path cannot escape that root -after host resolution. The virtual namespace does not expose symbolic-link identity, permission bits, or arbitrary host paths -as part of the portable contract. +An unsupported direction must state why it is unsupported. A definition never substitutes for the missing executable +contract. + +## Driver definitions and third-party extension + +`defineDriver()` is the smallest third-party extension seam. It validates structured driver metadata without registering +global state. + +```ts +import { defineDriver } from "@okikio/opfs/driver"; + +const driver = defineDriver({ + name: "example", + kind: "record", + provides: ["get", "set", "delete", "list"], + ownership: "borrowed", + requirements: [ + { code: "database", state: "available" }, + ], + limits: [ + { + code: "value-bytes", + kind: "hard", + source: "provider", + unit: "bytes", + value: 64 * 1024, + }, + ], + optimizations: [ + { + code: "partition", + enabled: true, + changesBehavior: true, + disableable: true, + }, + ], +}); +``` -Record stores and object stores need different translation layers ------------------------------------------------------------------ +`provides` records stable backend capability names for inspection. It is an open vocabulary so a provider-specific +driver can report operations beyond the three core driver families. `ownership` reports the long-lived backend resource +relationship: -A value store naturally answers "what value is stored at this key?" It does not naturally answer filesystem questions such as -"what are the direct children of this directory?" `RecordStoreType` supplies the reusable record translation for that family. +```text +none no disposable external backend resource is owned by this driver +borrowed the caller retains ownership of the injected backend resource +owned the driver owns the backend resource and can release it +``` -The complete persisted record has a canonical `path` plus a separate `parent`. Direct directory listing can therefore use an -index or prefix query over `parent` instead of scanning and parsing every path. File bytes are base64 so the fallback shape can -survive JSON, Web Storage, document databases, and SQL text. The extra storage and encoding work is accepted only for this -complete-record path. Native file and object adapters do not use that representation. +This report is separate from method presence. Typed file, record, and object driver contracts remain the operational +API. -The record contract also has optional byte lanes. A store can provide metadata-only stat, direct ranges, streaming reads, direct -materialized writes for selected modes, or streaming writes for selected modes. The generic record adapter advertises only the -lanes the store declares. Deno KV uses these lanes so a partitioned file is not reconstructed into one base64 record for stat, -listing, range reads, streaming reads, materialized append/update, or streamed replacement. Its append/update lane constructs a -new immutable generation one part at a time. This still performs provider I/O for untouched bytes, but it keeps JavaScript memory -bounded by the configured part/concurrency policy. Simpler record stores keep the small complete-record contract. +A concrete storage implementation should normally use `defineFileDriver()`, `defineRecordDriver()`, or +`defineObjectDriver()` so its operational contract is type-checked as well as its metadata. -An object store has a different strength: large objects, byte ranges, prefix listing, whole-object replacement, conditional -requests, and provider-side copy. `ObjectStoreType` preserves those concepts before `createObjectAdapter()` translates them into -filesystem operations. +No process-global driver or adapter registry is required. A package can export a definition and normal constructors. The +application chooses and composes them explicitly. -Files map directly to object keys. Empty directories need a marker object because a pure prefix does not exist until at least one -child exists: +## Requirements describe availability + +A requirement is structured data with a stable `code` and one state: ```text -/photos -> photos/ marker -/photos/a.jpg -> photos/a.jpg -/photos/2026/b.jpg -> photos/2026/b.jpg +available +missing +unknown ``` -The adapter also accepts implicit directories inferred from foreign prefixes. This matters when the bucket/container is not -created exclusively by this library. +A requirement can describe facts such as: + +```text +Deno KV database supplied +IndexedDB exposed in the current realm +S3 credentials resolved +browser OPFS root acquired +transaction capability supplied by a database integration +``` -An object namespace can contain both `mixed` and `mixed/child`. A real filesystem cannot. The filesystem view resolves an exact -`mixed` object as the file, because exact `stat()` already does that. Reads and writes follow the same rule. This creates one -stable interpretation for a foreign namespace instead of making `stat()` and `writeFile()` disagree. +A driver should not run hidden provider I/O from `inspect()` merely to turn every unknown into a known value. Dynamic +facts can remain unknown until the caller runs an explicit probe or performs the operation. -Streaming stays native only when the backend really streams ------------------------------------------------------------ +## Limits have provenance -`writeFile()` accepts strings, Blob, ArrayBuffer, typed-array views, ReadableStream, and AsyncIterable input. +A numeric limit is not meaningful unless the caller can tell where it came from. -When the selected adapter supports native streaming for the requested write mode, the facade forwards a byte stream directly. -When it does not, the facade collects the stream below `maxBufferedWriteBytes` and then calls the materialized adapter write. -Crossing the limit cancels the producer and returns `too-large`. +`LimitType` records: ```text -ReadableStream - | - +-- native mode supported ------> adapter.writeStream() - | - `-- no direct adapter stream lane - | - v - bounded collector - | | - | +-- over limit -> cancel producer -> too-large - v - Uint8Array - | - v - adapter.writeFile() +code +kind hard | policy | dynamic +source provider | implementation | user | probe +unit bytes | count | milliseconds | operations +value optional for a dynamic unknown ``` -This makes memory behavior visible. A simple record-backed adapter can accept streamed input through the public API while still -reporting an emulated stream route because the complete record is materialized before storage. A specialized record store can -report a partitioned stream lane when its own physical layout preserves backpressure. Deno KV does exactly that for -replacement streams when partitioning is enabled. +Examples: -Partitioning is not hidden as an optimization. It changes durable physical layout, so the adapter publishes `mode`, part size, -threshold, maximum parts, and layout identity. Deno KV exposes `never | auto | always`. Its parts are written under a new -generation and the manifest is committed last. A pre-manifest crash can leak unreachable parts but cannot publish a partial new -logical file. +```text +Deno KV serialized value ceiling + kind: hard + source: provider -Multipart and block-upload clients use `@std/async/pool` for bounded request admission. The surrounding client still owns the -provider lifecycle: S3 waits for already-started part requests before it sends AbortMultipartUpload, while Azure documents that -uncommitted blocks have no equivalent abort operation. The pool limits concurrent work; it does not become authority for remote -commit, cleanup, or the terminal provider failure. +Deno KV conservative inline decoded-body budget + kind: policy + source: implementation/user -HTTP retry policy is separate from body replayability. Direct clients rebuild authorization on every retry and use exponential -backoff with jitter, but a mechanically replayable request can still be semantically non-idempotent. S3 multipart initiation and -completion therefore disable automatic request retry. Uploaded parts use stable part numbers and can use the normal retry policy. -The low-level S3/Azure request APIs expose `retry: false` so a caller can make the same decision for provider-specific operations. -One-shot `ReadableStream` request bodies are never retried automatically. +maxParts chosen by the application + kind: policy + source: user -Object append and update are optimistic read-modify-write --------------------------------------------------------- +available browser storage quota + kind: dynamic + source: probe +``` -Object stores do not expose a portable in-place byte update. Append and update therefore use the current object as the starting -image, modify that image, and replace the object. +Missing limits mean unknown. They never mean unlimited. -Without a precondition, two writers can both read version A and then publish different replacements; the later one silently -loses the earlier write. When the object client advertises conditional writes, the adapter uses the current ETag as `If-Match`. -A concurrent change then fails the replacement instead of becoming silent data loss. +## Optimizations are inspectable policy + +Every optimization that can change observable behavior must be independently disableable. + +Observable behavior includes more than returned bytes. It includes: + +- request count; +- failure timing; +- storage layout; +- atomicity or visibility points; +- consistency/caching behavior; +- provider-side resource lifetime; +- retry/cancellation timing; +- memory use when the alternate route has different materialization behavior. + +`DriverOptimizationSchema` enforces the critical invariant: + +> `changesBehavior: true` requires `disableable: true`. + +Current examples include: ```text -writer A: HEAD ETag=A -> GET A ---------> PUT if-match A -> succeeds -writer B: HEAD ETag=A -> GET A -------------------------> PUT if-match A -> fails +S3 delayed multipart promotion +S3 derived signing-key cache +Azure block upload +Azure server-side copy +Deno KV partition layout ``` -If a provider claims conditional writes but does not return an ETag for an existing object, the adapter refuses append/update. -That is safer than publishing an unconditional write while the capability record says optimistic protection exists. +The facade has its own route switches for streaming, ranges, native copy, and native move. Driver and facade switches +remain separate because they control different layers. -The provider client can disable `conditionalWrite` when a compatible protocol implementation does not support the required -precondition. The library does not choose provider behavior from a provider-name table. +## Planning is deterministic -Copy and move preserve their real commit behavior -------------------------------------------------- +A driver planner accepts the concrete operation shape: -Native host filesystems use their copy and rename operations when available. Object stores use provider-side copy when the -client can do it. The facade removes/rejects the destination according to its own overwrite contract before invoking native copy, -so the adapter does not have to invent another overwrite policy. +```text +operation +canonical path +canonical destination when relevant +logical size when known +input bytes when different from final size +bytes vs stream source +write mode +range flag +``` -When no native copy exists, file bytes move through a stream when both ends support streaming or through bounded materialization -otherwise. +`plan()` performs no storage or network I/O. It returns: -A native move can be atomic or near-atomic according to the host/provider operation. The portable fallback is explicitly: +```text +supported +support native | partitioned | unsupported at the driver layer +partBytes +parts +problems[] +actions[] +``` + +Problems and actions are structured. Their human messages are not the identity used by application logic. + +The facade planner then adds adapter and filesystem facts. One final plan can therefore explain all of these at once: ```text -copy source -> destination - | - +-- copy failed -> source remains - | - `-- copy succeeded -> remove source +driver: Deno KV key exceeds a provider serialized-key ceiling +adapter: no native range route +filesystem: requested stream would exceed maxBufferedWriteBytes +``` + +The caller can distinguish each cause and choose a concrete action. + +## File drivers + +A file driver is closest to native OPFS semantics. Node, Deno, Bun, and browser OPFS can expose direct ranges, streams, +native copy/move, asynchronous positional files, or synchronous random access when the runtime supports them. + +The adapter above a file driver is intentionally close to delegation: + +```text +Node file APIs + | +Node file driver + | +file adapter + | +FileSystemType +``` + +The host-path mapper lives with drivers. It maps virtual `/` below one configured host directory and rejects escape from +that host root. + +## Record drivers + +Record storage naturally addresses complete values or documents instead of files. `RecordDriverType` therefore defines +logical record operations plus optional stronger byte lanes. + +Portable record methods: + +```text +get(path) +set(record) +delete(path) +list(parent) +``` + +Optional stronger methods: + +```text +stat(path) metadata without body reconstruction +readFile(path, range) direct byte/range access +openReadStream(path) direct streaming +writeFile(path, bytes) direct write modes +writeStream(path, stream) direct streaming write modes +``` + +The driver declares replacement semantics, binary support, and transaction availability separately. This lets a SQLite +or Deno KV driver preserve stronger behavior without pretending localStorage has it. + +The generic record format remains a portable fallback. It uses base64 file bodies because JSON/document/text-column +stores can all preserve that representation. A specialized driver is free to use native BLOB/byte storage internally and +expose the same logical record contract above it. + +## Object drivers + +Object storage preserves object semantics before the adapter translates them into files/directories. + +An object driver can retain: + +- ranged GET; +- conditional writes; +- validators/ETags; +- provider object versions; +- metadata; +- native/server-side copy; +- multipart/block upload; +- continuation tokens; +- provider request metrics. + +The filesystem adapter does not remove these concepts from the driver. It uses the subset required to provide canonical +filesystem primitives. + +## Partitioning belongs to drivers + +Partitioning changes physical storage layout, so it belongs at the backend driver layer. + +Examples: + +```text +Deno KV one logical file -> manifest + value parts +S3 one object upload -> multipart upload parts +Azure Blob one blob upload -> blocks + committed block list +SQL possible future file row -> part rows / BLOB segments +``` + +These systems have different visibility, cleanup, atomicity, and retry rules. A universal facade chunker would hide +those provider-specific guarantees. + +A partitioning strategy should describe: + +```text +whether it changes durable layout +its activation policy +part size +part count ceiling +visibility/commit point +cleanup behavior +streaming capability +memory behavior +whether callers can disable it +``` + +## Deno KV reference layout + +Deno KV demonstrates the full model. + +The provider documents serialized key/value limits. The driver also chooses smaller decoded-body budgets because a raw +byte count is not equal to serialized value size. + +The large-file layout uses an immutable generation and manifest-last publication: + +```text +old manifest -> old generation + +write new part 0 +write new part 1 +... +write new part N + | + v +write new manifest logical visibility point + | + v +remove old reachable generation ``` -That fallback is not atomic. A failure after copy and before remove can leave both entries. The API documents this instead of -claiming POSIX rename semantics on every backend. +If part writing fails, the new manifest is not published. The previous generation remains visible. The driver +best-effort removes parts from the failed generation. -Before copy or move, source and destination are checked for overlap. The library never removes an overwrite destination that is -an ancestor or descendant of the source. +A process crash before publication can still leave unreachable physical parts. That is storage leakage, not a partially +visible logical file. `DenoKvDriverType.collect()` exposes explicit, age-gated reclamation. The default one-hour grace +period avoids ordinary collection racing a recent unpublished generation, and `maxDeletes` bounds one maintenance pass. +Background deletion is not hidden inside ordinary reads or writes. -Coordination protects cooperating callers, not the whole storage system ------------------------------------------------------------------------- +The Deno KV planner also estimates physical tuple size from the concrete logical path. A file can be small enough to fit +by byte count while its physical key is too large. The planner reports that condition before provider I/O. -There are two classes of mutation. +## Filesystem path invariant -A file mutation acquires a shared tree lock plus an exclusive lock for that canonical file path: +Every adapter and driver path that participates in the filesystem seam is canonical: + +```text +/ +/a +/a/b.txt +``` + +The public facade can accept normalizable input, but `normalizePath()` runs before backend calls. Root escape, +backslashes, and NUL are rejected. + +The virtual path namespace is not an operating-system path namespace. Host file drivers map the canonical path below one +configured host root. + +## Streaming and memory invariant + +Large file size must not automatically become JavaScript heap size. + +A native streaming route is used only when the selected driver and adapter expose it and the corresponding optimization +is enabled. Otherwise the facade can materialize an input only up to `maxBufferedWriteBytes`. + +```text +stream + | + +-- native driver route --------------------> bounded backend streaming + | + `-- facade fallback -> bounded collector + | + +-- under limit -> materialized adapter write + `-- over limit -> cancel producer + too-large +``` + +The capability report distinguishes those routes. It does not label a buffered fallback as native streaming. + +## Writable-file invariant + +OPFS-shaped `createWritable()` stages a logical file image and commits on close. Abort discards the staged image. + +This is useful compatibility behavior, not the preferred large sequential write path. A caller that can use +`writeFile()` gives the facade a chance to select a true streaming adapter route. + +## Synchronous-file lifetime + +A synchronous file has two coupled resources: + +```text +facade path lock <------ same lifetime ------> driver sync file + | | + +---------------- close() ----------------+ +``` + +The path lock must remain held for the native file lifetime. `writeAll()` repeats partial writes until the complete +input is written or the backend reports no progress. + +## Coordination invariant + +The facade coordinates callers that use the same library lock namespace. + +File mutation: ```text shared tree lock | -exclusive /a/file lock +exclusive file-path lock | -write or sync-file lifetime +write / writable file / sync file lifetime ``` -A structural mutation such as recursive copy, move, remove, or empty-directory work acquires the exclusive tree lock: +Structural mutation: ```text exclusive tree lock | -structural mutation +copy / move / recursive remove / emptyDir ``` -Independent files can therefore make progress concurrently while a tree mutation cannot race an active library file mutation. -The in-realm lock implementation queues new readers behind an already-waiting writer so a busy read/write workload does not -starve structural work. +`local` coordination only spans one JavaScript realm. `web-locks` can coordinate cooperating browser realms that share +the lock namespace. `none` performs no library coordination. + +Database or distributed applications that require cross-process serialization must use the database/provider's real +transaction, lease, advisory-lock, or equivalent primitive. A local facade lock cannot provide that guarantee. -`coordination: "web-locks"` uses the browser Web Locks API. `auto` uses Web Locks when exposed and falls back to in-realm FIFO -coordination. `local` is one-realm coordination only. `none` preserves cancellation and adapter semantics but makes the caller -responsible for concurrency. +## Copy and move invariant -None of these modes becomes a distributed lock. Separate Node processes, browser profiles, hosts, or independent applications -need provider/database coordination when same-path atomicity matters across those processes. +Native copy and native move are separate capabilities. -Synchronous and asynchronous writable resources own locks for their complete lifetime --------------------------------------------------------------------------------------- +A native copy can avoid routing bytes through JavaScript. Object stores commonly provide this even when they cannot +provide rename semantics. -A synchronous file is not one short method call. It owns both the adapter file resource and the facade path lock until close: +When native move is absent: ```text -facade path lock <-------- same lifetime --------> adapter sync file - | | - +------------------- close() ------------------+ +source -> copy -> destination + | + `---------- remove source after successful copy ``` -This prevents an asynchronous write through the same facade from entering while synchronous random access is active. -`writeAll()` loops over partial native writes until the complete input is written or the backend reports no progress. +This fallback is not atomic. A failure after copy can leave both paths. Inspection and planning identify the route as +emulated. + +Source/destination overlap is checked before recursive structural work. An overwrite cannot delete an ancestor or +descendant that contains the source. -The OPFS-shaped `createWritable()` facade stages one file image and commits it on close. Abort discards the staged image. This is -useful for compatibility with File System API write commands, including seek and truncate. It is not the recommended path for -very large sequential files because the staged image is materialized. `FileSystemType.writeFile()` can use an adapter's native -streaming path instead. +## Database topology invariant -Integration direction is explicit ---------------------------------- +Two database directions must remain distinct. -An adapter is `ecosystem -> OPFS`. A driver is `OPFS -> ecosystem`. A bridge is only a descriptor that groups those directions; -it does not add a third translation layer to each operation. +Database-backed filesystem: ```text -ecosystem/client -> adapter -> FileSystemType -> driver -> ecosystem contract +Drizzle/db0/RxDB/SQLite database + | + record driver + | + record adapter + | + FileSystemType ``` -Some ecosystems are genuinely bidirectional. unstorage has a storage contract that can be consumed as a record backend and a -driver contract that can be implemented over `FileSystemType`. RxDB, db0, and Drizzle are not symmetric: their reverse -contracts require query, conflict, dialect, schema, or change-stream semantics a filesystem does not own. `defineBridge()` -therefore requires an explicit reason for unsupported directions instead of encouraging a false adapter. +SQLite database stored on OPFS: + +```text +application + | + Drizzle + | +SQLite engine + | +SQLite VFS + | +FileSystemType / native OPFS +``` -Cancellation and disposal are different operations --------------------------------------------------- +The current `driver/sqlite` and `adapter/sqlite` implement the first topology. They do not implement a SQLite VFS. A +future VFS must implement the SQLite engine's real file/VFS contract. -An `AbortSignal` asks active work to stop before a commit when possible. Closing a filesystem ends ownership of the facade. -Closing the facade does not dispose the adapter unless `disposeAdapter: true` transferred that ownership. +## Resource ownership -The same rule continues below the adapter: +Injected resources are borrowed by default. ```text -caller creates database/client/cache - | - +--> adapter borrows it - | | - | +--> filesystem closes - | `--> resource remains open - | - `--> caller still owns resource +caller creates database/client/filesystem + | + +--> driver/adapter/bridge borrows it + | + `--> caller remains owner ``` -An adapter option such as `disposeDatabase`, `disposeStore`, or another explicit ownership flag changes that lifecycle. The -option exists because connection pools, RxDB collections, unstorage instances, object clients, and caches are commonly shared by -more than one subsystem. +Ownership transfers only through an explicit option such as: -Errors normalize the portable category without erasing the provider cause -------------------------------------------------------------------------- +```text +disposeDatabase +disposeStorage +disposeDriver +disposeAdapter +disposeFileSystem +``` + +Disposal is idempotent at the owning layer where the public contract promises idempotency. A library must not close a +shared connection or filesystem merely because a facade closes. A driver only exposes backend disposal when its +construction options transferred ownership, so higher layers cannot accidentally dispose a borrowed database or storage +instance. + +Read-only policy is also a driver property for record backends. A read-only driver reports `write: false`, omits +optional write primitives, and rejects direct mutations before backend I/O. Adapters preserve that state rather than +inventing a second write policy that can disagree with the driver. + +## Cancellation invariant + +Cancellation asks active work to stop. Disposal releases owned resources. They are different operations. + +Long-running drivers check the signal before expensive work and between bounded chunks. When the facade aborts a stream +write, it cancels the producer when practical so upstream work does not continue after the file operation has become +terminal. + +Provider cleanup can need a separate bounded signal. For example, canceling an S3 multipart write must not use the +already aborted caller signal for the `AbortMultipartUpload` cleanup request. + +## Error invariant + +Backends fail with different error types. The facade normalizes known filesystem conditions to `FileSystemError` codes +while retaining the original cause. -Browsers use DOMException names. Node/Deno/Bun expose host error codes. Databases and cloud providers have their own errors. -`toFileSystemError()` maps known failures to stable package categories while retaining the original `cause`. +```text +DOMException / Node error code / provider error + | + v + FileSystemError + code + operation + path + cause +``` + +Protocol clients keep their own rich errors where provider-specific data matters. Translation into a filesystem error +happens at the storage/filesystem layer, not by deleting provider information at the client. + +## Metrics are layered + +Logical and physical work are not the same metric. + +`MetricsType` belongs to `FileSystemType` and records logical operations, logical bytes, route selection, facade +buffering, and optional facade timing. + +`DriverMetricsType` belongs to a driver and can record physical work such as: + +```text +provider requests +retries +responses/failures +logical payload bytes +physical bytes +parts/blocks/chunks +peak active provider work +backend duration +cleanup duration +``` -S3 and Azure clients also retain provider request identities on their own errors. Those IDs matter when a service returns an -unexpected result and the provider support logs are the only authoritative trace. +The benchmark matrix should compare each layer independently rather than attributing every cost to the facade. -Import safety follows the package graph ---------------------------------------- +## Import-safety invariant -The root package exports the portable facade, native browser OPFS convenience path, schemas, errors, handles, and capability -probes. It does not export every adapter from the root. +The root package is browser-safe. Runtime/provider code remains on explicit subpaths. ```text -@okikio/opfs browser-safe core + native OPFS -@okikio/opfs/adapter/node node:fs imports -@okikio/opfs/adapter/deno Deno runtime APIs -@okikio/opfs/adapter/bun Bun + Node-compatible APIs -@okikio/opfs/s3 Web Fetch/Crypto S3 client -@okikio/opfs/azure Web Fetch Azure Blob client -@okikio/opfs/adapter/drizzle optional drizzle-orm peer +@okikio/opfs +@okikio/opfs/driver/node +@okikio/opfs/driver/deno +@okikio/opfs/driver/bun +@okikio/opfs/driver/s3 +@okikio/opfs/driver/azure +@okikio/opfs/adapter/* +@okikio/opfs/bridge/* ``` -Importing a module does not read environment variables, configure logs, connect to providers, start workers, or mutate a global -adapter registry. +Importing a module does not connect to storage, read environment variables, start a worker, configure global logging, or +mutate a process registry. -Schemas are executable contracts, not duplicated type declarations -------------------------------------------------------------------- +## Review rules -Project-owned structural values use Zod schemas and inferred TypeScript output types. Public schema constants end in `Schema`. -Serializable project-owned types normally end in `Type`. +A storage change is not complete until these questions have concrete answers: -Zod 4 implements Standard Schema. The exported Zod value is therefore also the Standard Schema value. Creating a second OPFS -schema wrapper would add maintenance without adding a stronger contract. +1. Which layer owns the behavior? +2. Is the provider/native contract preserved below the adapter? +3. Are provider, implementation, user, and dynamic limits distinguishable? +4. Can an observable optimization be disabled? +5. Does planning use the actual path/size/source shape needed to detect known limits? +6. Is growing work bounded by bytes, parts, concurrency, retries, or time? +7. Who owns cancellation and who owns disposal? +8. Does an emulated route state its weaker atomicity, consistency, or memory behavior? +9. Does a reverse bridge implement the ecosystem's real contract? +10. Do tests and benchmarks exercise the layer being claimed rather than bypassing it? diff --git a/docs/ecosystems.md b/docs/ecosystems.md index 73fcfde..d3f95a3 100644 --- a/docs/ecosystems.md +++ b/docs/ecosystems.md @@ -1,158 +1,189 @@ -Ecosystem integrations -====================== +# Ecosystem integrations -`@okikio/opfs` integrates with another storage ecosystem at the highest stable abstraction the application already owns. It does -not duplicate every provider driver from that ecosystem. +## Purpose -That rule gives two complementary directions: +Storage ecosystems can connect to `@okikio/opfs` in two different directions. The package names those directions +explicitly and does not force symmetry where the upstream contract cannot be implemented correctly. ```text -existing storage resource existing OPFS filesystem - | | - v v - adapter reverse driver - | | - v v - FileSystemType KV / unstorage API +ecosystem/native resource FileSystemType + | | + driver bridge + | | + adapter v + | ecosystem API + v + FileSystemType ``` -The forward path lets filesystem-shaped application code use another storage system. The reverse path lets another ecosystem -consume any backend already reachable through `FileSystemType`. +`integration` definitions describe the two directions. They are metadata, not bridges. -Bridge descriptors make asymmetry part of the contract ------------------------------------------------- +## Direction inventory -`@okikio/opfs/bridge` groups the existing forward adapter and reverse driver for an ecosystem. It does not force every -integration to be symmetric. +| Integration | ecosystem -> OPFS | OPFS -> ecosystem | Reverse status | +| ----------- | ----------------- | ----------------- | -------------------------------------------------------------------------- | +| unstorage | yes | yes | `bridge/unstorage` implements a real Driver-shaped contract | +| RxDB | yes | no | a complete reverse path must implement `RxStorage` | +| db0 | yes | no | a filesystem is not a SQL database/query engine | +| Drizzle | yes | no | a filesystem is not a Drizzle dialect/schema/query engine | +| generic KV | n/a | yes | `bridge/kv` exposes the small contract the filesystem can actually satisfy | -| Bridge | ecosystem -> OPFS | OPFS -> ecosystem | Why the reverse side is absent when unsupported | -| --- | --- | --- | --- | -| `UnstorageBridge` | yes | yes | both stable shapes exist | -| `RxDbBridge` | yes | no | `RxStorage` also owns queries, conflicts, change streams, cleanup, and storage-instance semantics | -| `Db0Bridge` | yes | no | a filesystem is not a SQL query/dialect engine | -| `DrizzleBridge` | yes | no | a filesystem does not own Drizzle schema, dialect, or query-builder behavior | -| `KeyValueBridge` | no | yes | the generic reverse KV shape does not define persistence semantics needed to build an adapter | +`defineIntegration()` validates that declared support and constructors agree. An unsupported direction must have a +reason. -`defineBridge()` validates that direction declarations agree with real constructors. An unsupported direction must include a -reason. Third-party integrations can therefore publish capability honestly without inventing a method that only works for a -small subset of the upstream contract. +## unstorage -Unstorage works in both directions without a provider explosion ---------------------------------------------------------------- +### unstorage as the backend -The forward adapter accepts the high-level unstorage `Storage` contract: +The forward direction accepts unstorage's high-level `Storage` object: -```ts -import { createStorage } from "unstorage"; -import memoryDriver from "unstorage/drivers/memory"; -import { createFileSystem } from "@okikio/opfs"; -import { createUnstorageAdapter } from "@okikio/opfs/adapter/unstorage"; - -const storage = createStorage({ driver: memoryDriver() }); -const fileSystem = createFileSystem(createUnstorageAdapter(storage)); +```text +unstorage Storage + | +createUnstorageDriver + | +createRecordAdapter + | +FileSystemType ``` -Unstorage remains responsible for its selected driver, mounts, provider SDKs, retry behavior, and provider-specific limits. The -OPFS adapter only uses the stable high-level operations it needs to persist records. +This means the package does not reimplement unstorage's provider catalogue. Memory, filesystem, Redis, S3, Azure, +Cloudflare, Deno KV, IndexedDB, db0, or another upstream driver remains unstorage's responsibility. + +The OPFS record driver uses a private prefix and reversible key encoding. Read-only storage can be declared read-only at +the adapter layer. -This is intentionally broader than maintaining separate OPFS adapters for every unstorage driver. Current unstorage drivers span -browser storage, Cloudflare, Azure, S3, Deno KV, filesystem, Redis, databases, blobs, HTTP, and other providers. Duplicating that -catalog here would create a second compatibility matrix that would drift from upstream. +### FileSystemType as an unstorage Driver -The reverse direction is more powerful after the generic key-value driver: +The reverse direction is a real bridge: ```text -unstorage Storage - | - v -@okikio/opfs unstorage Driver - | - v -KeyValueDriverType - | - v +unstorage + | +createUnstorageBridge(fileSystem) + | FileSystemType - | - +-- native OPFS - +-- Node / Deno / Bun - +-- S3 / Azure Blob - +-- localStorage / IndexedDB / Cache - +-- Deno KV / SQLite - +-- RxDB / db0 / Drizzle / unstorage - `-- custom adapter ``` -`createKeyValueDriver()` owns the collision-safe filesystem mapping. `createUnstorageDriver()` only translates that contract to -unstorage's driver method names and `maxDepth` flag. Both reverse views retain `inspect()`, `plan()`, and `getMetrics()` from the -backing filesystem. An ecosystem caller can therefore reject a value above `maxFileBytes`, see when a streamed write would -buffer or partition, disable a native route on the filesystem, and observe the same counters without a second capability table. +The bridge supports values, raw bytes, metadata, keys, clear, and disposal behavior used by the stable unstorage Driver +surface implemented here. -The extra key directory is required because a KV store can contain both `foo` and `foo:bar`: +A private directory+leaf layout is necessary because unstorage can contain both: ```text -foo -> /key-foo/value -foo:bar -> /key-foo/key-bar/value +foo +foo:bar ``` -A naive `:` to `/` conversion would try to make `/foo` both a file and a directory. The private `value` leaf removes that -conflict while reversible segment encoding keeps `%`, `~`, spaces, slashes inside a key segment, and other URI-sensitive text -distinct. +A normal filesystem cannot make `/foo` both a file and a directory, so the bridge stores each exact value in a private +`value` leaf below its encoded hierarchy directory. -RxDB stays above RxStorage --------------------------- +Literal percent/tilde and separator-like characters are encoded reversibly so logical keys cannot collide. -RxDB already defines `RxStorage` as its storage-engine abstraction. An RxCollection adds document behavior, indexes, conflict -handling, and the selected RxStorage implementation. +## RxDB -`createRxDbAdapter()` therefore accepts an existing collection instead of implementing another RxStorage engine. The exported -`RxDbRecordJsonSchema` uses canonical `path` as the primary key and indexes `parent` for direct-child listing. +The forward driver targets an injected `RxCollection`: ```text -RxDB application +chosen RxStorage | - v -RxCollection + RxCollection | - +--> selected RxStorage + createRxDbDriver | - v -createRxDbAdapter() + record adapter | - v + FileSystemType +``` + +RxDB remains responsible for the selected `RxStorage`, document revision/conflict behavior, wrappers, replication, +multi-instance coordination, and licensing. + +`RxDbRecordJsonSchema` defines the collection shape expected by this integration: + +```text +path primary key +parent indexed direct-parent path +name +kind +data +size +lastModified +mediaType +``` + +The path fields have an explicit length ceiling because RxDB indexed string fields need a declared maximum length in the +supported schema shape. The driver rejects an oversized path before it asks the collection to query or write it. + +### Why there is no reverse RxDB bridge yet + +RxDB's storage-engine contract is `RxStorage`, not a key/value object. A real implementation must own semantics such as: + +- bulk writes with per-document conflict results; +- prepared Mango queries and counts; +- attachments; +- changed-document checkpoints; +- change streams; +- cleanup of deleted documents; +- multi-instance behavior where applicable; +- close and remove lifecycle. + +`FileSystemType` does not provide those semantics automatically. A future `bridge/rxdb` should therefore be a +substantial RxStorage implementation with its own tests and performance model, not a wrapper that renames filesystem +methods. + +## db0 + +The db0 direction begins with an already connected `Database`: + +```text +db0 connector + | + Database + dialect + | +createDb0Driver + | +record adapter + | FileSystemType ``` -This keeps the adapter compatible with the collection regardless of whether the application selected memory, IndexedDB, OPFS, -filesystem, SQLite, remote, worker, or another RxStorage family. RxDB keeps ownership of replication, multi-instance behavior, -licensing, storage wrappers, and conflicts. +The current dialect branches are: -db0 and direct SQLite share the same SQL record model ------------------------------------------------------- +```text +sqlite +libsql +postgresql +mysql +``` -`createDb0Adapter()` targets db0's high-level `Database` contract and its reported dialect. The SQL generation currently covers -SQLite, libSQL, PostgreSQL, and MySQL branches. It depends on database behavior rather than connector names, so a new db0 -connector does not need a new OPFS adapter when it presents the same database contract. +The driver generates the table/CRUD SQL required for the selected dialect. It targets db0's portable database/statement +contract rather than connector names. -`createSqliteAdapter()` is the focused direct SQLite path for applications that already own a small connected statement API. It -reuses the SQLite branch of the same SQL record contract instead of maintaining a second table layout and upsert implementation. +The path primary key uses a SHA-256 identity while the original path is retained separately. That avoids assuming +arbitrary long text is a portable primary-key type across the supported SQL families. -The default SQL table is `opfs_entries`. The path identity and parent path are stored separately so direct-child listing can use -a provider-appropriate index. The db0 path uses portable parameter placeholders and lets the db0 connector translate them where -its dialect requires a different native parameter shape. +The table is initialized only when requested. The injected database remains caller-owned unless `disposeDatabase` is +true. -Drizzle keeps schema ownership with the application ---------------------------------------------------- +There is no reverse db0 bridge because a filesystem cannot implement arbitrary SQL parsing, query planning, +transactions, dialects, schema metadata, or connector behavior. -Drizzle spans several SQL dialects and runtime drivers. A universal OPFS-owned Drizzle table would either choose one dialect or -hide dialect-specific DDL details. +## Drizzle -`createDrizzleAdapter()` therefore receives: +Drizzle is also a forward record driver, but the schema stays caller-owned: -1. an already-connected Drizzle database; -2. a table built for that database dialect; -3. the required logical columns. +```text +Drizzle database + table + | + createDrizzleDriver + | + record adapter + | + FileSystemType +``` -Required logical fields are: +The caller provides columns for: ```text path @@ -165,102 +196,123 @@ lastModified mediaType ``` -`path` must be unique or primary. `size` and `lastModified` must round-trip JavaScript safe integers. The bridge uses the common -select/insert/delete builder shape and keeps Drizzle an optional peer dependency. +`path` must be unique or a primary key. `size` and `lastModified` must round-trip JavaScript safe integers. -Inside one `FileSystemType`, normal coordination serializes same-path mutations. Separate processes or hosts are not serialized -by an in-memory facade lock. A database-backed deployment that needs cross-process replacement atomicity must use transactions, -leases, advisory locks, or another mechanism provided by its actual database/driver. +The generic driver uses Drizzle's common CRUD surface. Replacement is delete then insert, so the driver reports +best-effort replacement. This route is serialized inside one cooperating `FileSystemType`, but it is not an atomic +cross-process database replacement. -S3-compatible providers are configured by capability, not brand guesses ------------------------------------------------------------------------- +A database-specific Drizzle driver can expose stronger upsert, transaction, binary, and partition behavior without +changing the portable generic integration. -The direct S3 client is meant to work with AWS S3 and compatible XML/SigV4 services, but "S3-compatible" is not a promise that -all operations, preconditions, limits, checksums, or control-plane features are identical. +### Drizzle-backed filesystem versus Drizzle over OPFS-backed SQLite -The client options deliberately separate the parameters that compatible services vary: +These architectures are opposite directions. + +**A. Database-backed filesystem** is implemented now: ```text -endpoint -bucket -region -addressing: path | virtual -headers -copy: boolean -conditionalWrite: boolean -partSize -copyPartSize -concurrency -credentials +Drizzle + | +database rows + | +record driver + | +FileSystemType ``` -The safe rule is to read the selected provider's current primary documentation and enable only the capabilities it actually -implements for the operations used by the filesystem. - -A few current examples show why this matters: - -| Provider family | Current nuance that affects this client | -| --- | --- | -| AWS S3 | baseline SigV4, multipart upload, CopyObject, UploadPartCopy, conditional completion | -| Cloudflare R2 | S3-compatible endpoint with its own supported-operation set; `auto` is a documented region value | -| DigitalOcean Spaces | implements a documented subset of the S3 API and its own published object/multipart limits | -| Google Cloud Storage XML API | S3-compatible multipart exists, but documented multipart precondition behavior differs | -| Backblaze B2 S3 API | S3-compatible surface with its own unsupported/changed AWS control-plane features | - -For a Google Cloud Storage XML multipart path that does not support the preconditions expected by optimistic object -read-modify-write, create the client with `conditionalWrite: false`. That does not make append/update magically atomic; it makes -the absence of that safety property explicit. - -Cloudflare R2 and other services can also disable `copy` if their selected endpoint/path does not provide the server-side copy -contract expected by the adapter. The filesystem then falls back to the normal streamed/materialized copy path instead of -calling a native capability that was never real. - -The S3 request escape hatch is intentional ------------------------------------------ - -A filesystem does not need to model every S3 object or bucket feature. The direct client therefore exposes signed -`request(options)` in addition to `ObjectStoreType`. - -Use the filesystem/object layer for portable file behavior. Use the lower-level request API when the application needs a -provider-specific control such as an object-lock header, tag operation, checksum policy, versioning call, or another S3 operation -whose semantics should not be flattened into a generic filesystem method. - -The same principle applies to Azure Blob ----------------------------------------- - -Azure Blob has enough differences that the package implements a native Azure REST client rather than translating Azure through -an S3 compatibility layer. - -`createAzureClient()` supports SAS, Microsoft Entra bearer tokens, Shared Key, and custom-header credentials. Its service version is explicit. Its streamed -writes use Azure block APIs, and its large server-side copies use Put Block From URL when synchronous Copy Blob From URL is too -small. - -The object adapter above Azure is still the same `createObjectAdapter()` used by S3. The provider client owns Azure-specific -HTTP mechanics; the object adapter owns the file/directory translation. - -Choose the integration that matches the resource you already own ------------------------------------------------------------------ - -| Existing application resource | Preferred integration | -| --- | --- | -| browser OPFS root | `createOpfsAdapter()` or `openFileSystem()` | -| host directory | Node, Deno, or Bun adapter | -| unstorage `Storage` | `createUnstorageAdapter()` | -| RxDB collection | `createRxDbAdapter()` | -| db0 `Database` | `createDb0Adapter()` | -| connected SQLite statement API | `createSqliteAdapter()` | -| Drizzle database + table | `createDrizzleAdapter()` | -| Deno KV database | `createDenoKvAdapter()` | -| localStorage/sessionStorage-like Web Storage | `createLocalStorageAdapter()` | -| IndexedDB | `createIndexedDbAdapter()` / `openIndexedDbAdapter()` | -| Cache Storage `Cache` | `createCacheAdapter()` | -| S3-compatible endpoint | `createS3Client()` + `createS3Adapter()` | -| Azure Blob container | `createAzureClient()` + `createAzureAdapter()` | -| custom KV/document layer | `createRecordAdapter()` | -| custom object storage | `createObjectAdapter()` | -| any `FileSystemType` needed as KV | `createKeyValueDriver()` | -| any `FileSystemType` needed by unstorage | `createUnstorageDriver()` | - -Adding an extra ecosystem layer only because this package already has an adapter for it usually makes the system harder to -reason about. Use the direct adapter/client/driver for the abstraction the application already owns, and use bridge descriptors -when code needs to inspect both directions as one integration. +**B. Drizzle over an OPFS-backed SQLite database** requires SQLite below Drizzle: + +```text +application + | +Drizzle ORM + | +SQLite engine + | +SQLite VFS + | +FileSystemType / native OPFS +``` + +The current `driver/sqlite` does not implement a VFS. It stores OPFS logical records inside SQLite rows. + +A future SQLite VFS should be designed against the SQLite engine's actual VFS/file contract. It can then use +`FileSystemType` where that contract can be mapped correctly, including sync-access and locking requirements. Drizzle +can sit above the SQLite engine normally. + +## SQLite + +The current SQLite driver is intentionally direct and small. It consumes an injected connected SQLite database that can +prepare and run statements. + +Use it when the application already has a SQLite database and wants to store a virtual filesystem in rows. + +Do not use it as evidence that arbitrary SQLite WASM engines can already store their database file on this package. That +second capability is future VFS work. + +## Generic key/value bridge + +`createKeyValueBridge(fileSystem)` is a reverse bridge with a deliberately small contract: + +```text +has +get / set +getRaw / setRaw +remove +meta +keys +clear +inspect +plan +getMetrics +dispose +``` + +It is useful for ecosystem adapters that need hierarchical string/raw values but do not need SQL, document queries, or +conflict semantics. + +It exposes the backing filesystem's inspection and plan results rather than inventing a separate capability system. + +## Direction metadata + +`integration/definition` exists so applications and third-party packages can reason about asymmetry without executing +storage constructors. + +```ts +import { defineIntegration } from "@okikio/opfs/integration/definition"; + +const integration = defineIntegration({ + name: "example", + directions: { + toOpfs: { supported: true }, + fromOpfs: { + supported: false, + reason: "The upstream reverse contract requires query semantics.", + }, + }, + toOpfs(source) { + return createExampleAdapter(source); + }, +}); +``` + +There is no global registration. Applications import the integration definitions they want to use. + +## Choosing an integration + +Start from the resource the application already owns. + +| Existing resource | Preferred path | +| --------------------------------------------- | ----------------------------------------- | +| unstorage `Storage` | `driver/unstorage` -> `adapter/unstorage` | +| RxDB `RxCollection` | `driver/rxdb` -> `adapter/rxdb` | +| db0 `Database` | `driver/db0` -> `adapter/db0` | +| Drizzle database + table | `driver/drizzle` -> `adapter/drizzle` | +| connected SQLite database | `driver/sqlite` -> `adapter/sqlite` | +| custom value/document storage | `driver/record` -> `adapter/record` | +| existing `FileSystemType` needed as KV | `bridge/kv` | +| existing `FileSystemType` needed by unstorage | `bridge/unstorage` | + +Do not wrap a resource through an unrelated ecosystem only to reach OPFS. Every additional abstraction adds semantics, +requirements, failure behavior, and measurable overhead. diff --git a/docs/environments.md b/docs/environments.md index 836078e..d78e71c 100644 --- a/docs/environments.md +++ b/docs/environments.md @@ -1,178 +1,193 @@ -Execution environments -====================== +# Execution environments -`@okikio/opfs` keeps the filesystem frontend separate from backend availability. The same core source can compile for Window, -workers, Deno, Bun, and Node while runtime-specific adapters remain on explicit subpaths. +## Purpose -The package does not maintain a browser-brand or runtime-brand behavior table. It asks the current realm or adapter what it can -actually do and preserves the resulting capability/failure information. +`@okikio/opfs` keeps the portable filesystem frontend separate from the runtime/backend driver. The root module remains +safe to import in Window, workers, Deno, Bun, and Node. Runtime-specific code stays on explicit driver/adapter subpaths. -Browser OPFS follows the storage key of the current realm ---------------------------------------------------------- +The package does not select behavior from a runtime-brand table. It probes actual browser capabilities and reads +configured driver capabilities/requirements. -Native browser OPFS requires `navigator.storage.getDirectory()`. +## Browser OPFS -In Window, use the asynchronous facade: +Native browser OPFS begins with `navigator.storage.getDirectory()`. + +The root convenience path is: ```ts -const fileSystem = await openFileSystem(); -await fileSystem.writeFile("/state.json", "{}", { parents: true }); -``` +import { openFileSystem } from "@okikio/opfs"; -Do not assume synchronous access from Window. `openSyncFile()` succeeds only when the actual native file handle exposes the sync -handle API and the adapter reports that capability. +await using fileSystem = await openFileSystem(); +``` -DedicatedWorker is the important worker case because browsers commonly expose synchronous OPFS access there. The library still -probes the handle instead of saying "DedicatedWorker means sync": +The explicit layers are: ```ts -const capabilities = await probeOpfs(); -if (capabilities.syncAccessHandleExposed) { - const file = await fileSystem.openSyncFile("/database.sqlite", { - create: true, - parents: true, - }); - try { - // synchronous random access - } finally { - file.close(); - } -} +import { createFileSystem } from "@okikio/opfs"; +import { createFileAdapter } from "@okikio/opfs/adapter/file"; +import { createOpfsDriver } from "@okikio/opfs/driver/opfs"; + +const root = await navigator.storage.getDirectory(); +const driver = createOpfsDriver(root); +const fileSystem = createFileSystem(createFileAdapter(driver)); ``` -SharedWorker and ServiceWorker use the same asynchronous filesystem APIs when storage is exposed. A ServiceWorker must keep the -browser event alive itself. The filesystem cannot call `event.waitUntil()` on behalf of the application. +### Window + +Use the asynchronous filesystem API. Do not assume sync access is available from Window. + +### DedicatedWorker + +Use asynchronous methods normally. `openSyncFile()` succeeds only when the actual OPFS file handle exposes synchronous +access in that realm. The library probes the method rather than inferring support from the worker type. + +### SharedWorker + +Use the same capability-driven approach. A SharedWorker does not automatically imply synchronous access. + +### ServiceWorker + +Use asynchronous methods and keep the event lifetime explicit: ```ts self.addEventListener("message", (event) => { - event.waitUntil(saveMessage(event.data)); + const work = (async () => { + const fileSystem = await openFileSystem(); + await fileSystem.writeFile("/events/latest.json", "{}", { parents: true }); + })(); + + event.waitUntil(work); }); ``` -Iframes need policy-aware tests, not a blanket promise ------------------------------------------------------ +The filesystem cannot extend a ServiceWorker event lifetime by itself. -A same-origin iframe normally observes storage under the same applicable storage key as its embedding context. +## Iframes and storage partitioning -A third-party iframe can be partitioned by browser privacy/storage policy. The package does not try to escape that policy -implicitly. The optional iframe API is separate because requesting unpartitioned storage, where the browser supports it, belongs -inside an explicit user-activation and permission flow. +A same-origin iframe opens storage for its current storage key normally. -An opaque sandbox can reject storage because it has no usable origin. `probeOpfs()` returns the actual root result and normalized -failure instead of inferring the outcome from the sandbox flag alone. +A third-party iframe can receive partitioned storage according to browser policy. Normal `openFileSystem()` does not +attempt to escape that policy. + +`@okikio/opfs/iframe` contains the explicit Storage Access API-related OPFS request helpers for browsers that expose +them: ```text -iframe starts - | - v -probe actual storage API - | - +-- root opens ------> use selected strategy - | - `-- root rejected ---> preserve normalized reason -> application fallback +supportsUnpartitionedOpfsRequest +requestUnpartitionedFileSystem ``` -Private browsing, packaged `file:` pages, enterprise browser policy, quota, and persistence can also change availability or -lifetime. The package deliberately does not fingerprint private mode or promise OPFS on `file:` URLs. Probe the realm you are -actually running in. +The application owns user activation and permission presentation. + +A sandboxed opaque-origin iframe can reject storage. Use `probeOpfs()` and the actual normalized error rather than +browser-name guessing. + +## Private browsing and quota + +Private/incognito modes can change availability, quota, persistence, and lifetime. The package does not fingerprint +private browsing. -Playwright tests the browser contexts directly ----------------------------------------------- +```text +probe actual storage capability + | + +-- available -> use selected route + `-- unavailable -> inspect problem/error and choose application fallback +``` -The canonical browser suite uses Playwright Test projects for Chromium, Firefox, and WebKit. The test matrix exercises: +Quota is a dynamic fact. A driver/facade should not advertise one fixed unlimited capacity when the browser/provider +does not guarantee it. -| Context or behavior | Chromium | Firefox | WebKit | -| --- | ---: | ---: | ---: | -| Window async OPFS | probe + execute | probe + execute | probe + execute | -| DedicatedWorker async | probe + execute | probe + execute | probe + execute | -| DedicatedWorker sync handle | probe + open | probe + open | probe + open | -| SharedWorker | probe + execute | probe + execute | probe + execute | -| ServiceWorker | black-box + instrumentation | black-box | black-box | -| same-origin iframe | probe + execute | probe + execute | probe + execute | -| cross-origin iframe | observe policy result | observe policy result | observe policy result | -| opaque sandbox | observe policy result | observe policy result | observe policy result | -| fresh context isolation | execute | execute | execute | -| persistent profile reopen | execute | execute | execute | -| abort before commit | execute | execute | execute | -| localStorage / IndexedDB / Cache adapters | execute | execute | execute | +## Web Locks -Playwright's deeper ServiceWorker inspection is Chromium-specific, so only that instrumentation is browser-specific. The actual -ServiceWorker OPFS behavior stays a black-box page-to-worker message test in every browser that exposes the API. +`coordination: "auto"` uses Web Locks when available and otherwise falls back to one-realm local FIFO locks. -Deno keeps the same library contracts with runtime-specific capabilities ------------------------------------------------------------------------- +`web-locks` can coordinate cooperating same-storage-key browser realms using the same lock names. `local` cannot +coordinate a separate tab/worker process. `none` disables library coordination. -Use `@okikio/opfs/adapter/deno` for a host directory. The runtime needs the filesystem permissions required by the configured -root. The adapter does not request permissions or inspect environment variables itself. +## Deno -Deno KV is a separate adapter because it is a record store, not a host filesystem: +Use the convenience adapter: ```ts -import { createFileSystem } from "@okikio/opfs"; -import { createDenoKvAdapter } from "@okikio/opfs/adapter/deno-kv"; +import { createDenoAdapter } from "@okikio/opfs/adapter/deno"; +``` + +or the explicit driver: -const kv = await Deno.openKv("./data.kv"); -const fileSystem = createFileSystem(createDenoKvAdapter(kv)); +```ts +import { createDenoDriver } from "@okikio/opfs/driver/deno"; ``` -Current Deno documentation still marks KV as unstable. Real Deno KV tests therefore use `--unstable-kv`. The production adapter -accepts a structural KV contract, so simply importing the module does not require a global `Deno` object. +The configured `root` becomes virtual `/`. The runtime needs the filesystem permissions selected by the host +application. + +Deno KV is a separate record driver under `driver/deno-kv`. It is not the host-filesystem driver. -Deno KV also has a much smaller physical value limit than an ordinary filesystem file. The adapter exposes that limit and a -partition policy through `inspect()`. With the default `partition: "auto"`, small materialized files stay inline while large -files and unknown-size replacement streams use a manifest plus raw byte parts. Callers that need a one-record layout can set -`partition: "never"`; the adapter then stops advertising its partitioned replacement-stream lane and large values fail -explicitly instead of changing layout. +## Node -Node and Bun use explicit host adapters ---------------------------------------- +Use `driver/node` and `adapter/node`. The driver uses `node:fs`/`node:fs/promises`/`node:stream` through the explicit +runtime subpath and maps virtual paths below one configured host root. -The Node adapter uses `node:fs` and `node:fs/promises`. It supports native streaming reads and writes, ranges, copy, rename, -synchronous random access, and flush. The configured host root is the only host directory intentionally exposed through the -virtual path namespace. +Node supports native streaming, ranges, copy, rename/move, positional writes, and synchronous random access where +implemented by the driver. -The Bun adapter uses Bun file APIs for the direct byte path and Bun's Node-compatible filesystem APIs for directory, update, -copy/move, and synchronous host-file behavior. The same public tests import `node:test`; Bun currently supports the in-process -`node:test` API when those files are run with `bun test`. +The package engine range starts at Node 22.18. A validation host below that version can provide supplemental evidence +but cannot stand in for the declared runtime matrix. -The Bun benchmark keeps two raw file-copy baselines: Node-compatible `copyFile()` and `Bun.write(destination, Bun.file(source))`. -The second path lets Bun select its file-backed Blob fast path. A Bun-only provider benchmark also compares Bun's native -`S3Client` with the AWS SDK baseline, this package's direct SigV4 client, the object adapter, and the filesystem facade. These -measurements are evidence for route selection; they do not make runtime brand part of the portable API contract. +## Bun -Electron can use the Node adapter in a trusted main-process layer. The OPFS-shaped API is not a reason to expose an arbitrary host -root directly to untrusted renderer content. Use an application-specific IPC/service contract or browser OPFS where that matches -the trust model. +Use `driver/bun` and `adapter/bun`. -Object clients are runtime-neutral Web clients, but deployment policy still matters ------------------------------------------------------------------------------------ +The driver resolves Bun lazily. It uses Bun's native file primitives for the paths where they provide a clear benefit +and the Node-compatible file driver for operations that require stronger host-filesystem behavior. -The S3 and Azure clients use Web Fetch, Web Crypto where signing is required, Web Streams, AbortSignal, and focused `@std/*` -packages for concurrency, byte assembly, stream limits, path mapping, and XML. They are -therefore usable across Deno, Bun, Node, browsers, and workers that expose those Web APIs. +Bun runtime tests are required before claiming Bun behavior complete. Structural TypeScript compatibility alone is not +runtime evidence. -That does not make every deployment equally appropriate. +## Electron -A browser calling S3 directly needs CORS rules that permit the required methods and headers. More importantly, long-lived cloud -storage secrets should not be shipped to untrusted browser code. Use short-lived scoped credentials or a trusted server design. +A trusted Electron main process can use the Node file driver. Do not expose an arbitrary host root directly to untrusted +renderer content merely because the public API looks like OPFS. A renderer should use a controlled IPC service or +browser OPFS according to the application's security model. -The same applies to Azure. SAS tokens and bearer tokens should be scoped to the actual client threat model. The library accepts a -refresh function so a long-lived process does not have to freeze one credential at client creation. +## Record/database environments -Provider endpoints can also have runtime-specific network rules. A Cloudflare Worker, browser extension, serverless host, or -corporate browser policy can allow or reject different origins. Those network policies are outside the filesystem abstraction. +localStorage, IndexedDB, Cache Storage, RxDB, unstorage, db0, Drizzle, and SQLite integrations work where their injected +upstream resource and the required Web primitives are available. -Coordination scope is part of the execution environment -------------------------------------------------------- +The driver describes backend requirements. The adapter/facade describes effective filesystem routes. Database-backed +record storage generally cannot claim native streaming unless the driver implements a dedicated byte lane. -`web-locks` can coordinate cooperating browser realms that share the relevant Web Locks namespace. `local` only coordinates one -JavaScript realm. Separate OS processes and hosts need backend-level coordination where same-path atomicity matters. +## Server coordination -Database/object adapters can use provider transactions, ETags, versions, advisory locks, or leases where the provider exposes -them. The facade does not pretend an in-memory lock became distributed merely because the persisted bytes live on a remote -service. +`coordination: "local"` is one JavaScript realm only. It does not serialize two Node/Deno/Bun processes or two machines. +When cross-process same-path atomicity is required, use the actual backend primitive: + +```text +database transaction +advisory lock +lease +provider conditional write +provider-specific lock/serialization +``` + +A facade-local lock cannot upgrade a best-effort database replacement into a distributed transaction. + +## Import safety + +The root package and provider-neutral core do not import runtime-specific implementations automatically. + +Use explicit subpaths for: + +```text +driver/node +driver/deno +driver/bun +driver/deno-kv +driver/s3 +driver/azure +adapter/* +``` -Provider protocol details are kept in [S3 client protocol](./s3.md), [Azure Blob client protocol](./azure.md), and the -[provider integration test guide](./providers.md). Shared Key is intended for trusted server/Azurite contexts because it exposes -the Azure storage account key to the runtime. +No driver reads environment variables or configures global application logging at import time. diff --git a/docs/providers.md b/docs/providers.md index e7d1d8d..f95b6bb 100644 --- a/docs/providers.md +++ b/docs/providers.md @@ -1,232 +1,233 @@ -Object-provider integration tests -================================= +# Provider and filesystem baseline tests -Purpose -------- +## Purpose -The direct S3 and Azure clients own HTTP protocol behavior that a pure mock cannot prove. This guide explains the local provider -fixtures used to test real signing, request routing, ranges, multipart/block state, copy, listing, and filesystem translation -without requiring cloud credentials for every maintainer run. +S3 and Azure Blob have two kinds of external baseline: -Provider containers supplement deterministic protocol tests. They do not replace the Amazon S3 or Azure Blob specifications. -They also do not become a production dependency or a storage abstraction. Testcontainers exists only in development tooling. +1. protocol/API clients such as the AWS SDK and Azure SDK; +2. filesystem clients such as AWS Mountpoint and Azure BlobFuse. -Testcontainers owns provider lifecycle --------------------------------------- +They answer different questions. SDK/provider tests validate object protocol behavior. Filesystem-client benchmarks +compare the performance and semantics of a mature provider-to-filesystem translation against this package's own +driver/adapter/facade stack. -`tests/provider/fixture.ts` uses Testcontainers for Node.js instead of a repository-owned Docker Compose and polling harness: +## Testcontainers owns disposable protocol providers + +`tests/provider/fixture.ts` starts local provider services with Testcontainers for Node.js. ```text -node:test - | - v -ProviderFixture - | - +--> Testcontainers GenericContainer - | | - | `--> SeaweedFS 4.41 - | S3-compatible API - | - `--> @testcontainers/azurite - | - `--> Azurite 3.36.0 - Azure Blob API +node:test / Mitata + | + ProviderFixture + | + +-- SeaweedFS + | S3-compatible endpoint + | + `-- Azurite + Azure Blob endpoint ``` -The images remain pinned: +The fixture uses mapped host ports and Testcontainers-owned readiness/cleanup. There is no repository-owned Docker +Compose lifecycle, fixed provider port, curl readiness loop, or shell trap for these services. + +The provider fixture is test infrastructure only. Public package code does not import Testcontainers. + +## SeaweedFS S3 target + +SeaweedFS is used as one independent S3-compatible implementation. The test config supplies endpoint, bucket, region, +and test credentials to: ```text -chrislusf/seaweedfs:4.41 -mcr.microsoft.com/azure-storage/azurite:3.36.0 +AWS SDK +project S3 client +project S3 driver +project S3 adapter +FileSystemType ``` -Testcontainers chooses free host ports and waits for the exposed service instead of requiring fixed `8333` and `10000` host -ports. The S3 fixture combines listening-port and HTTP readiness. The official Azurite module owns its emulator-specific startup -contract, uses in-memory persistence, skips only Azurite's API-version allow-list check, and exposes the mapped Blob endpoint. +Tests cover the implemented shared contract, including: -SeaweedFS still uses `GenericContainer` because the Testcontainers Node catalog does not provide a SeaweedFS module. The code -uses the official Azurite module because Testcontainers recommends a focused module when one exists instead of duplicating that -container's configuration in every project. +- signed put/head/get/delete; +- byte ranges; +- conditional create/replace where supported by the fixture; +- multipart streamed replacement; +- server-side copy; +- prefix/delimiter listing; +- filesystem directory/object translation. -`ProviderFixture` owns every container it starts. `close()` is idempotent and attempts to stop every owned container even if one -cleanup operation fails. Partial construction also stops SeaweedFS if Azurite cannot start. This keeps acquisition and cleanup in -one place instead of spreading lifecycle work across shell traps, readiness polling, and the test body. +SeaweedFS does not define Amazon S3 semantics. AWS-specific limits and canonical signing behavior remain covered by +deterministic protocol tests and AWS primary documentation. -Testcontainers is an interim compute layer ------------------------------------------- +## Azurite target -The project deliberately does not treat Testcontainers as the final runtime/provider abstraction. Testcontainers Node currently -centers Docker-compatible container runtimes. Its documentation covers Docker directly and documents setup/limitations for -Podman, Colima, and Rancher Desktop. +The official `@testcontainers/azurite` module supplies a Blob endpoint and test account credentials. -That is sufficient for the current local service fixtures. It is not the model for future Apple `container`, WSL containers, -microVMs, cloud VMs, Kubernetes, or other compute providers. A future environment/provider layer can replace how fixtures are -started while preserving this test contract: +The provider suite exercises: ```text -provider fixture - | - +--> endpoint - +--> credentials - +--> lifecycle ownership - `--> diagnostics - -protocol/client tests consume only those facts +Azure SDK +project Azure client +project Azure driver +project Azure adapter +FileSystemType ``` -The provider test therefore does not inspect Docker container IDs or Docker networks after startup. Those are fixture mechanics, -not S3/Azure test semantics. +Coverage includes: -Run the provider suite ----------------------- +- Shared Key authentication against the emulator; +- blob put/head/range/delete; +- block upload and block-list commit; +- server-side copy; +- container listing/prefix translation; +- filesystem translation. -The canonical command is: +Azurite is an emulator. A green Azurite result is interoperability evidence, not proof of every cloud Azure service +version or feature. -```sh -mise run test-providers +## Provider benchmark staircase + +`bench/provider.bench.ts` keeps every layer visible. + +S3: + +```text +AWS SDK + -> project S3 client + -> project S3 driver + -> project object adapter + -> facade metrics:none + -> facade metrics:basic ``` -The task is intentionally small: +Azure: ```text -mise - | - +--> deno ci - | - `--> node --test tests/provider.test.ts - | - `--> Testcontainers owns startup/readiness/cleanup +Azure SDK + -> project Azure client + -> project Azure driver + -> project object adapter + -> facade metrics:none + -> facade metrics:basic ``` -`node:test` remains the repository test runner. Testcontainers supplies resources to the test; it does not become a second test -framework. +These are separate benchmark samples, not one chain executed for each operation. The staircase wording means each result +adds one project layer over the same configured provider. -GitHub Actions calls the same mise task. The workflow installs mise, asks mise for the Deno and Node versions required by the -provider job, and then runs `mise run test-providers`. GitHub Actions owns only job topology, permissions, runner selection, -timeouts, and secrets. It does not duplicate provider startup commands. +`bench/bun-provider.bench.ts` adds Bun's native S3 client as another S3 baseline when Bun is available. -Playwright owns browser lifecycle separately --------------------------------------------- +Container startup and image pull are completed before measured Mitata cases begin. -Testcontainers and Playwright solve different lifecycle problems: +## Physical versus logical metrics -```text -node:test + Testcontainers - S3/Azure service interoperability +The provider client/driver can report physical counters such as requests, retries, responses, failures, and +multipart/block work. The facade reports logical filesystem operations and facade buffering. -Playwright Test - Chromium/Firefox/WebKit runtime interoperability - Window/Worker/ServiceWorker/iframe/storage lifecycle -``` +A benchmark should not infer provider requests from logical operations. One logical large write can become many physical +parts. + +## AWS Mountpoint baseline -The browser matrix stays under `tests/browser/`. Playwright owns browser installation, isolated `BrowserContext` instances, -persistent profiles, traces, retries, and Vite fixture-server lifecycle. Provider tests do not launch browsers, and browser tests -do not launch provider containers merely to share a framework. +AWS Mountpoint is a separate filesystem-client comparator. It translates file operations to S3 and deliberately supports +a subset of ordinary filesystem semantics. -Provider benchmarks keep startup outside timed work ----------------------------------------------------- +The package does not assume Mountpoint is a generic S3-compatible implementation. The benchmark target should use Amazon +S3 when that is the intended conformance/performance comparison. A custom endpoint can be used for local experimentation +when Mountpoint supports the selected endpoint configuration, but that does not turn the result into an AWS-supported +compatibility claim. -The same Testcontainers fixture owns provider startup for: +Mountpoint setup/mount lifecycle is external to the normal Testcontainers provider fixture because it is a host +FUSE/system client. After the mount exists, set: ```sh -mise run bench-providers +OPFS_MOUNTPOINT_S3_ROOT=/path/to/mount +mise run bench-filesystem-clients ``` -`bench/providers.ts` starts SeaweedFS and Azurite once, obtains their random host endpoints, and passes those endpoints to the -actual benchmark programs. Container startup, image pull, and readiness time therefore do not enter a Mitata sample. - -S3 is measured as: +The benchmark then measures: ```text -AWS SDK v3 baseline -Bun native S3Client baseline -@okikio/opfs direct SigV4 client -direct ObjectStore adapter -FileSystemType with metrics none -FileSystemType with metrics basic +raw Node fs against the mount +Node file driver against the mount +file adapter against the driver +FileSystemType against the adapter ``` -Azure is measured as: +This shows the project overhead when the provider-to-filesystem translation is owned by Mountpoint rather than by the +package's S3 client/driver. + +## Azure BlobFuse baseline + +BlobFuse is the analogous Azure Blob filesystem-client comparator. It also has caching/configuration semantics that can +change observable filesystem behavior and performance. + +Provision/mount BlobFuse externally, then run: + +```sh +OPFS_BLOBFUSE_ROOT=/path/to/mount +mise run bench-filesystem-clients +``` + +The same raw -> driver -> adapter -> facade staircase is measured. + +Caching mode and write mode must be recorded with benchmark results. A cached BlobFuse read is not semantically +equivalent to a direct uncached REST read simply because both return the same bytes in one sample. + +## Why filesystem-client mounts are not a default CI job + +Mountpoint and BlobFuse need operating-system packages, FUSE support, mount permissions, and cleanup. Those requirements +are materially different from a disposable HTTP test container. + +The repository therefore provides a canonical benchmark program and mise task, while the runner owns system-level mount +setup. A dedicated privileged benchmark runner can automate installation/mounting without making ordinary pull-request +CI depend on FUSE privileges. + +## Comparable operations only + +Filesystem clients do not necessarily implement all POSIX operations, and object stores do not naturally have POSIX +semantics. The benchmark therefore starts from supported operations rather than treating an unsupported operation as a +slow operation. + +For each baseline, record: ```text -@azure/storage-blob baseline -@okikio/opfs direct Azure REST client -direct ObjectStore adapter -FileSystemType with metrics none -FileSystemType with metrics basic +operation +native/support status +cache mode +write mode +payload size +concurrency +elapsed/throughput +request/physical metrics when available ``` -Small replacement and multipart/block cases remain separate. Different request plans must not be reported as facade overhead. -Loopback results are diagnostics about client/abstraction cost, not cloud-throughput claims. - -What the live tests prove -------------------------- - -The S3 path validates: - -- a real SigV4 HTTP request accepted by an independent S3-compatible server; -- PUT and HEAD; -- byte-range GET; -- create-only conditional replacement; -- multipart stream upload with a legal non-final part size; -- provider-side copy; -- prefix listing; -- delete cleanup; -- `ObjectStoreType -> AdapterType -> FileSystemType` translation. - -The Azure path validates: - -- Shared Key accepted by Azurite; -- explicit container creation; -- Put Blob and Get Blob Properties; -- byte-range GET; -- create-only conditional replacement; -- Put Block / Put Block List streaming upload; -- same-account server-side copy; -- prefix listing; -- delete cleanup; -- `ObjectStoreType -> AdapterType -> FileSystemType` translation. - -The provider receives the actual headers, query fields, bytes, XML, and signatures created by the library. - -What the live tests do not prove --------------------------------- - -A compatible server or emulator cannot prove every detail of a cloud service. The suite does not use it as the oracle for: - -- exact AWS canonical-request text; -- AWS-only HTTP-200 embedded error bodies; -- every S3 service limit; -- Azure Shared Key construction independently of Azurite; -- every historical Azure service-version size limit; -- cloud role/identity acquisition; -- region routing and redirects; -- provider throttling; -- cloud durability or consistency guarantees; -- billing, retention, encryption, replication, or versioning. - -Those cases belong to deterministic protocol tests where possible and opt-in real-cloud suites when a local provider cannot -represent the behavior faithfully. - -Why the matrix keeps both test styles -------------------------------------- - -| Test style | Strong at | Weak at | -| ---------- | --------- | ------- | -| Deterministic request test | Exact canonical text, headers, limits, branch selection | Real HTTP parser/auth integration | -| Testcontainers provider | Real socket/HTTP/auth/protocol interoperability | Complete cloud parity | -| Optional real cloud | Actual provider behavior | Cost, credentials, availability, reproducibility | - -A client change that affects signing, multipart/block state, copy, conditions, retries, cancellation, or provider errors should -update the deterministic test and the provider integration when the local implementation can represent that behavior. - -Future provider breadth ------------------------ - -The current fixture is deliberately small. Additional S3-compatible services should be added only when they exercise a materially -different contract, not to increase a provider count. Useful differences include addressing, copy/condition support, multipart -errors, non-AWS region behavior, and pagination. - -Network-fault fixtures are also a useful next layer. Testcontainers provides a Toxiproxy module, which can be used to prove -retry, timeout, cancellation, and cleanup behavior against a real socket path without adding unreliable sleeps to the tests. -That belongs in a focused failure suite rather than the normal happy-path provider test. +A performance comparison is valid only when the operation semantics being compared are close enough to answer the same +question. + +## Real cloud suites + +Local provider fixtures should be supplemented by opt-in cloud suites before high-confidence releases of protocol +changes. + +Amazon S3: + +- short-lived credentials; +- disposable bucket/prefix; +- explicit cleanup; +- same deterministic operation set used locally where service semantics match. + +Azure Blob: + +- short-lived identity/SAS credentials; +- disposable container/prefix; +- explicit cleanup; +- service-version coverage relevant to the changed code. + +Cloud credentials must never become required for ordinary portable tests. + +## Failure injection + +Network fault injection is still future work. Toxiproxy or another socket-level test service can add latency, reset, +timeout, retry, cancellation, and multipart/block cleanup scenarios around the real provider clients. + +It should remain test-fixture infrastructure. The public OPFS/client/driver architecture must not depend on a +fault-injection service. diff --git a/docs/releasing.md b/docs/releasing.md index fa75192..758d9ca 100644 --- a/docs/releasing.md +++ b/docs/releasing.md @@ -1,111 +1,185 @@ -Release process -=============== +# Release process -`@okikio/opfs` publishes one Git release to JSR and npm. Git history decides the version. `deno.json` owns the authored package graph and public exports. semantic-release does not commit generated version files back to `main`. +## Purpose -Conventional commits --------------------- +The repository has one command authority: mise. GitHub Actions decides when a release/publish operation should run and +supplies GitHub-specific permissions, refs, secrets, and outputs. The actual project commands live under `.mise/tasks/`. -Commits merged to `main` use Conventional Commits 1.0.0. Release-relevant examples are: +```text +GitHub Actions + | + jdx/mise-action + | + mise task + | +Deno / Node / Bun / npm / semantic-release +``` + +## Conventional commits + +Commits merged to `main` use Conventional Commits. ```text -fix: correct a public behavior -> patch -feat: add a compatible capability -> minor -feat!: replace a public contract -> major +fix: ... patch +feat: ... minor +feat!: ... major -BREAKING CHANGE: describe the consumer impact -> major +BREAKING CHANGE: ... ``` -`build`, `chore`, `ci`, `docs`, `refactor`, `style`, and `test` are valid types. They do not cause a release unless the analyzed commit carries a breaking change. Pull requests should normally squash to one conventional commit so `main` remains a clear release input. +Commit validation runs through: -The checked-in `0.0.1` is development metadata, not the first public release decision. semantic-release does not support selecting `0.0.1` as the initial stable version. With no earlier release tag, the normal first release on `main` is `1.0.0`. If the project is not ready for a stable `1.0.0`, configure a prerelease branch before enabling the release workflow. Registry commands receive the semantic-release version through `--set-version`, so they do not publish the development placeholder. +```sh +BASE_SHA=... HEAD_SHA=... mise run commits +``` -Dependency graph ----------------- +The mise task pins the commitlint CLI package used by CI. Workflow YAML does not own a second commitlint command. -Dependency changes are not complete until both package-manager views are reproducible. After changing `deno.json` or -`package.json`, update the Deno lockfile intentionally with: +## Quality before release + +The normal release evidence includes: + +```sh +mise run quality +mise run test-deno +mise run test-node +mise run test-bun +mise run test-browser +mise run test-providers +mise run verify-npm +``` + +CI also runs benchmark smoke jobs so gross performance/runtime regressions are visible before a semantic release can +run. + +The quality task includes strict checks, lint, documentation lint, formatting, stress, coverage, and package dry-runs. + +## Lockfiles + +Dependency changes are incomplete until both package-manager views are reproducible. + +After dependency/import changes, regenerate with the canonical managers: ```sh deno install --frozen=false pnpm install ``` -Review and commit both `deno.lock` and `pnpm-lock.yaml`. CI uses `deno ci`, which deliberately rejects a missing or stale Deno -lockfile instead of resolving an unseen dependency graph during the build. +Review and commit `deno.lock` and `pnpm-lock.yaml`. Do not hand-edit integrity/resolution data merely because one +validation host cannot access the registry. + +## Semantic release + +The release workflow starts only after the successful `CI` workflow for `main`. + +It checks out the tested commit, verifies that the checked-out `main` commit still equals the CI commit, then runs: + +```sh +mise run release +``` +The mise task invokes pinned semantic-release tooling. semantic-release determines the version from Git history, creates +the Git tag, and creates the GitHub release according to repository configuration. -Release flow ------------- +The repository uses tags shaped as: ```text -push to main - | - v -CI - | frozen dependencies, format, lint, docs, types, - | runtime tests, browser matrix, stress, package dry-runs - v -Release workflow - | - +-- verify main still equals the tested CI SHA - +-- semantic-release reads commits since opfs@ - +-- semantic-release creates the Git tag and GitHub Release - | - `-- only when a new tag appeared: - workflow_dispatch Publish Registries - | - +----> JSR - `----> npm +opfs@1.2.3 ``` -The registry workflow is dispatched explicitly instead of relying on the GitHub `release` event. A release created with the repository `GITHUB_TOKEN` does not normally trigger another workflow from that release event. GitHub does allow `workflow_dispatch` events created with `GITHUB_TOKEN`, so the release workflow can start the separate publisher without a long-lived release PAT. +## Registry publication uses the immutable tag -A rerun on a commit that already has an `opfs@` tag does not dispatch another automatic publication. Manual registry retries remain available through `Publish Registries` and require the existing immutable tag. +The publish workflow receives an existing release tag. It never publishes arbitrary `main` state. -semantic-release owns only version analysis, tag creation, release notes, and the GitHub Release. It does not publish either registry and it does not create a release commit. +```text +release workflow + | +created opfs@X.Y.Z + | +dispatch publish.yml(tag=opfs@X.Y.Z) + | +checkout exact tag + | +JSR and/or npm +``` -npm packaging -------------- +A partial registry failure can be retried against the same immutable tag by selecting one target. -JSR consumes the authored TypeScript package directly. npm receives generated JavaScript and declaration files. +## JSR -`.mise/tasks/npm` runs `deno pack --no-deno-shim --set-version`. Deno derives the npm graph and exports from the same `deno.json` used by JSR. Drizzle needs one npm-only correction: it is an optional integration, so the task removes `drizzle-orm` from generated normal dependencies and writes it as an optional peer before the final `npm pack`. +JSR publication runs: -The generated tarball is then installed into a clean consumer. The verifier checks exports and declarations, Node/Deno/Bun imports when those runtimes are present, browser bundling, and the optional Drizzle subpath after the peer is installed. +```sh +RELEASE_VERSION=X.Y.Z mise run publish-jsr +``` -Registry setup --------------- +The task performs a frozen Deno install, a JSR dry-run, then the actual publish with the release version supplied +explicitly. -Before the first release: +## npm -1. Create `@okikio/opfs` on JSR and link it to `okikio/opfs` for GitHub OIDC publication. -2. Bootstrap the first npm publication with a granular/automation token in `NPM_TOKEN` when trusted publisher settings are not yet available for the package. -3. After npm has the package, configure trusted publishing for `.github/workflows/publish.yml`. -4. Remove `NPM_TOKEN` when the bootstrap fallback is no longer wanted. +npm publication runs: -The publisher uses GitHub OIDC for JSR and for normal npm trusted publication. npm is upgraded to a current npm 11 release in the publish job before trusted publishing. +```sh +RELEASE_VERSION=X.Y.Z mise run publish-npm +``` -Partial failure ---------------- +The task: -If one registry succeeds and the other fails, do not create another version. Run `Publish Registries` manually with the same existing `opfs@` tag and select only the failed registry. +1. resolves bootstrap token versus trusted-publishing mode; +2. runs `mise run verify-npm`; +3. builds the npm package from the Deno package graph; +4. verifies the tarball from Node/Deno/Bun consumers; +5. publishes the exact generated tarball with provenance. -Local release checks --------------------- +`drizzle-orm` remains an optional peer in the generated npm package rather than being converted into a mandatory runtime +dependency. -A Deno-capable machine can run: +## Bootstrap versus trusted npm publishing -```sh -mise install -mise run check -mise run test -mise run bench -deno task release:check +The publish workflow can choose: + +```text +auto +trusted +token ``` -The GitHub release and registry workflows use mise for their Node, Deno, and Bun toolchain too. Release-only policy remains in -GitHub Actions: OIDC permissions, immutable tag resolution, semantic-release, registry authentication, and the final publish -commands are deployment concerns rather than reusable repository tasks. +`auto` checks whether `@okikio/opfs` already exists on npm through `mise run npm-exists`. + +- first publication: token mode; +- later publications: trusted mode when configured. + +The workflow owns secret selection. The npm publish command still lives in the mise task. + +## Package verification + +`mise run verify-npm` builds a test-version tarball and executes the package consumer verifier. The verifier must +exercise the published package, not source-tree relative imports. + +A release is not accepted because `npm pack` returned success. The generated manifest, exports, optional peers, and +runtime consumer entrypoints must also be checked. + +## Version metadata + +The checked-in development version is not the release decision. The immutable Git tag and semantic-release result define +the published version. Packaging receives that version explicitly instead of committing generated version edits back to +`main`. + +## Recovery rules + +If GitHub release creation succeeds but one registry publish fails: + +1. do not create a replacement tag; +2. do not rebuild from a newer `main` commit; +3. re-run `Publish Registries` with the same immutable tag; +4. select only the failed registry when appropriate. + +If validation fails before release creation, fix the source and produce a new commit. Do not weaken a quality task +solely to make the release workflow progress. + +## Trusted npm publishing -`deno task release:check` performs package dry-runs. It does not upload a release. +The publish job installs `npm:npm@11.18.0` through mise. npm trusted publishing requires npm CLI 11.5.1 or later and +Node 22.14.0 or later. The explicit npm pin prevents the release path from depending on the bundled npm version of the +selected Node release. diff --git a/docs/s3.md b/docs/s3.md index 349162d..8537658 100644 --- a/docs/s3.md +++ b/docs/s3.md @@ -1,25 +1,24 @@ -S3 client protocol guide -======================== +# S3 client protocol guide -Purpose -------- +## Purpose -This document defines the S3 protocol contract implemented by -`@okikio/opfs/s3`. It is written for maintainers who need to change signing, -upload, copy, listing, conditional-write, or compatibility behavior without -silently changing the public filesystem guarantees. +This document defines the S3 protocol contract implemented by `@okikio/opfs/s3`. It is written for maintainers who need +to change signing, upload, copy, listing, conditional-write, or compatibility behavior without silently changing the +public filesystem guarantees. -The client is intentionally not a replacement for the complete AWS SDK. It -implements the S3 REST operations needed by the object-store adapter and keeps a -low-level signed `request()` method for S3-compatible features that do not belong -in the portable filesystem API. +The client is intentionally not a replacement for the complete AWS SDK. It implements the S3 REST operations needed by +the S3 object driver and keeps a low-level signed `request()` method for S3-compatible features that do not belong in +the portable filesystem API. The implementation is direct: ```text -ObjectStoreType / S3ClientType - | - v +S3ClientType + | + +--> direct protocol use + | + `--> S3 object driver -> object adapter -> FileSystemType + request construction | v @@ -32,40 +31,35 @@ Web Fetch `--> S3-compatible endpoint ``` -The protocol code uses Web Crypto for SHA-256 and HMAC-SHA256. `@std/encoding` -owns hexadecimal encoding, `@std/async/pool` owns bounded multipart concurrency, -`@std/xml` owns XML parsing/serialization, and `@std/bytes` supports the shared -chunk layer. No AWS SDK package is imported. +The protocol code uses Web Crypto for SHA-256 and HMAC-SHA256. `@std/encoding` owns hexadecimal encoding, +`@std/async/pool` owns bounded multipart concurrency, `@std/xml` owns XML parsing/serialization, and `@std/bytes` +supports the shared chunk layer. No AWS SDK package is imported. This document distinguishes three evidence classes: - - **Implemented** means the current source contains the behavior. - - **Protocol** means the behavior is required or described by current AWS S3 - documentation. - - **Provider-dependent** means an S3-compatible service can intentionally - differ and the client exposes configuration for that difference. - -The current implementation was reviewed against AWS S3 REST documentation on -August 14, 2026. The authoritative source remains AWS documentation when this -file and the service specification disagree. +- **Implemented** means the current source contains the behavior. +- **Protocol** means the behavior is required or described by current AWS S3 documentation. +- **Provider-dependent** means an S3-compatible service can intentionally differ and the client exposes configuration + for that difference. +The current implementation was reviewed against AWS S3 REST documentation on August 14, 2026. The authoritative source +remains AWS documentation when this file and the service specification disagree. -The client owns a focused S3 contract -------------------------------------- +## The client owns a focused S3 contract -`S3ClientType` extends the package `ObjectStoreType`, so it provides the object -operations the filesystem adapter needs: +`S3ClientType` implements the object-backend operations consumed by `driver/s3`, while retaining S3-specific protocol +methods. The object operations used by the driver are: -| Object operation | S3 REST operation | Current behavior | -| ---------------- | ----------------- | ---------------- | -| `head()` | `HeadObject` | Exact-key metadata or `null` for `404` | -| `get()` | `GetObject` | Complete object or one `Range` | -| `put(Uint8Array)` | `PutObject` | One materialized request up to 5 GB | -| `put(stream)` | Multipart upload | Bounded concurrent parts and explicit completion | -| `delete()` | `DeleteObject` | Missing object is already removed | -| `list()` | `ListObjectsV2` | Prefix, delimiter, limit, continuation token | -| `copy()` small | `CopyObject` | Server-side copy through the 5 GB single-copy limit | -| `copy()` large | Multipart `UploadPartCopy` | Server-side ranged copy without JS body transfer | +| Object operation | S3 REST operation | Current behavior | +| ----------------- | -------------------------- | --------------------------------------------------- | +| `head()` | `HeadObject` | Exact-key metadata or `null` for `404` | +| `get()` | `GetObject` | Complete object or one `Range` | +| `put(Uint8Array)` | `PutObject` | One materialized request up to 5 GB | +| `put(stream)` | Multipart upload | Bounded concurrent parts and explicit completion | +| `delete()` | `DeleteObject` | Missing object is already removed | +| `list()` | `ListObjectsV2` | Prefix, delimiter, limit, continuation token | +| `copy()` small | `CopyObject` | Server-side copy through the 5 GB single-copy limit | +| `copy()` large | Multipart `UploadPartCopy` | Server-side ranged copy without JS body transfer | The S3-specific surface also exposes: @@ -77,14 +71,45 @@ completeUpload() abortUpload() ``` -Those operations are public because an S3 consumer can need storage class, -checksums, encryption, object lock, tagging, or another provider extension that -is not a filesystem concern. The low-level request method signs the request but -does not interpret every S3 feature on the caller's behalf. +Those operations are public because an S3 consumer can need storage class, checksums, encryption, object lock, tagging, +or another provider extension that is not a filesystem concern. The low-level request method signs the request but does +not interpret every S3 feature on the caller's behalf. + +The driver remains a separate public layer: + +```ts +import { createS3Client } from "@okikio/opfs/s3"; +import { createS3DriverFromClient } from "@okikio/opfs/driver/s3"; +import { createObjectAdapter } from "@okikio/opfs/adapter/object"; + +const client = createS3Client(options); +const driver = createS3DriverFromClient(client); +const adapter = createObjectAdapter(driver); +``` + +The driver adds provider requirements, hard limits, optimization state, deterministic planning, and physical request +metrics. The adapter then translates object prefixes and values into canonical filesystem primitives. + +## Optimizations are independently controllable + +The S3 client currently exposes two optimization switches. Both are visible through the S3 driver inspection. + +`delayedMultipart` : Defaults to true. For an unknown-length stream, the client buffers the first bounded multipart +chunk before creating an upload. If EOF arrives while the object can use one `PutObject`, the client avoids multipart +initiation. If more data arrives, or the first chunk is already above the single-PUT ceiling, it starts multipart and +replays the retained bytes into the multipart lane. Set this to false when exact multipart request lifecycle is required +for testing or application policy. + +`signingKeyCache` : Defaults to true. The client caches the derived SigV4 signing key for the current +credentials/date/region/service tuple. The cache is per client and invalidates when the tuple changes. Disabling it +forces key derivation on every signed request. The optimization changes CPU work but not the canonical request or +resulting signature. +The client reports both switches through `client.optimizations`, while `driver.inspect().optimizations` exposes the same +state in the generic driver model. Any future optimization that changes request count, provider resource lifetime, +failure timing, or storage layout must also be independently disableable. -Addressing and canonical request construction ---------------------------------------------- +## Addressing and canonical request construction The client supports path-style and virtual-hosted-style addressing. @@ -100,30 +125,28 @@ Virtual-hosted style: https://bucket.endpoint.example/path/to/object ``` -`addressing: "path"` is the compatibility-oriented default because many local -and third-party S3 implementations expose one HTTP endpoint without wildcard -DNS for buckets. +`addressing: "path"` is the compatibility-oriented default because many local and third-party S3 implementations expose +one HTTP endpoint without wildcard DNS for buckets. -The client preserves slash separators in object keys while percent-encoding each -path segment. Signature Version 4 is sensitive to exact path and query -serialization. The implementation therefore does not use locale-sensitive -sorting and does not normalize the canonical URI after object-key construction. +The client preserves slash separators in object keys while percent-encoding each path segment. Signature Version 4 is +sensitive to exact path and query serialization. The implementation therefore does not use locale-sensitive sorting and +does not normalize the canonical URI after object-key construction. For the canonical query string, the client: -1. expands repeated query values; -2. URI-encodes each name and value; -3. sorts the encoded names and then encoded values by code-unit order; -4. joins the pairs with `&`. +1. expands repeated query values; +2. URI-encodes each name and value; +3. sorts the encoded names and then encoded values by code-unit order; +4. joins the pairs with `&`. For canonical headers, the client: -1. lowercases header names; -2. normalizes internal whitespace; -3. sorts the signed header names; -4. includes the request authority as `host` even though browser Fetch controls - the actual Host or HTTP/2 `:authority` field; -5. includes required `x-amz-*` headers. +1. lowercases header names; +2. normalizes internal whitespace; +3. sorts the signed header names; +4. includes the request authority as `host` even though browser Fetch controls the actual Host or HTTP/2 `:authority` + field; +5. includes required `x-amz-*` headers. The signing flow is: @@ -148,64 +171,52 @@ kDate -> kRegion -> kService(s3) -> kSigning Authorization signature ``` -Credential sources can be static or refreshable. Refreshable credentials are -resolved immediately before each signed request. Session credentials add -`x-amz-security-token` before canonical signing. +Credential sources can be static or refreshable. Refreshable credentials are resolved immediately before each signed +request. Session credentials add `x-amz-security-token` before canonical signing. -The client uses Web Crypto directly because browser-compatible SHA-256 and HMAC -are already platform APIs. `@std/crypto` does not implement S3 Signature -Version 4, so adding it would not remove protocol code or improve ownership. +The client uses Web Crypto directly because browser-compatible SHA-256 and HMAC are already platform APIs. `@std/crypto` +does not implement S3 Signature Version 4, so adding it would not remove protocol code or improve ownership. ### Payload hashes -Replayable Web request bodies are SHA-256 hashed before signing. The client -hashes strings, `ArrayBuffer`, `ArrayBufferView`, `Blob`, and -`URLSearchParams` values. Requests without bodies use the standard SHA-256 of -an empty payload. +Replayable Web request bodies are SHA-256 hashed before signing. The client hashes strings, `ArrayBuffer`, +`ArrayBufferView`, `Blob`, and `URLSearchParams` values. Requests without bodies use the standard SHA-256 of an empty +payload. -A low-level caller can supply `payloadHash`. This exists for S3 modes such as -`UNSIGNED-PAYLOAD` where the selected provider accepts that contract. The -client does not silently choose an unsigned payload when the exact request -bytes can be determined without consuming or re-encoding caller state. +A low-level caller can supply `payloadHash`. This exists for S3 modes such as `UNSIGNED-PAYLOAD` where the selected +provider accepts that contract. The client does not silently choose an unsigned payload when the exact request bytes can +be determined without consuming or re-encoding caller state. -A streamed low-level body cannot be consumed once merely to calculate a hash and -then consumed again by Fetch. `FormData` has a related problem because Fetch -owns its multipart boundary and exact wire encoding. Those two low-level body -forms therefore use `UNSIGNED-PAYLOAD` unless the caller supplies an explicit -`payloadHash`. Callers that need AWS streaming-signature chunk framing must -implement that S3-specific streaming mode above `request()`. The normal -high-level streamed `put()` avoids this ambiguity by using multipart parts, -each of which is materialized before its individual signed request. +A streamed low-level body cannot be consumed once merely to calculate a hash and then consumed again by Fetch. +`FormData` has a related problem because Fetch owns its multipart boundary and exact wire encoding. Those two low-level +body forms therefore use `UNSIGNED-PAYLOAD` unless the caller supplies an explicit `payloadHash`. Callers that need AWS +streaming-signature chunk framing must implement that S3-specific streaming mode above `request()`. The normal +high-level streamed `put()` avoids this ambiguity by using multipart parts, each of which is materialized before its +individual signed request. - -Object and multipart limits are part of planning ------------------------------------------------- +## Object and multipart limits are part of planning `S3_LIMITS` records the protocol limits that affect this implementation: -| Limit | Value used by the client | Why it matters | -| ----- | ------------------------ | -------------- | -| Maximum object | 53,687,091,200,000 bytes | Exact 10,000 x 5 GiB multipart ceiling (48.8 TiB) | -| Single `PutObject` | 5,000,000,000 bytes | Larger replacement must use multipart upload | -| Single `CopyObject` | 5,000,000,000 bytes | Larger copy must use multipart copy | -| Minimum multipart part | 5 MiB | Every non-final upload part must meet the S3 minimum | -| Maximum multipart part | 5 GiB | Client part-size configuration cannot exceed it | -| Maximum multipart parts | 10,000 | Known-size streams must choose a large enough part size | - -The distinction between GB and GiB is intentional. AWS documents the -single-request PUT/copy threshold in decimal GB, while multipart part limits use -binary-sized MiB/GiB values. AWS product documentation often calls the maximum -object size 50 TB, while the current object guide and multipart arithmetic make -the exact ceiling 10,000 x 5 GiB = 53,687,091,200,000 bytes (48.8 TiB, about -53.7 TB). The client uses the exact multipart-derived value so it does not +| Limit | Value used by the client | Why it matters | +| ----------------------- | ------------------------ | ------------------------------------------------------- | +| Maximum object | 53,687,091,200,000 bytes | Exact 10,000 x 5 GiB multipart ceiling (48.8 TiB) | +| Single `PutObject` | 5,000,000,000 bytes | Larger replacement must use multipart upload | +| Single `CopyObject` | 5,000,000,000 bytes | Larger copy must use multipart copy | +| Minimum multipart part | 5 MiB | Every non-final upload part must meet the S3 minimum | +| Maximum multipart part | 5 GiB | Client part-size configuration cannot exceed it | +| Maximum multipart parts | 10,000 | Known-size streams must choose a large enough part size | + +The distinction between GB and GiB is intentional. AWS documents the single-request PUT/copy threshold in decimal GB, +while multipart part limits use binary-sized MiB/GiB values. AWS product documentation often calls the maximum object +size 50 TB, while the current object guide and multipart arithmetic make the exact ceiling 10,000 x 5 GiB = +53,687,091,200,000 bytes (48.8 TiB, about 53.7 TB). The client uses the exact multipart-derived value so it does not reject objects that S3 can legally assemble. -`partSize` defaults to 8 MiB. `copyPartSize` defaults to 1 GiB. Both are -validated when the client is created. +`partSize` defaults to 8 MiB. `copyPartSize` defaults to 1 GiB. Both are validated when the client is created. -If `ObjectPutOptionsType.size` is supplied, the upload planner calculates the -minimum part size required to stay at or below 10,000 parts and uses the larger -of that value and the configured `partSize`. +If `ObjectPutOptionsType.size` is supplied, the upload planner calculates the minimum part size required to stay at or +below 10,000 parts and uses the larger of that value and the configured `partSize`. ```text known body size @@ -220,26 +231,21 @@ ceil(size / 10,000) `--> reject if > 5 GiB ``` -When a stream size is unknown, the configured part size remains authoritative. -The chunk iterator fails before part 10,001 instead of creating an invalid -multipart request. A caller with a very large stream should supply `size` so -the planner can choose a safe part size before network work begins. - -The final upload byte count is compared with the declared `size`. A mismatch is -a caller/data-source error and the multipart upload is aborted. +When a stream size is unknown, the configured part size remains authoritative. The chunk iterator fails before part +10,001 instead of creating an invalid multipart request. A caller with a very large stream should supply `size` so the +planner can choose a safe part size before network work begins. -Failure and cancellation do not transfer cleanup authority to the caller. Once -`CreateMultipartUpload` succeeds, the high-level streamed `put()` owns that -upload until completion or abort. If a part fails or the caller cancels, the -client stops consuming the source, waits for admitted part requests to settle, -and then sends `AbortMultipartUpload` with a separate cleanup signal. -`abortTimeoutMs` bounds that best-effort cleanup and defaults to 30 seconds. -Using a separate signal matters because the caller's cancellation signal is -already aborted at exactly the time cleanup becomes necessary. +The final upload byte count is compared with the declared `size`. A mismatch is a caller/data-source error and the +multipart upload is aborted. +Failure and cancellation do not transfer cleanup authority to the caller. Once `CreateMultipartUpload` succeeds, the +high-level streamed `put()` owns that upload until completion or abort. If a part fails or the caller cancels, the +client stops consuming the source, waits for admitted part requests to settle, and then sends `AbortMultipartUpload` +with a separate cleanup signal. `abortTimeoutMs` bounds that best-effort cleanup and defaults to 30 seconds. Using a +separate signal matters because the caller's cancellation signal is already aborted at exactly the time cleanup becomes +necessary. -Multipart upload is a commit protocol -------------------------------------- +## Multipart upload is a commit protocol A streamed replacement follows four protocol stages: @@ -262,53 +268,42 @@ CompleteMultipartUpload `--> failure -> AbortMultipartUpload after active parts settle ``` -`@std/async/pool` owns the concurrency admission. The client does not start an -unbounded Promise for every part. +`@std/async/pool` owns the concurrency admission. The client does not start an unbounded Promise for every part. -Each successful `UploadPart` must return an ETag. The completion document -contains one ordered `PartNumber`/`ETag` record per uploaded part. Before the -client sends that XML, it verifies: +Each successful `UploadPart` must return an ETag. The completion document contains one ordered `PartNumber`/`ETag` +record per uploaded part. Before the client sends that XML, it verifies: - - at least one part exists; - - no more than 10,000 parts exist; - - every part number is an integer in the legal range; - - a part number is not duplicated; - - every ETag is non-empty; - - the final order is ascending by part number. +- at least one part exists; +- no more than 10,000 parts exist; +- every part number is an integer in the legal range; +- a part number is not duplicated; +- every ETag is non-empty; +- the final order is ascending by part number. -The XML document is built with `@std/xml/stringify`; protocol escaping is not a -hand-written string replacement. +The XML document is built with `@std/xml/stringify`; protocol escaping is not a hand-written string replacement. -The completion request can carry destination `If-Match`, `If-None-Match`, and -`x-amz-mp-object-size`. Preconditions belong to the commit stage rather than -multipart initiation because commit is the point where the destination object +The completion request can carry destination `If-Match`, `If-None-Match`, and `x-amz-mp-object-size`. Preconditions +belong to the commit stage rather than multipart initiation because commit is the point where the destination object becomes authoritative. -S3 has an unusual completion failure mode: the service can send HTTP `200 OK` -before final assembly finishes and then put an `` document in the -response body. `completeUpload()` therefore parses the success body and treats -an embedded `` as a failed commit. - -If any streamed part operation fails, the client waits for the pool's already -started requests to settle before it sends `AbortMultipartUpload`. This order -prevents a late `UploadPart` from racing after the abort request. +S3 has an unusual completion failure mode: the service can send HTTP `200 OK` before final assembly finishes and then +put an `` document in the response body. `completeUpload()` therefore parses the success body and treats an +embedded `` as a failed commit. -An abort failure does not replace the original upload failure. The original -operation remains the terminal error because it is what caused cleanup. -Unfinished multipart state can still remain at the provider when both the -operation and cleanup request fail, so production buckets should also use an S3 -lifecycle rule for stale multipart uploads. +If any streamed part operation fails, the client waits for the pool's already started requests to settle before it sends +`AbortMultipartUpload`. This order prevents a late `UploadPart` from racing after the abort request. +An abort failure does not replace the original upload failure. The original operation remains the terminal error because +it is what caused cleanup. Unfinished multipart state can still remain at the provider when both the operation and +cleanup request fail, so production buckets should also use an S3 lifecycle rule for stale multipart uploads. -Server-side copy has two distinct paths ---------------------------------------- +## Server-side copy has two distinct paths -`copy()` first performs `HeadObject` on the source. This confirms existence, -obtains size for planning, and supplies metadata needed by multipart initiation. +`copy()` first performs `HeadObject` on the source. This confirms existence, obtains size for planning, and supplies +metadata needed by multipart initiation. -For a source at or below the single-copy limit, the client sends `CopyObject`. -For a larger source, it creates a multipart upload at the destination and sends -one `UploadPartCopy` request per byte range. +For a source at or below the single-copy limit, the client sends `CopyObject`. For a larger source, it creates a +multipart upload at the destination and sends one `UploadPartCopy` request per byte range. ```text HEAD source @@ -325,73 +320,61 @@ HEAD source CompleteMultipartUpload ``` -Multipart copy chooses a range size large enough to keep the destination at or -below 10,000 parts and rejects a plan that would require a part above 5 GiB. -The range is inclusive because `x-amz-copy-source-range` uses inclusive byte +Multipart copy chooses a range size large enough to keep the destination at or below 10,000 parts and rejects a plan +that would require a part above 5 GiB. The range is inclusive because `x-amz-copy-source-range` uses inclusive byte positions. Source preconditions map to S3 copy-source headers: -| Package option | S3 header | -| -------------- | --------- | -| `sourceIfMatch` | `x-amz-copy-source-if-match` | -| `sourceIfNoneMatch` | `x-amz-copy-source-if-none-match` | -| `sourceIfModifiedSince` | `x-amz-copy-source-if-modified-since` | +| Package option | S3 header | +| ------------------------- | --------------------------------------- | +| `sourceIfMatch` | `x-amz-copy-source-if-match` | +| `sourceIfNoneMatch` | `x-amz-copy-source-if-none-match` | +| `sourceIfModifiedSince` | `x-amz-copy-source-if-modified-since` | | `sourceIfUnmodifiedSince` | `x-amz-copy-source-if-unmodified-since` | -Destination `ifMatch` and `ifNoneMatch` apply directly to `CopyObject` or to the -multipart completion request. They are deliberately removed from individual -`UploadPartCopy` requests because a destination object does not become the +Destination `ifMatch` and `ifNoneMatch` apply directly to `CopyObject` or to the multipart completion request. They are +deliberately removed from individual `UploadPartCopy` requests because a destination object does not become the completed value until commit. -`CopyObject` and `UploadPartCopy` can also encode a service error inside an HTTP -200 response. Both paths use the same success-XML inspection as multipart -completion. +`CopyObject` and `UploadPartCopy` can also encode a service error inside an HTTP 200 response. Both paths use the same +success-XML inspection as multipart completion. - -Listing preserves provider pagination -------------------------------------- +## Listing preserves provider pagination `list()` uses `ListObjectsV2` with these mappings: -| Package field | S3 query field | -| ------------- | -------------- | -| `prefix` | `prefix` | -| `delimiter` | `delimiter` | -| `limit` | `max-keys` | -| `cursor` | `continuation-token` | - -`Contents` records become `ObjectEntryType`. `CommonPrefixes` become child -prefixes. `NextContinuationToken` is returned as the next opaque cursor. +| Package field | S3 query field | +| ------------- | -------------------- | +| `prefix` | `prefix` | +| `delimiter` | `delimiter` | +| `limit` | `max-keys` | +| `cursor` | `continuation-token` | -The object-store filesystem adapter is responsible for consuming pages and -interpreting directory-marker metadata. The S3 client itself does not pretend -that prefixes are native directories. +`Contents` records become `ObjectEntryType`. `CommonPrefixes` become child prefixes. `NextContinuationToken` is returned +as the next opaque cursor. +The S3 object driver and object adapter are responsible for consuming pages and interpreting directory-marker metadata. +The S3 client itself does not pretend that prefixes are native directories. -Conditional writes are capability claims ------------------------------------------ +## Conditional writes are capability claims -The client advertises `conditionalWrite: true` by default because Amazon S3 -honors the preconditions used by this implementation. An S3-compatible service -that ignores or only partially implements these conditions must set +The client advertises `conditionalWrite: true` by default because Amazon S3 honors the preconditions used by this +implementation. An S3-compatible service that ignores or only partially implements these conditions must set `conditionalWrite: false` in `createS3Client()`. -That flag changes filesystem behavior. The object adapter will not claim that a -read-modify-write append/update is protected from concurrent replacement when -the backend cannot enforce the ETag precondition. - -`copy: false` similarly disables server-side copy for an endpoint whose S3 API -does not implement the required copy operations correctly. +That flag changes filesystem behavior. The S3 driver/object adapter will not claim that a read-modify-write +append/update is protected from concurrent replacement when the backend cannot enforce the ETag precondition. -These overrides are explicit because compatibility means "uses the S3 protocol" -not "implements every Amazon S3 behavior". +`copy: false` similarly disables server-side copy for an endpoint whose S3 API does not implement the required copy +operations correctly. +These overrides are explicit because compatibility means "uses the S3 protocol" not "implements every Amazon S3 +behavior". -Failures retain provider evidence ---------------------------------- +## Failures retain provider evidence -Non-success responses are parsed as S3 XML when possible. `S3Error` retains: +Non-success responses are parsed as S3 XML when possible. `S3Error` retains: ```text status @@ -401,114 +384,93 @@ hostId original Response ``` -The original `Response` remains available so a caller can inspect headers and -provider-specific diagnostics that the stable error fields do not model. +The original `Response` remains available so a caller can inspect headers and provider-specific diagnostics that the +stable error fields do not model. -A response body that is not valid S3 XML still produces a failure based on HTTP -status and available text. The client does not convert an unknown provider -response into a fake known S3 error code. +A response body that is not valid S3 XML still produces a failure based on HTTP status and available text. The client +does not convert an unknown provider response into a fake known S3 error code. -Cancellation uses `AbortSignal` on each Fetch request. Multipart cancellation -is not a distributed transaction: cancellation can stop local admission and -abort HTTP work, while an already accepted provider request may still have -created remote multipart state. Cleanup is therefore explicit. +Cancellation uses `AbortSignal` on each Fetch request. Multipart cancellation is not a distributed transaction: +cancellation can stop local admission and abort HTTP work, while an already accepted provider request may still have +created remote multipart state. Cleanup is therefore explicit. -Request retry is explicit and operation-aware. `request` in `S3ClientOptionsType` configures retries, exponential delay, jitter, -and an optional per-attempt timeout. The implementation uses `@std/async/retry` for 408, 429, 5xx, and transport failures. -Authorization is rebuilt for every attempt so refreshable credentials and SigV4 timestamps are current. Signed redirects are -manual and are returned to the caller instead of being followed to another authority. +Request retry is explicit and operation-aware. `request` in `S3ClientOptionsType` configures retries, exponential delay, +jitter, and an optional per-attempt timeout. The implementation uses `@std/async/retry` for 408, 429, 5xx, and transport +failures. Authorization is rebuilt for every attempt so refreshable credentials and SigV4 timestamps are current. Signed +redirects are manual and are returned to the caller instead of being followed to another authority. -Body replayability is only one admission condition. A one-shot `ReadableStream` receives one attempt. A mechanically replayable -body can still belong to a non-idempotent protocol operation, so low-level `request()` also accepts `retry: false`. The high-level -client disables automatic retry for `CreateMultipartUpload` and `CompleteMultipartUpload` because a lost response can make the -server-side outcome ambiguous. Stable part-number PUTs, reads, deletes, lists, and ordinary replacements use the configured -policy. `request: { retries: 0 }` disables automatic retry client-wide. +Body replayability is only one admission condition. A one-shot `ReadableStream` receives one attempt. A mechanically +replayable body can still belong to a non-idempotent protocol operation, so low-level `request()` also accepts +`retry: false`. The high-level client disables automatic retry for `CreateMultipartUpload` and `CompleteMultipartUpload` +because a lost response can make the server-side outcome ambiguous. Stable part-number PUTs, reads, deletes, lists, and +ordinary replacements use the configured policy. `request: { retries: 0 }` disables automatic retry client-wide. -`getMetrics()` returns direct HTTP request counts, retry counts, terminal failures, response counts, and optional Fetch duration. -Set `metrics: "none"` when measuring the raw protocol path, `basic` for counters, or `timing` for counters plus durations. +`getMetrics()` returns direct HTTP request counts, retry counts, terminal failures, response counts, and optional Fetch +duration. Set `metrics: "none"` when measuring the raw protocol path, `basic` for counters, or `timing` for counters +plus durations. +## Provider compatibility and known non-goals -Provider compatibility and known non-goals ------------------------------------------- - -The direct client supports custom endpoint, region, headers, path/virtual -addressing, copy capability, and conditional-write capability. This is enough -to use Amazon S3 and many S3-compatible products while keeping compatibility -choices visible. +The direct client supports custom endpoint, region, headers, path/virtual addressing, copy capability, and +conditional-write capability. This is enough to use Amazon S3 and many S3-compatible products while keeping +compatibility choices visible. The current client does **not** claim complete coverage of: - - SigV4 streaming chunk signatures; - - presigned URL creation; - - S3 Express directory-bucket session management; - - access points, Object Lambda, or Outposts host construction; - - Multi-Region Access Point SigV4A; - - SSE-C/SSE-KMS convenience APIs; - - checksum negotiation beyond the payload hash needed for signing; - - object tagging, ACLs, retention, legal hold, replication, or lifecycle APIs; - - version-ID aware filesystem paths; - - bucket creation or bucket policy management; - - adaptive throttling and provider-specific `Retry-After` scheduling beyond the shared exponential retry policy. - -A low-level signed request can still reach some provider features when the -caller knows the exact S3 REST contract. A feature should receive a typed -high-level API only after the library can document and test its semantics. +- SigV4 streaming chunk signatures; +- presigned URL creation; +- S3 Express directory-bucket session management; +- access points, Object Lambda, or Outposts host construction; +- Multi-Region Access Point SigV4A; +- SSE-C/SSE-KMS convenience APIs; +- checksum negotiation beyond the payload hash needed for signing; +- object tagging, ACLs, retention, legal hold, replication, or lifecycle APIs; +- version-ID aware filesystem paths; +- bucket creation or bucket policy management; +- adaptive throttling and provider-specific `Retry-After` scheduling beyond the shared exponential retry policy. +A low-level signed request can still reach some provider features when the caller knows the exact S3 REST contract. A +feature should receive a typed high-level API only after the library can document and test its semantics. -Validation strategy -------------------- +## Validation strategy S3 validation is intentionally split into protocol and provider tests. -`tests/s3.test.ts` is deterministic. It uses a controlled Fetch implementation -to inspect exact requests and covers: - - - Signature Version 4 canonicalization; - - deterministic timestamps and credentials; - - minimum/maximum multipart part configuration; - - multipart part ordering and duplicate rejection; - - `x-amz-mp-object-size` at completion; - - HTTP 200 embedded service errors; - - source and destination copy preconditions; - - large-copy multipart planning; - - part-count behavior. - -`tests/provider.test.ts` opens the pinned SeaweedFS S3-compatible server through -Testcontainers. It proves a real HTTP implementation can accept our signed -requests for PUT, HEAD, range GET, conditional create, multipart streaming -upload, copy, listing, delete, and the object-store filesystem adapter. The -fixture uses a random mapped host port so concurrent local runs do not share one -fixed endpoint. - -The provider container is SeaweedFS, not a statement that SeaweedFS defines the -S3 specification. The container proves interoperability with one independent -S3-compatible implementation. AWS-specific wire details remain covered by the +`tests/s3.test.ts` is deterministic. It uses a controlled Fetch implementation to inspect exact requests and covers: + +- Signature Version 4 canonicalization; +- deterministic timestamps and credentials; +- minimum/maximum multipart part configuration; +- multipart part ordering and duplicate rejection; +- `x-amz-mp-object-size` at completion; +- HTTP 200 embedded service errors; +- source and destination copy preconditions; +- large-copy multipart planning; +- part-count behavior. + +`tests/provider.test.ts` opens the pinned SeaweedFS S3-compatible server through Testcontainers. It proves a real HTTP +implementation can accept our signed requests for PUT, HEAD, range GET, conditional create, multipart streaming upload, +copy, listing, delete, the S3 driver, and the object filesystem adapter. The fixture uses a random mapped host port so +concurrent local runs do not share one fixed endpoint. + +The provider container is SeaweedFS, not a statement that SeaweedFS defines the S3 specification. The container proves +interoperability with one independent S3-compatible implementation. AWS-specific wire details remain covered by the deterministic protocol tests and the AWS documentation listed below. -Before release, the maintainer test matrix should also run an opt-in real Amazon -S3 suite with short-lived credentials when CI secret policy permits it. That -suite must use a dedicated disposable bucket/prefix and explicit cleanup. - - -Primary specification sources ------------------------------ - -The implementation and this guide should be checked against these primary AWS -sources when S3 behavior changes: - - - AWS Signature Version 4 canonical request: - https://docs.aws.amazon.com/AmazonS3/latest/API/sig-v4-header-based-auth.html - - Multipart upload limits: - https://docs.aws.amazon.com/AmazonS3/latest/userguide/qfacts.html - - CompleteMultipartUpload: - https://docs.aws.amazon.com/AmazonS3/latest/API/API_CompleteMultipartUpload.html - - CopyObject: - https://docs.aws.amazon.com/AmazonS3/latest/API/API_CopyObject.html - - UploadPartCopy: - https://docs.aws.amazon.com/AmazonS3/latest/API/API_UploadPartCopy.html - - ListObjectsV2: - https://docs.aws.amazon.com/AmazonS3/latest/API/API_ListObjectsV2.html - -Secondary S3-compatible provider documentation can explain provider-specific -configuration, but it must not override the Amazon S3 wire contract when the -client claims Amazon S3 behavior. +Before release, the maintainer test matrix should also run an opt-in real Amazon S3 suite with short-lived credentials +when CI secret policy permits it. That suite must use a dedicated disposable bucket/prefix and explicit cleanup. + +## Primary specification sources + +The implementation and this guide should be checked against these primary AWS sources when S3 behavior changes: + +- AWS Signature Version 4 canonical request: + https://docs.aws.amazon.com/AmazonS3/latest/API/sig-v4-header-based-auth.html +- Multipart upload limits: https://docs.aws.amazon.com/AmazonS3/latest/userguide/qfacts.html +- CompleteMultipartUpload: https://docs.aws.amazon.com/AmazonS3/latest/API/API_CompleteMultipartUpload.html +- CopyObject: https://docs.aws.amazon.com/AmazonS3/latest/API/API_CopyObject.html +- UploadPartCopy: https://docs.aws.amazon.com/AmazonS3/latest/API/API_UploadPartCopy.html +- ListObjectsV2: https://docs.aws.amazon.com/AmazonS3/latest/API/API_ListObjectsV2.html + +Secondary S3-compatible provider documentation can explain provider-specific configuration, but it must not override the +Amazon S3 wire contract when the client claims Amazon S3 behavior. diff --git a/docs/sources.md b/docs/sources.md index 528baa6..721e232 100644 --- a/docs/sources.md +++ b/docs/sources.md @@ -1,10 +1,10 @@ -Research and source register -============================ +# Research and source register Research date: 2026-08-15. -This register records the external contracts used to design and test the implementation. Source code and provider behavior can -change, so a release review should recheck current primary sources rather than assuming this date remains current. +This register records the external contracts used to design and test the implementation. Source code and provider +behavior can change, so a release review should recheck current primary sources rather than assuming this date remains +current. Use this authority order when sources disagree: @@ -13,11 +13,10 @@ Use this authority order when sources disagree: 3. current package implementation and tests; 4. older experiments and secondary performance reports. -The package intentionally distinguishes implemented behavior from provider claims and proposals. A compatibility note in this -file is not evidence that an adapter passed a live integration test against that provider. +The package intentionally distinguishes implemented behavior from provider claims and proposals. A compatibility note in +this file is not evidence that an adapter passed a live integration test against that provider. -Browser File System and OPFS ----------------------------- +## Browser File System and OPFS Primary sources: @@ -46,8 +45,7 @@ Secondary OPFS/performance context reviewed earlier in the project: These sources informed performance questions. They do not override the File System Standard or real browser tests. -Playwright ----------- +## Playwright Primary documentation: @@ -63,8 +61,7 @@ The canonical browser test architecture uses Playwright Test for Chromium, Firef ServiceWorker instrumentation because Playwright documents that inspection surface as Chromium-specific; observable ServiceWorker behavior remains a black-box test in the other browsers. -Testcontainers --------------- +## Testcontainers Primary documentation and source reviewed: @@ -76,18 +73,17 @@ Primary documentation and source reviewed: - Toxiproxy module: - repository: -The provider suite keeps `node:test` as the test runner and uses Testcontainers only for service lifecycle. The official Azurite -module is used instead of reproducing its container flags. SeaweedFS has no focused Testcontainers module, so the fixture uses -`GenericContainer` with the pinned image and a composite listening-port/HTTP wait. Random mapped host ports avoid collisions -between concurrent local runs. +The provider suite keeps `node:test` as the test runner and uses Testcontainers only for service lifecycle. The official +Azurite module is used instead of reproducing its container flags. SeaweedFS has no focused Testcontainers module, so +the fixture uses `GenericContainer` with the pinned image and a composite listening-port/HTTP wait. Random mapped host +ports avoid collisions between concurrent local runs. -Testcontainers currently documents Docker directly and Docker-compatible configuration for Podman, Colima, and Rancher Desktop. -Those runtimes have provider-specific caveats, including Ryuk behavior under Podman and delayed port forwarding under -Colima/Rancher. The package therefore treats Testcontainers as an interim test-resource implementation, not as the future -compute-provider API for OPFS. +Testcontainers currently documents Docker directly and Docker-compatible configuration for Podman, Colima, and Rancher +Desktop. Those runtimes have provider-specific caveats, including Ryuk behavior under Podman and delayed port forwarding +under Colima/Rancher. The package therefore treats Testcontainers as an interim test-resource implementation, not as the +future compute-provider API for OPFS. -Deno, Node, and Bun -------------------- +## Deno, Node, and Bun Primary runtime sources: @@ -99,111 +95,92 @@ Primary runtime sources: - Bun benchmarking guidance: - Node documentation: -Current Deno documentation treats `node:test` as a first-class test API and currently marks Deno KV unstable. The real Deno KV -suite therefore uses `--unstable-kv` without making that flag part of unrelated source imports. Current Deno KV documentation -states a 2 KiB serialized key limit, a 64 KiB serialized value limit, 1,000 mutations per atomic operation, and an 800 KiB total -atomic-operation limit. The Deno KV adapter exposes these constraints and uses a configurable manifest/part layout instead of -pretending one logical file must fit in one 64 KiB value. - -Current Bun compatibility documentation says its in-process `node:test` API works when files run under `bun test`, while some -advanced Node test-runner/reporting features remain incomplete. The repository uses the common `describe`/`it`/hooks subset and -keeps the test API itself as `node:test`. - -Bun's current File I/O documentation says `Bun.write(destination, Bun.file(source))` selects fast platform system calls for -file-to-file copies. The current Bun Rust source also keeps file-backed Blob state distinct so file-to-file paths can avoid a -naive user-space read/write loop. The benchmark therefore compares Bun's direct copy shape with Node-compatible `copyFile` -before changing the adapter implementation. - -Bun's S3 documentation exposes `S3Client`, `S3File`, `write`, `stat`, `stream`, and multipart `writer()` APIs. A Bun-only provider -benchmark now uses that native implementation as a second S3 baseline beside AWS SDK v3. The project does not treat Bun main -branch implementation work as proof about a released runtime; the mise pin remains the current released version selected by the -repository until a deliberate toolchain update. - -Deno standard libraries and Standard Schema -------------------------------------------- - -Primary sources reviewed from the current `denoland/std` repository and JSR -packages: - - - `@std/async`: - - `@std/bytes`: - - `@std/encoding`: - - `@std/expect`: - - `@std/fs`: - - `@std/http`: - - `@std/path`: - - `@std/streams`: - - `@std/xml`: - - `@std/crypto`: - - Standard Schema: - - Zod 4: - -The review was operation-led. A standard package replaces project code only -when its contract matches the filesystem or provider requirement without hiding -a stronger invariant. - -`@std/async/pool` owns bounded multipart and block concurrency. The stable -`pooledMap()` contract limits active requests and lets already-started requests -settle after one item fails. S3 cleanup waits for that settlement before it -sends `AbortMultipartUpload`, so a late part cannot arrive after the cleanup -request. - -`@std/async/retry` owns the direct clients' exponential backoff, jitter, AbortSignal, and retriable-error loop. The protocol -layer still classifies whether a request may enter that loop. One-shot streams are not replayed, and S3 multipart initiation and -completion disable automatic retry because a lost response can make the remote lifecycle outcome ambiguous. The low-level -request APIs also expose `retry: false` for provider-specific operations. - -`@std/bytes/concat` owns byte-array concatenation used by bounded chunk -assembly. The package does not maintain another concatenation implementation. - -`@std/streams` owns bounded materialization through -`LimitedBytesTransformStream` and final stream collection through `toBytes()`. -The current `FixedChunkStream` API is still marked unstable, so fixed-size -provider chunks remain in the package's small streaming adapter until that -standard contract is suitable for a public dependency. - -`@std/encoding` owns Base64 and hexadecimal encoding through their direct -subpaths. S3 uses hexadecimal SHA-256 output, Azure Shared Key and block IDs -use Base64, and record stores use Base64 for portable byte persistence. - -`@std/path` owns host path normalization and resolution for the Deno, Bun, and -Node adapters. The OPFS virtual path model remains project-owned because it -rejects and normalizes a different namespace than an operating-system path. - -`@std/fs` was reviewed for copy, move, walk, ensure, and host filesystem -operations. Those are intentionally not used inside the primitive Deno/Bun/Node -adapters. The public filesystem facade already owns recursive copy/move/walk, -overwrite, cancellation, and adapter-neutral semantics. Calling `@std/fs` from -one host adapter would duplicate that layer and introduce host-only symlink and -filesystem assumptions. `@std/path`, by contrast, directly replaces custom host -path manipulation without changing facade semantics. - -`@std/http/etag` was reviewed for conditional request support. The clients keep -provider ETags opaque instead of generating or evaluating them locally. S3 -multipart ETags and Azure ETags are provider tokens, not hashes that this -library should reinterpret. The package therefore forwards `If-Match` and -`If-None-Match` values to the provider rather than applying `@std/http/etag` in -the client. The unstable HTTP message-signature utilities also do not implement -AWS Signature Version 4 or Azure Shared Key. - -`@std/xml` owns provider control-document parsing and serialization. S3 list, -error, multipart, and copy responses and Azure list/error/block-list documents -use the standard XML tree instead of regular expressions or hand-written XML +Current Deno documentation treats `node:test` as a first-class test API and currently marks Deno KV unstable. The real +Deno KV suite therefore uses `--unstable-kv` without making that flag part of unrelated source imports. Current Deno KV +documentation states a 2 KiB serialized key limit, a 64 KiB serialized value limit, 1,000 mutations per atomic +operation, and an 800 KiB total atomic-operation limit. The Deno KV adapter exposes these constraints and uses a +configurable manifest/part layout instead of pretending one logical file must fit in one 64 KiB value. + +Current Bun compatibility documentation says its in-process `node:test` API works when files run under `bun test`, while +some advanced Node test-runner/reporting features remain incomplete. The repository uses the common +`describe`/`it`/hooks subset and keeps the test API itself as `node:test`. + +Bun's current File I/O documentation says `Bun.write(destination, Bun.file(source))` selects fast platform system calls +for file-to-file copies. The current Bun Rust source also keeps file-backed Blob state distinct so file-to-file paths +can avoid a naive user-space read/write loop. The benchmark therefore compares Bun's direct copy shape with +Node-compatible `copyFile` before changing the adapter implementation. + +Bun's S3 documentation exposes `S3Client`, `S3File`, `write`, `stat`, `stream`, and multipart `writer()` APIs. A +Bun-only provider benchmark now uses that native implementation as a second S3 baseline beside AWS SDK v3. The project +does not treat Bun main branch implementation work as proof about a released runtime; the mise pin remains the current +released version selected by the repository until a deliberate toolchain update. + +## Deno standard libraries and Standard Schema + +Primary sources reviewed from the current `denoland/std` repository and JSR packages: + +- `@std/async`: +- `@std/bytes`: +- `@std/encoding`: +- `@std/expect`: +- `@std/fs`: +- `@std/http`: +- `@std/path`: +- `@std/streams`: +- `@std/xml`: +- `@std/crypto`: +- Standard Schema: +- Zod 4: + +The review was operation-led. A standard package replaces project code only when its contract matches the filesystem or +provider requirement without hiding a stronger invariant. + +`@std/async/pool` owns bounded multipart and block concurrency. The stable `pooledMap()` contract limits active requests +and lets already-started requests settle after one item fails. S3 cleanup waits for that settlement before it sends +`AbortMultipartUpload`, so a late part cannot arrive after the cleanup request. + +`@std/async/retry` owns the direct clients' exponential backoff, jitter, AbortSignal, and retriable-error loop. The +protocol layer still classifies whether a request may enter that loop. One-shot streams are not replayed, and S3 +multipart initiation and completion disable automatic retry because a lost response can make the remote lifecycle +outcome ambiguous. The low-level request APIs also expose `retry: false` for provider-specific operations. + +`@std/bytes/concat` owns byte-array concatenation used by bounded chunk assembly. The package does not maintain another +concatenation implementation. + +`@std/streams` owns bounded materialization through `LimitedBytesTransformStream` and final stream collection through +`toBytes()`. The current `FixedChunkStream` API is still marked unstable, so fixed-size provider chunks remain in the +package's small streaming adapter until that standard contract is suitable for a public dependency. + +`@std/encoding` owns Base64 and hexadecimal encoding through their direct subpaths. S3 uses hexadecimal SHA-256 output, +Azure Shared Key and block IDs use Base64, and record stores use Base64 for portable byte persistence. + +`@std/path` owns host path normalization and resolution for the Deno, Bun, and Node adapters. The OPFS virtual path +model remains project-owned because it rejects and normalizes a different namespace than an operating-system path. + +`@std/fs` was reviewed for copy, move, walk, ensure, and host filesystem operations. Those are intentionally not used +inside the primitive Deno/Bun/Node adapters. The public filesystem facade already owns recursive copy/move/walk, +overwrite, cancellation, and adapter-neutral semantics. Calling `@std/fs` from one host adapter would duplicate that +layer and introduce host-only symlink and filesystem assumptions. `@std/path`, by contrast, directly replaces custom +host path manipulation without changing facade semantics. + +`@std/http/etag` was reviewed for conditional request support. The clients keep provider ETags opaque instead of +generating or evaluating them locally. S3 multipart ETags and Azure ETags are provider tokens, not hashes that this +library should reinterpret. The package therefore forwards `If-Match` and `If-None-Match` values to the provider rather +than applying `@std/http/etag` in the client. The unstable HTTP message-signature utilities also do not implement AWS +Signature Version 4 or Azure Shared Key. + +`@std/xml` owns provider control-document parsing and serialization. S3 list, error, multipart, and copy responses and +Azure list/error/block-list documents use the standard XML tree instead of regular expressions or hand-written XML escaping. Storage payloads themselves do not pass through XML parsing. -`@std/crypto` was reviewed but is not used for provider signing. Web Crypto -already exposes browser-compatible SHA-256 and HMAC-SHA256, while the standard -crypto package does not implement AWS Signature Version 4 or Azure Shared Key -canonicalization. Adding it would introduce a wrapper without removing the -protocol code that actually carries the risk. +`@std/crypto` was reviewed but is not used for provider signing. Web Crypto already exposes browser-compatible SHA-256 +and HMAC-SHA256, while the standard crypto package does not implement AWS Signature Version 4 or Azure Shared Key +canonicalization. Adding it would introduce a wrapper without removing the protocol code that actually carries the risk. -`@std/expect` remains the assertion API on top of `node:test`. Zod 4 implements -Standard Schema, so the repository exports the Zod schemas directly instead of -maintaining a second validation wrapper for Standard Schema consumers. +`@std/expect` remains the assertion API on top of `node:test`. Zod 4 implements Standard Schema, so the repository +exports the Zod schemas directly instead of maintaining a second validation wrapper for Standard Schema consumers. - -S3 and Signature Version 4 --------------------------- +## S3 and Signature Version 4 AWS primary references: @@ -222,7 +199,8 @@ AWS primary references: Implementation details derived from these contracts include: -- canonical signing includes `host` even though browser Fetch does not let application code set the Host header directly; +- canonical signing includes `host` even though browser Fetch does not let application code set the Host header + directly; - multipart upload parts are bounded and the destination publishes on CompleteMultipartUpload; - conditional `If-Match`/`If-None-Match` behavior belongs to multipart completion for the commit path used here; - CompleteMultipartUpload can return an HTTP 200 response whose XML body later reports an error; @@ -230,11 +208,10 @@ Implementation details derived from these contracts include: - CopyObject has a 5 GB source limit, so larger provider-side copies use UploadPartCopy; - multipart uploads permit at most 10,000 parts and have defined part-size limits. -The package uses Web Crypto and Web Fetch rather than the AWS SDK so the direct client remains small, runtime-neutral, and -explicit about the S3 protocol surface it actually implements. +The package uses Web Crypto and Web Fetch rather than the AWS SDK so the direct client remains small, runtime-neutral, +and explicit about the S3 protocol surface it actually implements. -S3-compatible providers ------------------------ +## S3-compatible providers Provider-specific primary sources reviewed for compatibility differences: @@ -258,11 +235,11 @@ Backblaze B2 S3-compatible API: - S3-compatible API: -These providers illustrate why capability overrides exist. Endpoint, region, addressing, copy support, multipart preconditions, -checksum behavior, and unsupported control-plane operations can differ even when basic object requests use the S3 protocol. +These providers illustrate why capability overrides exist. Endpoint, region, addressing, copy support, multipart +preconditions, checksum behavior, and unsupported control-plane operations can differ even when basic object requests +use the S3 protocol. -Azure Blob Storage ------------------- +## Azure Blob Storage Microsoft primary references: @@ -274,16 +251,16 @@ Microsoft primary references: - Copy Blob From URL: - Put Block From URL: - List Blobs: -- Versioning for Azure Storage services: +- Versioning for Azure Storage services: + -The implementation keeps the service version explicit because accepted block sizes and Shared Key canonicalization depend on -the service version. Shared Key support starts at the augmented Blob format introduced in `2009-09-19`; zero-length -`Content-Length` signing changes after `2014-02-14`, and empty `x-ms-*` header canonicalization changes at `2016-05-31`. -Current copy behavior uses synchronous Copy Blob From URL for the smaller path and Put Block From URL ranges for large -provider-side copies. +The implementation keeps the service version explicit because accepted block sizes and Shared Key canonicalization +depend on the service version. Shared Key support starts at the augmented Blob format introduced in `2009-09-19`; +zero-length `Content-Length` signing changes after `2014-02-14`, and empty `x-ms-*` header canonicalization changes at +`2016-05-31`. Current copy behavior uses synchronous Copy Blob From URL for the smaller path and Put Block From URL +ranges for large provider-side copies. -Unstorage ---------- +## Unstorage Primary sources: @@ -293,12 +270,11 @@ Primary sources: - custom drivers: - built-in driver catalog: -The forward bridge targets `Storage`, not individual unstorage drivers. The reverse driver implements the stable Driver subset -needed for values, raw bytes, metadata, keys, clear, and disposal. `maxDepth` is advertised because the reverse driver applies -the depth filter itself. +The forward driver targets unstorage `Storage`, not individual unstorage backend implementations. The reverse bridge +implements the stable unstorage `Driver` subset needed for values, raw bytes, metadata, keys, clear, and disposal. +`maxDepth` is advertised because the bridge applies the depth filter itself. -RxDB ----- +## RxDB Primary sources: @@ -306,11 +282,10 @@ Primary sources: - RxStorage interface: - RxCollection implementation: -The bridge accepts an RxCollection. RxDB retains responsibility for the selected RxStorage, replication, conflicts, +The driver accepts an RxCollection. RxDB retains responsibility for the selected RxStorage, replication, conflicts, multi-instance behavior, wrappers, and licensing. -db0 and Drizzle ---------------- +## db0 and Drizzle Primary sources: @@ -319,35 +294,38 @@ Primary sources: - Drizzle ORM: - Drizzle repository: -The db0 bridge targets the Database/dialect contract rather than connector names. Direct SQLite reuses that same record schema. -Drizzle keeps table/DDL ownership with the application because its schema builders and database behavior are dialect-specific. +The db0 record driver targets the Database/dialect contract rather than connector names. Direct SQLite reuses that same +record schema. The Drizzle record driver keeps table/DDL ownership with the application because its schema builders and +database behavior are dialect-specific. -Upstream issue and pull-request review --------------------------------------- +## Upstream issue and pull-request review -Current upstream issue/PR review was used to find failure modes that happy-path API docs do not reveal. The implementation does -not copy another library's behavior blindly; the issues are evidence for tests and invariants. +Current upstream issue/PR review was used to find failure modes that happy-path API docs do not reveal. The +implementation does not copy another library's behavior blindly; the issues are evidence for tests and invariants. -Bun S3/Rust work reviewed included fixes for retry coverage, exponential backoff, timeouts, manual redirect handling, option -propagation, multipart abort on writer error, long SigV4 inputs, in-place multipart part assembly, XML parsing, proxy handling, -and worker-termination lifetime safety. The repeated lessons are: signed redirects must not be followed automatically, remote -cleanup has its own lifecycle, part concurrency needs a memory budget, and retry policy must not be inferred from body type alone. +Bun S3/Rust work reviewed included fixes for retry coverage, exponential backoff, timeouts, manual redirect handling, +option propagation, multipart abort on writer error, long SigV4 inputs, in-place multipart part assembly, XML parsing, +proxy handling, and worker-termination lifetime safety. The repeated lessons are: signed redirects must not be followed +automatically, remote cleanup has its own lifecycle, part concurrency needs a memory budget, and retry policy must not +be inferred from body type alone. -AWS SDK v3 issues reviewed included very large upload memory growth, unknown-size multipart completion hangs, empty-stream lockups, -stream chunk-integrity regressions, conditional-header gaps in `lib-storage`, browser decompression/checksum mismatches, socket -exhaustion, and S3-compatible provider deserialization/endpoint regressions. The project benchmark keeps the AWS SDK as a -baseline while retaining a smaller direct protocol client with independently testable semantics. +AWS SDK v3 issues reviewed included very large upload memory growth, unknown-size multipart completion hangs, +empty-stream lockups, stream chunk-integrity regressions, conditional-header gaps in `lib-storage`, browser +decompression/checksum mismatches, socket exhaustion, and S3-compatible provider deserialization/endpoint regressions. +The project benchmark keeps the AWS SDK as a baseline while retaining a smaller direct protocol client with +independently testable semantics. -Azure SDK issues reviewed included paused-stream abort hangs, invalid upload buffer arguments producing zero-byte blobs, large -buffer/block-size constraints, historical stream/file data corruption, copy polling request noise, and concurrency/default-size -questions. These reinforce explicit size/concurrency limits, bounded block admission, real abort tests, and provider request-count -benchmarks. +Azure SDK issues reviewed included paused-stream abort hangs, invalid upload buffer arguments producing zero-byte blobs, +large buffer/block-size constraints, historical stream/file data corruption, copy polling request noise, and +concurrency/default-size questions. These reinforce explicit size/concurrency limits, bounded block admission, real +abort tests, and provider request-count benchmarks. -Unstorage issues reviewed included non-atomic filesystem writes, S3 pagination/prefix bugs, XML entity decoding, file/prefix -collisions, SQL disposal, binary Redis storage, and Cloudflare Cache method binding. RxDB issues reviewed included OPFS/Expo file -truncation after crashes or rapid writes, large-replication corruption, and concurrency/benchmark questions. db0 issues reviewed -included connector/dialect exposure, caller-owned connections, deprecated sqlite3, and Drizzle result-shape mismatches. These are -why the OPFS project keeps ownership, collision semantics, partial-result failure, and backend capability differences explicit. +Unstorage issues reviewed included non-atomic filesystem writes, S3 pagination/prefix bugs, XML entity decoding, +file/prefix collisions, SQL disposal, binary Redis storage, and Cloudflare Cache method binding. RxDB issues reviewed +included OPFS/Expo file truncation after crashes or rapid writes, large-replication corruption, and +concurrency/benchmark questions. db0 issues reviewed included connector/dialect exposure, caller-owned connections, +deprecated sqlite3, and Drizzle result-shape mismatches. These are why the OPFS project keeps ownership, collision +semantics, partial-result failure, and backend capability differences explicit. Deno KV issue review also covered historical reports about large prefix-list cost and selector/transaction limits: @@ -355,10 +333,11 @@ Deno KV issue review also covered historical reports about large prefix-list cos - - -The Deno KV physical key layout therefore indexes a logical entry by `(namespace, "entry", parentPath, name)`. Listing one -directory uses `(namespace, "entry", parentPath)` as the provider prefix, so descendants of a child directory are not part of -that prefix result. Physical body parts use the complete canonical path as one tuple component rather than expanding each path -segment into the provider prefix. This keeps exact lookup and direct-child enumeration aligned with the filesystem contract. +The Deno KV physical key layout therefore indexes a logical entry by `(namespace, "entry", parentPath, name)`. Listing +one directory uses `(namespace, "entry", parentPath)` as the provider prefix, so descendants of a child directory are +not part of that prefix result. Physical body parts use the complete canonical path as one tuple component rather than +expanding each path segment into the provider prefix. This keeps exact lookup and direct-child enumeration aligned with +the filesystem contract. Recent Drizzle issue review included SQLite/libSQL transaction-lifetime failures and migration/data-loss cases: @@ -368,12 +347,12 @@ Recent Drizzle issue review included SQLite/libSQL transaction-lifetime failures - - -These are not all adapter-runtime bugs, but they reinforce a deliberate contract here: the generic Drizzle bridge does not -claim universal cross-process atomic replacement or own application migrations. The caller keeps dialect/driver/table lifecycle -and can provide a stronger database-specific transaction strategy when that concrete driver proves the required semantics. +These are not all driver-runtime bugs, but they reinforce a deliberate contract here: the generic Drizzle record driver +does not claim universal cross-process atomic replacement or own application migrations. The caller keeps +dialect/driver/table lifecycle and can provide a stronger database-specific transaction strategy when that concrete +driver proves the required semantics. -Project architecture and writing sources ----------------------------------------- +## Project architecture and writing sources The implementation was reviewed against the attached/current project guides covering: @@ -386,13 +365,49 @@ The implementation was reviewed against the attached/current project guides cove - runtime-neutral TypeScript and explicit runtime subpaths; - verification against real runtimes and extracted release artifacts. -Older OPFS experiments were treated as intent/history only. The current repository, current project rules, and current upstream -contracts are the implementation authority for this pass. +Older OPFS experiments were treated as intent/history only. The current repository, current project rules, and current +upstream contracts are the implementation authority for this pass. + +## Client protocol handoffs + +The detailed implementation contracts live in [s3.md](./s3.md) and [azure.md](./azure.md). The Testcontainers-backed +interoperability matrix is documented in [providers.md](./providers.md). These files separate protocol requirements from +emulator evidence and record the unsupported surface explicitly. + +## Filesystem-client baselines + +AWS Mountpoint primary sources reviewed on 2026-08-15: + +- Amazon S3 Mountpoint overview: +- Mountpoint usage: +- Mountpoint configuration source: + +Mountpoint is a high-throughput S3 filesystem client, not a full POSIX filesystem. AWS documents that it can list/read +existing objects and create new files, while operations such as modifying existing files, symbolic links, and file +locking are not general Mountpoint capabilities. The benchmark therefore compares only operations whose semantics are +close enough to answer the same question. `--endpoint-url` and `--force-path-style` can support controlled endpoint +experiments, but an alternate endpoint is not an AWS compatibility guarantee. + +Azure BlobFuse primary sources reviewed on 2026-08-15: + +- BlobFuse repository: +- Microsoft limitations/known issues: + +BlobFuse translates Linux FUSE operations to Azure Blob requests. Its cache modes and unsupported/altered filesystem +operations are part of benchmark semantics. In particular, cache configuration can change freshness and request count, +while operations such as file locking and several extended/POSIX operations are not supported. Benchmark output must +therefore record cache/write mode rather than treating every mounted file operation as equivalent to direct REST access. + +## Mise and registry publishing +Primary tool/release sources reviewed on 2026-08-15: -Client protocol handoffs ------------------------- +- mise npm backend: +- mise settings: +- npm trusted publishing: +- npm provenance: -The detailed implementation contracts live in [s3.md](./s3.md) and [azure.md](./azure.md). The Testcontainers-backed interoperability -matrix is documented in [providers.md](./providers.md). These files separate protocol requirements from emulator evidence and -record the unsupported surface explicitly. +Mise can install npm-distributed CLI tools directly through the `npm:` backend. The release configuration therefore pins +the npm CLI through mise instead of depending on the npm version bundled with a selected Node release. npm's current +trusted-publishing requirements call for npm CLI 11.5.1 or later and Node 22.14.0 or later; the repository pins npm +11.18.0 for the publish job. diff --git a/docs/validation.md b/docs/validation.md index 52aa57c..564d061 100644 --- a/docs/validation.md +++ b/docs/validation.md @@ -1,330 +1,419 @@ -Validation strategy -=================== +# Validation strategy -The test architecture separates portable filesystem semantics from the runtimes and providers that supply concrete storage. -This is deliberate. A fast memory test should not be the evidence for browser OPFS interoperability, and a browser test should -not be the only evidence for a deterministic path or copy invariant. +## Purpose -The canonical layers are: +Validation follows the storage layers. A memory test is not evidence for browser OPFS interoperability. A fake HTTP test +is not evidence that an S3-compatible server accepts the request. A facade benchmark is not enough to identify whether +overhead came from the protocol client, driver, adapter, metrics, or provider. + +The canonical model is: ```text node:test + @std/expect - portable schemas, paths, facade behavior, record/object translations, - ecosystem bridges, S3/Azure protocol behavior with deterministic fakes - -real server runtimes - Deno host filesystem - Deno KV - Node host filesystem + node:sqlite - Bun host filesystem + schemas / paths / driver contracts / adapter translation / facade semantics + deterministic S3/Azure protocol tests + record/object/database contract tests -Testcontainers + node:test - SeaweedFS S3 compatibility / Azurite Blob interoperability - random host ports / readiness / owned cleanup +Deno / Node / Bun runtime tests + actual host filesystem behavior + Deno KV runtime behavior Playwright Test Chromium / Firefox / WebKit - Window / Worker / ServiceWorker / iframe / persistence / browser storage + Window / Worker / iframe / ServiceWorker / persistence -Mitata - raw backend baseline -> adapter primitive -> filesystem facade +Testcontainers + node:test + disposable SeaweedFS and Azurite provider interoperability + +Mitata + Playwright benchmarks + native/client -> driver -> adapter -> facade -> facade+metrics ``` -A test states the contract it protects. Avoid tests that only mirror the current implementation line by line. - -Portable tests protect filesystem semantics -------------------------------------------- - -The portable suite uses `node:test` with `describe` and `it`, plus `@std/expect` for expectations. Deno runs these same source -files directly. Node runs the same source. Bun runs the same `node:test` API through `bun test`. - -The suite covers: - -- schema acceptance/rejection and Standard Schema exposure; -- canonical path normalization and root-escape rejection; -- file and directory handle semantics; -- replace, append, update, truncate, and byte-range behavior; -- staged writable close versus abort; -- stream cancellation after the operation becomes terminal; -- bounded stream materialization for simple record adapters; -- Deno KV partitioned large-file stat/list/range/stream behavior, bounded append/update patching, and manifest-last visibility; -- optimization-disabled differential paths and matching preflight plans; -- filesystem route/peak-buffer metrics; -- copy/move overwrite and source/destination overlap protection; -- file mutation versus structural mutation coordination; -- queued cancellation recovery; -- sync-file lock lifetime and partial-write looping; -- adapter disposal ownership; -- record-store semantics; -- generic object-store directories, ranges, streaming replacement, optimistic read-modify-write, and native copy; -- foreign object layouts where an exact file key and a child prefix coexist; -- unstorage forward and reverse integration; -- the generic reverse key-value driver and collision-safe keys; -- RxDB, db0 dialect, Drizzle, and direct SQLite translation; -- S3 Signature Version 4, XML list/error parsing, multipart commit preconditions, HTTP-200 embedded failures, multipart - server-side copy, retry/backoff, timeout, manual redirects, one-shot body admission, and non-idempotent multipart lifecycle retry guards; -- Azure list/error parsing, large server-side range copy, bearer/SAS source authorization, provider request identities, - retry/backoff, and one-shot/explicit no-retry behavior. - -Focused commands: +## Repository command authority + +Mise owns tool versions and repository commands. ```sh -deno task test:portable -deno task test:node -deno task test:bun +mise install +mise run check +mise run test +mise run test-node +mise run test-deno +mise run test-bun +mise run test-browser +mise run test-providers +mise run bench +mise run bench-providers +mise run bench-browser +mise run bench-filesystem-clients +mise run quality ``` -The deterministic stress run shuffles and repeats the portable suite so hidden test order does not become a dependency: +GitHub Actions owns only GitHub-specific orchestration: triggers, permissions, matrices, secrets, outputs, immutable +release refs, and calls into those mise tasks. -```sh -deno task test:stress -``` +## Portable tests -Runtime suites prove runtime adapters against the real API ----------------------------------------------------------- +Portable tests use `node:test` with `describe`/`it` and `@std/expect`. Deno and Node consume the same TypeScript source. +Bun uses the same test contracts where its runner/runtime supports them. -`tests/deno.test.ts` exercises the real Deno host filesystem adapter. `tests/node.test.ts` exercises real Node host filesystem -operations and runs the SQL bridge against Node's built-in SQLite engine. `tests/bun.test.ts` exercises the Bun adapter against -Bun's actual runtime. +Important portable suites: -Deno KV has a separate real integration test because current Deno requires the unstable KV flag: +```text +tests/path.test.ts + canonical path parsing / root escape / names -```sh -deno task test:deno-kv -``` +tests/driver.test.ts + driver definition validation + requirement/limit provenance + behavior-changing optimization disableability + direct third-party driver planning -The adapter module is still type-checked with the server set. The unstable flag belongs to the real Deno KV execution, not to -unrelated package imports. +tests/memory.test.ts + deterministic record driver/adapter/facade behavior -The normal server-runtime matrix is: +tests/filesystem.test.ts + locks / staged writes / copy / move / cancellation / lifecycle -```sh -deno task test:deno -deno task test:deno-kv -deno task test:node -deno task test:bun -``` +tests/ecosystems.test.ts + unstorage / RxDB / db0 / Drizzle + integration direction metadata + real reverse unstorage bridge -The pinned mise task runs these after installing the frozen dependency graph: +tests/object.test.ts + generic object driver -> object adapter -> facade contract -```sh -mise run test +tests/s3.test.ts + deterministic S3 REST/SigV4/multipart/copy/retry behavior + +tests/azure.test.ts + deterministic Azure REST/auth/block/copy/retry behavior + +tests/deno-kv-partition.test.ts + partition layout using an in-memory Deno KV contract double + +tests/sqlite.test.ts + direct SQLite row-driver behavior ``` -GitHub Actions does not recreate this runtime setup with separate Node, Deno, and Bun setup actions. `jdx/mise-action` installs -the pinned mise release and only the tools required by the current job. The job then calls the same focused mise task a -maintainer can run locally, such as `mise run test-deno`, `mise run test-node`, or `mise run test-bun`. This keeps tool versions -and test commands in the repository instead of duplicating them in workflow YAML. +A test should identify the contract it protects. Avoid tests that merely restate private implementation steps. -Testcontainers owns provider-service lifecycle ---------------------------------------------- +## Driver tests -`tests/provider.test.ts` remains a `node:test` suite. `tests/provider/fixture.ts` uses Testcontainers only to supply real local -services. SeaweedFS runs through `GenericContainer`; Azurite uses the official `@testcontainers/azurite` module. Testcontainers -selects free host ports, applies readiness checks, and owns container cleanup. No Docker Compose subprocess or project polling -loop is required. +A driver is a public extension seam and therefore receives direct tests before an adapter exists. -```sh -mise run test-providers -``` +The generic suite proves: -The provider job in GitHub Actions installs Deno and Node through mise and calls that same task. Docker-compatible runtime -selection remains Testcontainers configuration, not package runtime logic. This is an interim fixture layer and does not constrain -a future provider abstraction to the Docker API. +- structured requirements are retained; +- provider, implementation, user, and probe limits keep their provenance; +- an optimization with `changesBehavior: true` cannot declare `disableable: false`; +- driver planning can reject a known impossible input without storage I/O; +- structured problems/actions are stable machine data. -Playwright owns browser installation and browser lifecycle ---------------------------------------------------------- +Backend-specific driver tests then protect physical rules. Deno KV is the most important reference because byte size, +path/key size, partition policy, and provider ceilings all affect admission. -The browser suite lives under `tests/browser/`. There is no custom browser-launch loop or custom test-result protocol. -Playwright owns browser installation, contexts, server lifecycle, traces, retries, and test attribution. +## Deno KV validation -Install the compatible browser builds, then run the matrix: +The portable Deno KV partition suite uses a deterministic contract double. It proves: -```sh -deno task test:browser:install -deno task test:browser -``` +- no stored test value crosses the documented serialized value ceiling enforced by the double; +- conservative `partBytes`/`inlineBytes` safety budgets reject unsafe configuration; +- a long physical key is rejected during driver preflight before provider I/O; +- `partition: "never"` returns structured `change-policy`/`select-driver` actions for oversized input; +- large logical files reconstruct exactly; +- directory listing does not load file body parts; +- range reads touch only overlapping parts; +- streamed replacement uses bounded partition writes without facade buffering; +- disabling the facade stream-write optimization forces the bounded facade fallback; +- append/update preserve untouched bytes; +- manifest-last replacement never publishes a partial new generation. -or: +The Deno-native suite uses the real Deno KV API when the runtime is available. It remains the release evidence for +actual serialization/provider behavior; the contract double does not replace it. -```sh -mise run test-browser -``` +## Host filesystem runtime tests -The same semantic tests run in Chromium, Firefox, and WebKit. Tests probe runtime capability and then assert the actual result. -They do not encode statements such as "Firefox has no sync OPFS" or "WebKit always rejects this iframe" into the test logic. -Those are exactly the assumptions an interoperability suite is supposed to detect when browser behavior changes. +Node, Deno, and Bun each run real host-file tests because a shared structural contract cannot prove runtime I/O +behavior. -The browser cases include: +The runtime suites cover the applicable routes: ```text -Window - OPFS probe - async write/read - abort before commit - -DedicatedWorker - async OPFS - synchronous handle probe and open attempt +replace / append / update +ranges +native streams +native copy +native move +asynchronous positional files +sync random access +flush +close/disposal +cancellation +``` -SharedWorker - async OPFS through the actual SharedWorker realm +Bun tests also verify the Bun-specific read/replace route rather than only the Node-compatible fallback. -ServiceWorker - black-box registration + postMessage in all browsers - deeper serviceWorkers() instrumentation in Chromium only +## Browser tests use Playwright -iframes - same-origin - cross-origin - opaque sandbox +Playwright owns browser lifecycle and cross-browser orchestration. The matrix covers Chromium, Firefox, and WebKit +instead of encoding a Chromium-only browser assumption. -storage lifecycle - fresh BrowserContext isolation - persistent profile close/reopen +Browser cases include: +```text +Window async OPFS +DedicatedWorker +SharedWorker +ServiceWorker observable behavior +same-origin iframe +cross-origin iframe +opaque sandbox iframe +persistent/reopen behavior browser record adapters - localStorage - IndexedDB - Cache Storage +locking/cancellation where the browser exposes the capability ``` -The iframe and ServiceWorker tests report unsupported runtime APIs as capability skips. A supported realm whose OPFS root is -rejected is not silently skipped; the test asserts that a normalized root failure is present. +Synchronous OPFS access is probed from the actual realm/handle. A test does not infer support from browser name or +worker type. + +Playwright's deeper ServiceWorker instrumentation is browser-specific, so cross-browser service-worker tests use a +page-owned registration/message path when direct runner instrumentation is unavailable. -Benchmarks measure overhead against the direct backend ------------------------------------------------------- +## Provider tests use Testcontainers -A benchmark without a raw baseline cannot tell whether the adapter is fast or merely whether one code path is faster than -another OPFS code path. The benchmark layout therefore keeps three layers visible: +`tests/provider/fixture.ts` owns disposable provider services through Testcontainers. ```text -raw backend - | - v -adapter primitive - | - v +ProviderFixture + | + +-- SeaweedFS S3-compatible endpoint + `-- Azurite Blob endpoint +``` + +Testcontainers selects mapped host ports and owns readiness/disposal. The repository does not keep a parallel Docker +Compose, fixed-port, curl-polling lifecycle. + +Provider tests exercise: + +```text +protocol client + -> provider + +driver + -> client/provider + +adapter + -> driver + FileSystemType - | - +-- coordination: none - `-- coordination: local + -> adapter ``` -`bench/memory.bench.ts` measures raw `Map`, direct `RecordStoreType`, direct memory adapter, and facade overhead. This exposes the -cost of record serialization separately from the higher-level filesystem contract. +The provider suite proves interoperability with SeaweedFS/Azurite. It does not redefine Amazon S3 or Azure Blob +specifications. Deterministic protocol tests continue to protect exact signing, conditions, limits, and error parsing. -`bench/node.bench.ts`, `bench/deno.bench.ts`, and `bench/bun.bench.ts` compare raw host filesystem reads/writes and copy with -the direct adapter and facade. Bun measures both Node-compatible `copyFile` and `Bun.write(destination, Bun.file(source))` so a -future adapter change has a runtime baseline instead of an assumption. `bench/deno-kv.bench.ts` does the same for real local Deno -KV, and `bench/sqlite.bench.ts` compares a raw SQLite BLOB row with the direct record adapter and facade. Metrics are disabled for -facade baseline measurements. +## Stress and lifecycle tests -```sh -deno task bench:memory -deno task bench:deno -deno task bench:deno-kv -deno task bench:node -deno task bench:sqlite -deno task bench:bun +`test:stress` runs the portable suite repeatedly with a fixed shuffle seed. Its purpose is to expose ordering, lock, +cleanup, and state-sharing defects that one deterministic order can hide. + +Lifecycle-sensitive code must test: + +- successful cleanup; +- cleanup after failure; +- caller cancellation; +- post-open stream cancellation; +- close exactly once; +- abort exactly once; +- use after close/abort; +- ownership transfer versus borrowed resources; +- cleanup with a separate signal when the caller signal is already aborted. + +## Coverage + +Coverage is useful evidence, not architectural proof. The coverage task exists to find unexecuted branches in portable +code. A high line percentage does not prove that provider limits, cancellation, or resource ownership are correct. + +## Benchmarks measure each layer + +Every benchmark should identify the cost added by one layer. + +For an object protocol: + +```text +official/native SDK baseline + | +project protocol client + | +project object driver + | +project object adapter + | +FileSystemType metrics:none + | +FileSystemType metrics:basic ``` -`mise run bench` runs the server/memory set with the pinned runtimes. +For a host/native filesystem: -The browser benchmark keeps the raw browser API, direct adapter, and facade visible in each real browser. Native OPFS uses 25 -replace/read iterations with a 64 KiB payload. localStorage, IndexedDB, and Cache Storage use 20 iterations with a 16 KiB -payload. Each sample records the raw, adapter, and facade durations plus the adapter/raw, facade/raw, and facade/adapter ratios -as a Playwright attachment. +```text +raw runtime filesystem API + | +project file driver + | +project file adapter + | +FileSystemType +``` -```sh -deno task bench:browser -# or -mise run bench-browser +For memory/record storage: + +```text +raw Map/value structure + | +record driver + | +record adapter + | +FileSystemType ``` -Microbenchmarks are evidence about overhead in the measured operation. They are not universal provider throughput numbers. -Object-store latency, geographical distance, TLS, provider multipart behavior, and connection reuse can dominate the small -client/facade cost. +The benchmark result should include throughput/latency plus semantic context. A faster route is not a valid substitute +if it has different supported operations, consistency, atomicity, or caching semantics. -`mise run bench-providers` starts the pinned SeaweedFS/Azurite services through Testcontainers and compares official SDK/native -runtime baselines with the direct protocol clients, object adapters, and filesystem facade. `bench/providers.ts` owns provider -startup before it launches benchmark programs, so image pull/readiness time is outside Mitata samples. S3 includes AWS SDK v3 and a Bun-native `S3Client` run; Azure -uses `@azure/storage-blob`. The small write baseline includes the same follow-up stat/properties request as the direct project -client, and multipart/block cases are separate. `metrics: "none"` versus `metrics: "basic"` makes instrumentation overhead -visible instead of hiding it. Real-cloud performance still requires an opt-in controlled provider benchmark. +## Provider benchmarks -Type, lint, format, and documentation gates stay separate ---------------------------------------------------------- +`bench/provider.bench.ts` uses the Testcontainers provider fixture and compares: -`deno task check` type-checks the code in environment-focused groups so unrelated ambient globals do not accidentally make an -invalid target look valid: +S3: ```text -check:core - root/core + provider-neutral adapters/clients + reverse drivers +AWS SDK +project S3 client +project S3 driver +project object adapter +facade metrics:none +facade metrics:basic +``` -check:browser - Window/browser storage adapters + Playwright specs/config +Azure: -check:workers - DedicatedWorker / SharedWorker / ServiceWorker fixtures with WebWorker libs +```text +Azure SDK +project Azure client +project Azure driver +project object adapter +facade metrics:none +facade metrics:basic +``` + +`bench/bun-provider.bench.ts` also compares Bun's native S3 client against the same project layers when Bun is +available. -check:server - Deno / Deno KV / Node / Bun / SQLite adapters and server benchmarks +Provider container startup/readiness happens before measured samples. Container pull/start time is not benchmark data. -check:tests - portable + runtime test source +## Filesystem-client baselines -check:deno-kv - Deno KV test source with the unstable KV flag +`bench/filesystem-provider.bench.ts` compares already-mounted provider filesystem clients through the same local-file +staircase: -check:providers - Testcontainers fixture + provider tests + provider benchmark orchestration +```text +raw mounted path + -> Node file driver + -> file adapter + -> FileSystemType ``` -The normal quality gates are: +Environment variables select mounted roots: ```sh -deno ci -deno task check -deno task lint -deno task doc -deno task fmt:check +OPFS_MOUNTPOINT_S3_ROOT=/mnt/s3 \ +OPFS_BLOBFUSE_ROOT=/mnt/azure \ +mise run bench-filesystem-clients ``` -`deno ci` is important because the committed manifests and lockfile must describe one dependency graph. A changed dependency is -not ready for release until the real lockfile has been regenerated and the frozen install succeeds. +The external mounts are intentionally not started inside the normal Testcontainers fixture. AWS Mountpoint and Azure +BlobFuse are FUSE/system clients with host privileges, installation, mount, and unmount lifecycle beyond a normal +application container. A dedicated benchmark runner can provision them and then call the same mise task. -Release validation checks the artifact, not only the source tree ---------------------------------------------------------------- +Only comparable operations should be measured. Unsupported filesystem operations are capability differences, not +benchmark failures. -Before publication, the repository runs the complete source-level checks plus registry dry-runs. npm packaging uses Deno's -package output and then adjusts Drizzle from a normal generated dependency to the optional peer relationship authored by this -project. +## Browser benchmarks -The artifact gate should verify: +The Playwright benchmark compares raw native OPFS calls against the package's OPFS driver, adapter, and facade where +practical. Each browser result is separate. A result from one browser is not generalized to another engine. -1. the public export map contains every intended subpath and no internal-only file; -2. the generated npm package has JavaScript/declarations that import in Node, Deno, and Bun; -3. browser-safe imports bundle without pulling server-only adapters into the root graph; -4. optional Drizzle remains optional until its subpath is imported; -5. package files exclude tests, benchmarks, coverage, temporary output, and repository-only tooling; -6. the extracted artifact passes the same checks that are meaningful after packaging. +## Metrics cost is measurable -The release command is: +Facade metrics support: -```sh -deno task release:check +```text +none +basic +timing ``` -Browser tests and browser benchmarks remain explicit matrix jobs because downloading three browser engines is a large operation -and should not be hidden inside every local unit-test invocation. +`none` is the instrumentation baseline. `basic` records counters without per-operation timing. `timing` adds monotonic +clock work. Benchmarks keep those modes separate so metrics overhead cannot hide inside the main facade result. + +Driver physical metrics are also distinct from facade metrics. S3/Azure request/retry/part work should not be inferred +from one logical filesystem write. + +## Quality gate + +`mise run quality` owns the Deno-centric release quality gate: + +```text +frozen dependency install +strict check graph +lint +public documentation lint +format check +stress tests +coverage tests +JSR dry-run +npm/deno package dry-run +``` + +`mise run test`, browser tests, provider tests, and runtime matrix jobs add the environment-specific evidence. + +## Agent validation + +A ChatGPT/agent host can lack Deno, Bun, Docker, mise, package registry access, or Playwright browsers. Temporary +validation support belongs under `.agents/` and never becomes production code. + +Allowed fallback rules: + +1. Keep production source Deno/browser/server-native. +2. Use the installed Node.js/TypeScript toolchain for supplemental strict checks. +3. Add narrow `.agents/` declarations/stubs only for dependencies unavailable in the host. +4. Do not change production imports merely to satisfy the agent host. +5. Report missing canonical runtime gates explicitly. + +A validation-only type stub can prove project TypeScript structure. It cannot prove the external dependency's real +runtime or full type contract. Release CI must run against the actual dependency graph. + +## Artifact verification + +Before delivering a modified ZIP: + +1. run every available strict/type/behavior/configuration check on the working tree; +2. inspect stale exports/imports and documentation terminology; +3. inspect package exports and publish payload; +4. remove generated validation/build dependency state; +5. create the ZIP; +6. extract that exact ZIP to a clean directory; +7. recreate only validation-side host declarations if needed; +8. rerun the available checks against the extracted artifact; +9. compare source/extracted file lists; +10. record SHA-256. + +The extracted artifact is the final thing that must pass the claimed checks. A green mutable working tree is not enough. -Testcontainers provider tests ------------------------------ +## Release evidence -`mise run test-providers` runs `tests/provider.test.ts` under Node. The suite opens pinned SeaweedFS and Azurite services through -Testcontainers, uses random mapped host ports, waits for provider readiness, and releases every owned container after the suite. -It proves real HTTP/signing/interoperability for the direct clients without a repository-owned Docker Compose lifecycle. It does -not replace deterministic request-shape tests or real-cloud conformance. See [providers.md](./providers.md) for the exact coverage -and limitations. `mise run bench-providers` starts the same fixture outside the timed benchmark programs. +A release-ready claim requires all applicable canonical gates, including Deno, Node, Bun, Playwright, provider +containers, package dry-runs, and lockfile validation. If the current host cannot run one of those environments, the +result is recorded as unverified rather than passed.