From e66923f7941b1549934f3ba9c456057eabb69ac5 Mon Sep 17 00:00:00 2001 From: Okiki Ojo Date: Tue, 18 Aug 2026 11:40:38 -0400 Subject: [PATCH] docs(architecture): explain driver and provider boundaries Previous documentation described the storage stack less clearly than the current refactor uses it in code. Document the driver -> adapter -> FileSystemType -> bridge model, update the API and design guides around limit provenance and planning, and refresh the provider docs so reviewers and users can understand the intent of the new architecture before reading implementation details. --- AGENTS.md | 60 ++- README.md | 545 +++++++++++++----------- docs/adapters.md | 587 +++++++++++++------------- docs/api.md | 973 +++++++++++++++++++++++++------------------ docs/azure.md | 471 ++++++++++----------- docs/design.md | 892 +++++++++++++++++++++++++++------------ docs/ecosystems.md | 448 +++++++++++--------- docs/environments.md | 261 ++++++------ docs/providers.md | 363 ++++++++-------- docs/releasing.md | 214 ++++++---- docs/s3.md | 566 ++++++++++++------------- docs/sources.md | 389 ++++++++--------- docs/validation.md | 565 ++++++++++++++----------- 13 files changed, 3538 insertions(+), 2796 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 3adf43d..5900947 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -1,19 +1,55 @@ # Repository implementation rules +Read `README.md` and `docs/design.md` before changing architecture or public APIs. + - Treat `@okikio/opfs` as a library programming model, not as an application runtime. - Keep the root entrypoint import-safe in Window, Worker, Deno, Bun, and Node contexts. -- Put concrete storage integrations under `src/adapter/` and expose them through explicit public subpaths. -- Put reverse ecosystem interfaces under `src/driver/`. -- Prefer one-word file and folder names. Use more words only when the precise concept requires them. +- Use the storage path literally: `client -> driver -> adapter -> FileSystemType -> bridge`. +- A client owns a wire protocol when one exists. It does not own OPFS semantics. +- A driver owns backend-native storage mechanics, requirements, limits, optimization policy, physical metrics, and + resources that it explicitly acquires. +- An adapter translates one driver into the small canonical OPFS primitive contract. It does not reimplement provider + mechanics. +- A bridge starts from `FileSystemType` and implements a real ecosystem contract. Do not call direction metadata or a + nominal wrapper a bridge. +- Put direction/support metadata under `src/integration/`. Definitions remain import-safe and require no global + registry. +- Do not preserve obsolete layer names or compatibility entrypoints. Update every current consumer, test, export, + document, and benchmark when a public contract changes. +- Prefer one-word file and folder names. Use two or three words only when the precise concept requires them. - All Zod schema constants end in `Schema`. - Project-owned data types normally end in `Type`. -- Prefer direct schema and type exports. Use namespace imports only when they improve short operation call sites. -- Prefer `get`, `create`, `open`, `save`, `inspect`, `plan`, `convert`, `read`, `write`, `close`, `remove`, `copy`, and `move` over vague verbs. -- Avoid `generate`, `execute`, `handle`, `process`, `manager`, `helper`, `common`, `shared`, and `misc` unless an external protocol requires the word. -- Keep Deno, Bun, Node, browsers, and Workers on the same core TypeScript source. Runtime-specific adapters can use runtime-specific APIs behind explicit subpaths. -- Prefer Web APIs and existing standard-library capabilities before custom infrastructure. -- Adapters never configure logging, read environment variables, or acquire unrelated global resources at import time. -- The caller owns injected database, collection, storage, and filesystem resources unless an adapter option explicitly transfers ownership. -- TSDoc and comments use plain technical English. They teach options, examples, impact, reasoning, ownership, limits, failure behavior, and necessary background. -- Document public schemas, types, properties, functions, classes, and adapter contracts. Document internal symbols when their invariant, lifecycle, or failure behavior is not obvious. +- Prefer direct schema and type exports. Use namespace imports only when they improve short coherent operation call + sites. +- Prefer `get`, `create`, `open`, `save`, `inspect`, `plan`, `read`, `write`, `close`, `remove`, `copy`, and `move` over + vague verbs. +- Avoid `generate`, `execute`, `handle`, `process`, `manager`, `helper`, `common`, `shared`, and `misc` unless an + external protocol requires the word. +- Keep Deno, Bun, Node, browsers, and Workers on the same core TypeScript source. Runtime-specific drivers stay behind + explicit subpaths. +- Prefer Web APIs and `@std/*` before custom infrastructure when they provide the required semantics. +- Importing a module must not connect to a provider, read credentials, configure logging, start workers, or perform + unrelated global work. +- The caller owns injected databases, collections, clients, stores, mounts, and filesystems unless an explicit option + transfers ownership. +- Every behavior-changing optimization must be independently disableable and visible through inspection. +- Keep provider hard limits, implementation safety limits, user policy, and dynamic probe results distinct. Do not + collapse them into one unexplained number. +- `plan()` is deterministic preflight. It must not perform provider I/O. A probe is a separate explicit operation. +- Partitioning belongs to the driver that owns the physical layout. The driver must describe visibility, part limits, + cleanup, and whether the layout changes observable behavior. +- Use `node:test` with `describe` and `it`; use `@std/expect` for expectations. Playwright owns real browser environment + tests. +- Use mise as the repository command authority. GitHub Actions owns triggers, permissions, matrices, outputs, and + secrets, then calls `mise run ...`. +- Testcontainers owns disposable provider fixtures. Do not restore fixed host ports, hand-written readiness polling, or + Compose lifecycle scripts for S3/Azure tests. +- Benchmarks compare the native/provider baseline, project client when present, driver, adapter, facade with metrics + disabled, and measured facade. Compare filesystem clients such as Mountpoint and BlobFuse only on operations their + documented semantics support. +- TSDoc and comments use plain technical English. Teach options, examples, impact, reasoning, ownership, limits, + cancellation, performance, and necessary background. +- Document important non-exported symbols when they own an invariant, lifecycle rule, serialization layout, or failure + rule. - Comments explain why a rule exists or what must remain true. Do not restate obvious syntax. +- Keep agent-only validation under `.agents/`. It must never become a production import or published artifact. diff --git a/README.md b/README.md index 55cac1e..075d33a 100644 --- a/README.md +++ b/README.md @@ -1,323 +1,384 @@ -@okikio/opfs -============ +# @okikio/opfs -`@okikio/opfs` is an OPFS-shaped filesystem programming model that can sit on top of browser OPFS, host filesystems, -object stores, key-value stores, browser storage, document databases, and SQL databases. +`@okikio/opfs` is an OPFS-shaped storage programming model for browser OPFS, host filesystems, object stores, key-value +stores, browser storage, document databases, and SQL databases. -The public filesystem owns the semantics that application code should not have to rebuild: canonical virtual paths, -OPFS-shaped handles, recursive copy and move, cancellation, staged writable files, bounded stream fallbacks, coordination, -normalized failures, and resource ownership. An adapter translates those operations into one concrete backend. +The package does not pretend those systems are identical. It separates protocol behavior, backend storage mechanics, +OPFS translation, portable filesystem semantics, and reverse ecosystem projections so applications can inspect the real +route and its limits before work starts. ```text -application - | - +-- path API -------------------+ - | readFile / writeFile | - | copy / move / walk | - | v - +-- OPFS-shaped handles --> FileSystemType - | - canonical adapter operations - | - +---------------------+---------------------+ - | | | - v v v - native files record stores object stores - OPFS / Node / Deno KV / DB rows S3 / Azure Blob - / Bun | - +-- localStorage - +-- IndexedDB - +-- Cache Storage - +-- Deno KV - +-- unstorage - +-- RxDB - +-- db0 / SQLite - `-- Drizzle +ecosystem / native API + | + v + client protocol client when a wire protocol exists + | + v + driver backend-native storage contract + | requirements / limits / optimizations / physical metrics + v + adapter driver -> canonical OPFS primitives + | + v + FileSystemType paths / handles / locks / fallbacks / logical metrics + | + v + bridge FileSystemType -> real ecosystem contract ``` -The reverse direction is useful too. `@okikio/opfs/driver/kv` exposes any `FileSystemType` as a small hierarchical key-value -store, and `@okikio/opfs/driver/unstorage` adapts that view to unstorage. This means an application can give unstorage an OPFS, -Node, Deno, Bun, S3, Azure Blob, IndexedDB, Deno KV, SQLite, or another custom OPFS backend without a second provider matrix. +A client is optional. Node, Deno, Bun, browser OPFS, IndexedDB, and SQLite can start at the driver layer. S3 and Azure +Blob have explicit protocol clients because their wire contracts are independently useful. -Install and start with the backend you actually own ----------------------------------------------------- +The filesystem facade owns the behavior application code should not rebuild for every backend: canonical virtual paths, +OPFS-shaped handles, recursive work, cancellation, staged writable files, bounded stream fallbacks, coordination, +normalized failures, resource ownership, and deterministic preflight planning. -Deno and JSR can import the package directly: +## Start with the storage you own -```ts -import { openFileSystem } from "jsr:@okikio/opfs"; - -const fileSystem = await openFileSystem(); -try { - await fileSystem.writeFile("/state/app.json", "{}", { parents: true }); -} finally { - await fileSystem.close(); -} -``` +Browser OPFS has a root convenience function: -npm-compatible runtimes use the same TypeScript API: +```ts +import { openFileSystem } from "@okikio/opfs"; -```sh -npm install @okikio/opfs +await using fileSystem = await openFileSystem(); +await fileSystem.writeFile("/state/app.json", "{}", { parents: true }); ``` -Server code selects a concrete adapter instead of importing a different filesystem API: +Server code can compose each layer explicitly: ```ts import { createFileSystem } from "@okikio/opfs"; import { createNodeAdapter } from "@okikio/opfs/adapter/node"; -const fileSystem = createFileSystem( +await using fileSystem = createFileSystem( createNodeAdapter({ root: "./data" }), { coordination: "local" }, ); ``` -The root entrypoint is intentionally import-safe in browsers, workers, Deno, Bun, and Node. Runtime-specific dependencies stay -on explicit subpaths. Importing `@okikio/opfs` does not import `node:fs`, inspect environment variables, connect to databases, -or configure application logging. - -The first-party backend set is deliberately broad, but the layers stay small: - -| Subpath | Backend or role | Important behavior | -| --- | --- | --- | -| `adapter/opfs` | native browser OPFS | native handles, streams, sync access when exposed | -| `adapter/node` | `node:fs` | streams, ranges, copy/rename, sync random access | -| `adapter/deno` | Deno filesystem | streams, ranges, copy/rename, sync random access | -| `adapter/bun` | Bun + Node-compatible fs | Bun read/write fast paths plus host filesystem operations | -| `adapter/memory` | in-memory records | deterministic tests and temporary state | -| `adapter/record` | `RecordStoreType` | common translation for value/document/SQL stores | -| `adapter/object` | `ObjectStoreType` | common translation for object stores without hiding object semantics | -| `adapter/s3` | `S3ClientType` | direct S3/S3-compatible storage, no AWS SDK | -| `adapter/azure` | `AzureClientType` | direct Azure Blob REST storage, no Azure SDK | -| `adapter/localstorage` | Web Storage | synchronous string store translated through records | -| `adapter/indexeddb` | IndexedDB | indexed record persistence with caller-controlled database ownership | -| `adapter/cache` | Cache Storage | record persistence in an injected Cache | -| `adapter/deno-kv` | Deno KV | record persistence over a caller-owned KV database | -| `adapter/sqlite` | connected SQLite | focused SQLite view over the same SQL record contract as db0 | -| `adapter/unstorage` | unstorage `Storage` | forward bridge above the selected unstorage driver | -| `adapter/rxdb` | RxDB collection | forward bridge above the selected RxStorage | -| `adapter/db0` | db0 `Database` | SQL bridge across db0 dialects/connectors | -| `adapter/drizzle` | Drizzle database + table | caller-owned schema and common CRUD bridge | -| `driver/kv` | reverse key-value view | collision-safe key hierarchy over any filesystem | -| `driver/unstorage` | reverse unstorage driver | lets unstorage consume any `FileSystemType` | - -Drizzle is an optional peer dependency because the integration is only loaded through its explicit subpath. - -Object storage is not flattened into a fake local disk -------------------------------------------------------- - -S3 and Azure Blob can both back the filesystem facade, but they remain object stores underneath. That distinction affects -performance and correctness. - -A complete replacement can stream to multipart/block upload. Append and update cannot normally mutate object bytes in place, -so the object adapter performs a read-modify-write operation. When the provider exposes conditional writes, the previous ETag -is used as an optimistic precondition so a concurrent writer fails rather than being silently overwritten. - -Native copy is also a separate capability. The filesystem asks the adapter to copy before it opens a source stream, so -provider-side copy stays inside S3/Azure instead of becoming an accidental download and re-upload. +`createNodeAdapter()` is a convenience composition. The explicit form is useful when an application wants to inspect or +use the backend driver before it creates a filesystem: + +```ts +import { createFileSystem } from "@okikio/opfs"; +import { createFileAdapter } from "@okikio/opfs/adapter/file"; +import { createNodeDriver } from "@okikio/opfs/driver/node"; + +const driver = createNodeDriver({ root: "./data" }); +console.log(driver.inspect()); +console.log(driver.plan({ operation: "write", path: "/large.bin", size: 1_000_000 })); + +const adapter = createFileAdapter(driver); +await using fileSystem = createFileSystem(adapter); +``` + +The root entrypoint remains import-safe in Window, Worker, Deno, Bun, and Node contexts. Runtime-specific code stays on +explicit subpaths. Importing the root package does not connect to a provider, read credentials, start workers, or +configure application logging. + +## Layer inventory + +The first-party integration set is broad, but each layer has one job. + +| Family | Client | Driver | Adapter | Reverse bridge | +| --------------- | --------------------- | --------------------- | ---------------------- | --------------------------------------- | +| browser OPFS | n/a | `driver/opfs` | `adapter/opfs` | filesystem itself | +| Node filesystem | n/a | `driver/node` | `adapter/node` | ecosystem-specific | +| Deno filesystem | n/a | `driver/deno` | `adapter/deno` | ecosystem-specific | +| Bun filesystem | n/a | `driver/bun` | `adapter/bun` | ecosystem-specific | +| memory | n/a | `driver/memory` | `adapter/memory` | generic KV bridge possible | +| Deno KV | n/a | `driver/deno-kv` | `adapter/deno-kv` | generic KV bridge possible | +| localStorage | n/a | `driver/localstorage` | `adapter/localstorage` | generic KV bridge possible | +| IndexedDB | n/a | `driver/indexeddb` | `adapter/indexeddb` | generic KV bridge possible | +| Cache Storage | n/a | `driver/cache` | `adapter/cache` | cache-specific bridge not yet provided | +| SQLite rows | engine-owned | `driver/sqlite` | `adapter/sqlite` | see database direction below | +| db0 | connector-owned | `driver/db0` | `adapter/db0` | no fake SQL projection | +| Drizzle | dialect/driver-owned | `driver/drizzle` | `adapter/drizzle` | no fake SQL projection | +| RxDB | RxStorage-owned | `driver/rxdb` | `adapter/rxdb` | full RxStorage bridge not yet provided | +| unstorage | upstream driver-owned | `driver/unstorage` | `adapter/unstorage` | `bridge/unstorage` | +| S3 | `s3` | `driver/s3` | `adapter/s3` | object-specific bridge not yet provided | +| Azure Blob | `azure` | `driver/azure` | `adapter/azure` | object-specific bridge not yet provided | + +`adapter/file`, `adapter/record`, and `adapter/object` are reusable translators for third-party drivers. `driver/file`, +`driver/record`, and `driver/object` are the corresponding backend-native contracts. + +## Drivers are independently useful + +A driver is not an adapter with a new name. It owns backend mechanics that remain meaningful without `FileSystemType`. + +A configured driver reports: ```text -filesystem.copy() - | - +-- nativeCopy ----> provider/server-side copy - | - `-- fallback ------> source stream -> bounded transfer -> destination +name and storage family +stable backend operations/capabilities it provides +backend resource ownership: none / borrowed / owned +requirements and current availability +provider hard limits +implementation safety limits +user-selected policy limits +dynamic limits that still require probing +behavior-changing and transparent optimizations +structured preflight problems and actions +physical metrics when available +owned-resource disposal when applicable ``` -The direct S3 client implements Signature Version 4, range reads, ListObjectsV2, multipart upload, conditional completion, -CopyObject, and multipart UploadPartCopy for objects above CopyObject's 5 GB source limit. It also checks S3's unusual -success-with-error-body responses for copy and multipart completion. +Limits always include provenance. For example, Deno KV can report its serialized provider ceiling separately from the +smaller raw payload budget this library chooses to leave room for serialization overhead. A caller can therefore +distinguish a provider rule from an implementation safety choice and from its own `maxParts` policy. + +`plan()` is deterministic and performs no provider I/O: ```ts -import { createFileSystem } from "@okikio/opfs"; -import { createS3Adapter } from "@okikio/opfs/adapter/s3"; -import { createS3Client } from "@okikio/opfs/s3"; - -const client = createS3Client({ - endpoint: "https://s3.us-east-1.amazonaws.com", - bucket: "my-bucket", - region: "us-east-1", - credentials: { accessKeyId, secretAccessKey }, +const plan = driver.plan({ + operation: "write", + path: "/archive/data.bin", + size: 80 * 1024, + source: "bytes", + mode: "replace", }); -const fileSystem = createFileSystem(createS3Adapter(client)); +for (const problem of plan.problems) console.log(problem.code, problem.limit); +for (const action of plan.actions) console.log(action.kind); ``` -S3 compatibility is a protocol family, not one identical product. The client therefore accepts endpoint, region, addressing, -headers, and capability overrides. For example, an S3-compatible service that does not support multipart preconditions should -set `conditionalWrite: false` instead of pretending the safety property exists. `client.request()` remains available for S3 -features that do not belong in the portable filesystem contract. - -Azure uses its own REST model instead of being forced through an S3 abstraction. It supports SAS, Microsoft Entra bearer -tokens, Shared Key, and caller-defined authorization headers. Large server-side copies use Put Block From URL after Azure's -smaller synchronous Copy Blob From URL path is no longer sufficient. +Dynamic facts such as available quota remain unknown until a separate explicit probe supplies them. The planner never +invents a quota or silently performs network/storage I/O. -The protocol clients are documented separately because their wire contracts are larger than the filesystem adapter surface: +## Capabilities are layered instead of flattened -- [S3 client protocol](./docs/s3.md) covers SigV4, request canonicalization, multipart upload/copy, conditions, limits, errors, - compatibility controls, and known non-goals. -- [Azure Blob client protocol](./docs/azure.md) covers REST versions, SAS/bearer/Shared Key authorization, block upload/copy, - conditions, limits, errors, and Azurite behavior. -- [Provider integration tests](./docs/providers.md) explains the Testcontainers-backed SeaweedFS and Azurite matrix and what those - local providers can and cannot prove. +`driver.inspect()` describes the backend. `FileSystemType.inspect()` adds the adapter translation and effective facade +routes: -The facade makes capability, limits, routing, and cost inspectable ----------------------------------------------------------------- +```ts +const inspection = fileSystem.inspect(); + +inspection.driver; // provides, ownership, requirements, limits, optimizations +inspection.adapter; // native OPFS translation capabilities +inspection.support; // effective native/emulated/partitioned/unsupported routes +inspection.optimizations; // facade route switches +inspection.metrics; // logical filesystem counters +inspection.driverMetrics; // physical backend counters when available +``` -`AdapterCapabilitiesType` describes immediate adapter behavior. `FileSystemType.inspect()` describes the configured stack after -facade fallbacks and optimization policy are applied. This distinction lets callers ask whether a route is `native`, `emulated`, -`partitioned`, or `unsupported` without guessing from the adapter name. +`FileSystemType.plan()` combines the driver preflight with adapter and facade policy. Problems remain structured and +retain the layer that identified them. Human-readable messages are presentation data, not the machine contract. ```ts -const fileSystem = createFileSystem(adapter, { - maxBufferedWriteBytes: 32 * 1024 * 1024, - metrics: "basic", - optimizations: { - nativeCopy: false, - }, -}); - -console.log(fileSystem.inspect()); -console.log(fileSystem.plan({ +const plan = fileSystem.plan({ operation: "write", + path: "/video.bin", source: "stream", - mode: "replace", size: 512 * 1024 * 1024, -})); +}); + +if (!plan.supported) { + // Example actions: partition, change-policy, reduce-input, select-driver. + console.log(plan.problems, plan.actions); +} ``` -`inspect()` includes native capabilities, effective support, hard limits known by the adapter, partition layout, resolved -optimization controls, the facade buffer ceiling, and a detached metrics snapshot. `plan()` is deterministic and does no I/O. -When size is known it can reject a request before work begins, show expected facade materialization, or explain the physical -part count selected by a partitioned adapter. Unknown provider limits remain unknown rather than being invented. +Every optimization that can change request count, failure timing, storage layout, consistency, atomicity, or another +observable property is independently disableable and visible through inspection. -Write planning separates the resulting logical file from the bytes supplied by the current call. `size` checks logical -file/partition limits. `inputBytes` checks whether a non-native input stream fits under `maxBufferedWriteBytes`. Replace usually -needs only `size`; append/update should provide both values when they are known. +## Object storage keeps object semantics -Optimizations that select a materially different route are independently disableable: native stream read/write, direct range -read, native/server-side copy, and native move. The fallback is used only when it can preserve the portable filesystem contract. -For example, disabling provider-side copy can force bytes through this process and cannot reproduce provider-private control-plane -metadata such as every ACL, tag, lock policy, or checksum policy. Portable file bytes and `mediaType` are preserved. +S3 and Azure Blob do not become fake POSIX disks. Their clients keep protocol-specific operations while object drivers +expose portable object mechanics to the adapter. -`metrics: "none"` removes facade counter updates for baseline benchmarks. `basic` counts operations, bytes, failures, route -selection, and peak facade materialization. `timing` adds monotonic durations. The direct S3 and Azure clients expose separate HTTP -request/retry metrics so protocol overhead and facade overhead can be measured independently. +```text +S3 REST / Azure Blob REST + | + v + client + | + v + object driver + | + v + object adapter + | + v + FileSystemType +``` -Large values are a backend capability, not a promise that every value store is unlimited. `adapter/deno-kv` is the first record -backend with a physical partition layout. Small files stay inline; large files use raw binary parts and a manifest-last commit. -Metadata lookup, directory listing, byte ranges, and stream reads do not reconstruct the complete logical file. The partition -policy is `never | auto | always`, so applications that do not want a changed durable layout can disable it explicitly. -Materialized append/update writes also build a new generation part-by-part, so the existing logical file is not joined into one -large base64 record before a small patch can be applied. +A complete replacement can stream through multipart/block upload. Append and update normally require read-modify-write. +Server-side copy remains a separate capability because routing bytes through JavaScript is not equivalent to a provider +control plane copy. -`streamWriteModes` remains mode-specific. A simple record adapter can have no native stream lane, while Deno KV can advertise a -partitioned replacement stream and an object store can advertise native replacement streaming. Append and update can still be -emulated or unsupported independently. Deno KV's materialized append/update lane is direct, but streamed append/update remains -emulated because the incoming stream must first fit under the facade buffer ceiling. +The S3 client exposes two optimization switches that are useful to inspect and benchmark: -Bridges make both integration directions explicit -------------------------------------------------- +- `delayedMultipart`: buffers the first bounded part so a small unknown-length stream can use `PutObject`. Disable it + when the multipart request lifecycle itself is required. +- `signingKeyCache`: reuses the derived SigV4 signing key for unchanged credentials/date/region/service. It does not + change the signed request semantics. -Adapters remain `ecosystem -> OPFS`. Drivers remain `OPFS -> ecosystem`. A bridge groups both directions and records an explicit -reason when one direction cannot honestly exist. +Azure exposes `blockUpload` and `serverCopy` independently. Disabling block upload removes native streamed-write support +rather than pretending a stream can still be sent without staging. Disabling server copy makes the object adapter/facade +choose an honest fallback when one can preserve the requested semantics. -```ts -import { UnstorageBridge, RxDbBridge } from "@okikio/opfs/bridge"; +See [docs/s3.md](./docs/s3.md) and [docs/azure.md](./docs/azure.md) for the protocol contracts. + +## Deno KV is the reference partitioned record driver -console.log(UnstorageBridge.directions); -// { toOpfs: { supported: true }, fromOpfs: { supported: true } } +Deno KV demonstrates why one flat `maxFileBytes` number is insufficient. The driver separates: -console.log(RxDbBridge.directions.fromOpfs); -// unsupported: a filesystem is not an RxStorage query/conflict/change-stream engine +```text +provider: serialized key/value and atomic-operation ceilings +implementation: conservative inline/part payload budgets +user policy: partition mode, part size, max parts, I/O concurrency +derived: logical file capacity for the selected layout ``` -The included bridge descriptors cover unstorage, RxDB, db0, Drizzle, and the generic reverse key-value view. Reverse KV and -unstorage drivers also expose the backing filesystem's `inspect()`, `plan()`, and `getMetrics()` methods so capability, size, -partition, optimization, and instrumentation decisions remain visible after the direction changes. Third parties can -use `defineBridge()` without a global registry. An unsupported direction must include a reason, which prevents a bridge from -silently pretending that asynchronous filesystem behavior can provide an unrelated synchronous or query-oriented contract. +Large files use immutable physical generations and a manifest-last visibility point: -Use schemas directly --------------------- +```text +old manifest -> old parts + +write new part 0..N + | + v +write new manifest visibility point + | + v +remove old reachable generation +``` -Project-owned structural data is defined by Zod schemas and inferred TypeScript types. Schema constants end in `Schema`, and -project-owned serializable types normally end in `Type`. +A failed write does not publish a partial logical file. Unknown-length streamed replacement uses the partitioned lane +when partitioning is enabled. `partition: "never"` disables that behavior and makes oversized/streaming requests fail or +use a bounded facade fallback instead of changing durable layout silently. -```ts -import { PathSchema, type PathType } from "@okikio/opfs/schema"; +The preflight planner also evaluates the concrete virtual path. Deno KV limits serialized keys, so file size alone is +not enough to decide whether an operation is admissible. + +## Database direction matters -const path: PathType = PathSchema.parse("/cache/result.bin"); +There are two different database architectures and the package documents them separately. + +### Database-backed filesystem + +The existing SQLite/db0/Drizzle/RxDB drivers store logical filesystem records in an existing database abstraction: + +```text +Database / collection + | + v +record driver + | + v +record adapter + | + v +FileSystemType +``` + +For Drizzle, the caller supplies a connected database and a table. The generic driver intentionally reports replacement +as best-effort because its portable CRUD route is delete then insert. A dialect-specific driver can expose stronger +binary, transaction, upsert, or partition behavior without weakening the generic contract. + +### Database file stored on OPFS + +Using OPFS as storage for a SQLite database is the opposite direction: + +```text +application + | + Drizzle + | +SQLite engine + | +SQLite VFS + | +FileSystemType / native OPFS ``` -Zod 4 schemas implement Standard Schema, so consumers that accept Standard Schema can use these exported schemas directly. The -package does not maintain a parallel wrapper layer that could drift from the executable Zod contract. +`adapter/sqlite` does **not** implement this topology. It stores virtual filesystem records inside an already connected +SQLite database. A future SQLite VFS integration must implement the database engine's real VFS contract. It must not be +represented as an SQL bridge that only renames filesystem methods. + +See [docs/ecosystems.md](./docs/ecosystems.md) for concrete Drizzle/db0/RxDB behavior. + +## Bridges are real reverse contracts -Test the semantics where they actually run ------------------------------------------- +A bridge starts from `FileSystemType` and implements an ecosystem contract. Direction metadata lives under +`integration/` and is not itself a bridge. -Portable filesystem contracts use `node:test` and `@std/expect`. Deno runs the same portable test source, Node runs the same -source, and Bun runs the same `node:test` API through its compatibility layer. Runtime-specific suites then prove the real host -filesystem adapters. +The package currently provides: -Playwright Test owns the browser matrix. The same tests run in Chromium, Firefox, and WebKit and exercise Window OPFS, -DedicatedWorker, SharedWorker, ServiceWorker, same-origin and cross-origin iframes, opaque sandbox behavior, fresh-context -isolation, persistent-profile reopen, cancellation, and browser storage adapters. Tests probe capabilities instead of selecting -behavior from browser names. +- `bridge/kv`: a small hierarchical key-value contract over any `FileSystemType`. +- `bridge/unstorage`: an unstorage Driver-shaped implementation over any `FileSystemType`. -Mitata benchmarks compare three layers where possible: +The unstorage layout keeps `foo` and `foo:bar` distinct even though a normal filesystem path cannot be both a file and a +directory. ```text -raw backend API - | - v -adapter primitive - | - v -FileSystemType facade - | - +-- coordination: none - `-- coordination: local +unstorage + | +bridge/unstorage + | +FileSystemType + | +any configured adapter/driver stack ``` -Browser benchmarks compare raw native APIs, direct adapters, and the facade for OPFS, localStorage, IndexedDB, and Cache -Storage in Chromium, Firefox, and WebKit. Node, Deno, and Bun benchmarks compare their raw filesystem APIs with direct adapters and the facade. Bun additionally compares -Node-compatible `copyFile` with `Bun.write(destination, Bun.file(source))` rather than assuming one host copy path is faster. -Deno KV and SQLite have the same raw-to-adapter-to-facade measurements. +`integration` definitions state which directions are real. RxDB, db0, and Drizzle currently remain honest one-way +integrations because a filesystem alone is not an RxStorage query/conflict/change-stream engine or a SQL engine. +Third-party packages can use `defineIntegration()` and the driver/adapter primitives without registering global state. + +## Testing and benchmarks follow the layers + +Portable behavior uses `node:test` with `@std/expect`. The same source is checked/run in the supported server runtimes +where the runtime capability exists. Playwright Test owns actual Window, Worker, iframe, ServiceWorker, persistence, and +browser-storage coverage. Testcontainers owns disposable SeaweedFS and Azurite provider fixtures. -Provider benchmarks use the same local provider fixture but keep each layer separate: official AWS/Azure SDK baseline, direct -protocol client, direct object adapter, facade with metrics disabled, and facade with basic metrics. A Bun-native S3 run compares -Bun's Rust-backed `S3Client` against the same project layers. Multipart/block cases are separate from single-request writes so a -different request plan is never presented as abstraction overhead. +Provider benchmarks compare: -With the pinned mise toolchain: +```text +official/native baseline + | +project protocol client + | +project driver + | +project adapter + | +facade metrics:none + | +facade metrics:basic +``` + +`bench/filesystem-provider.bench.ts` adds the same staircase above already-mounted AWS Mountpoint and Azure BlobFuse +filesystems. Set `OPFS_MOUNTPOINT_S3_ROOT` and/or `OPFS_BLOBFUSE_ROOT` before `mise run bench-filesystem-clients`. The +benchmark uses only file operations that can be compared through the mounted filesystem contract. It does not count an +unsupported filesystem operation as a performance failure. + +Mise is the repository command authority: ```sh mise install mise run check mise run test -mise run bench mise run test-browser -mise run bench-browser +mise run test-providers +mise run bench mise run bench-providers +mise run bench-browser +mise run bench-filesystem-clients # requires external mounts ``` -GitHub Actions uses the same tool declarations and mise tasks. The workflow installs mise once per job, asks mise to install -only the runtimes that job needs, and then calls `mise run ...`. Runtime matrix jobs override one configured version with -`MISE__VERSION`; the Node matrix uses this to test Node 22, 24, and 26 without introducing a second tool-version -source. Third-party actions are pinned to immutable commit SHAs, and the mise binary version is pinned separately. - -The focused Deno tasks are documented in [docs/validation.md](./docs/validation.md). - -Read the rest by the question you have --------------------------------------- - -- [Public API](./docs/api.md) explains the developer-facing filesystem and handle contracts. -- [Adapters](./docs/adapters.md) explains every first-party backend and the contracts for custom storage. -- [Architecture](./docs/design.md) explains invariants, streaming, copy/move, locks, ownership, and failure behavior. -- [Ecosystems](./docs/ecosystems.md) explains unstorage, RxDB, db0, Drizzle, S3-compatible services, and reverse drivers. -- [Environments](./docs/environments.md) explains Window, workers, iframes, Deno, Bun, Node, and provider clients. -- [Validation](./docs/validation.md) defines the canonical test and benchmark matrix. -- [Sources](./docs/sources.md) records the standards and upstream contracts that the implementation follows. -- [Releasing](./docs/releasing.md) explains JSR/npm packaging and release checks. +GitHub Actions owns triggers, permissions, matrices, outputs, and secrets. It then invokes the same mise tasks. Release +and publish commands also live under `.mise/tasks/` rather than becoming a second command layer in workflow YAML. + +## Read next + +- [Public API](./docs/api.md) explains filesystem, inspection, planning, driver, adapter, and bridge entrypoints. +- [Architecture](./docs/design.md) explains layer ownership, invariants, partitioning, metrics, and lifecycle. +- [Adapters and drivers](./docs/adapters.md) explains the first-party translation/backend matrix and extension + contracts. +- [Ecosystems](./docs/ecosystems.md) explains unstorage, RxDB, db0, Drizzle, SQLite direction, and reverse bridge + constraints. +- [Environments](./docs/environments.md) explains browser realms, Deno, Bun, Node, and server coordination. +- [Providers](./docs/providers.md) explains Testcontainers and provider/filesystem baseline benchmarks. +- [Validation](./docs/validation.md) defines the release gates and this repository's test matrix. +- [Sources](./docs/sources.md) records the standards and upstream contracts used by the implementation. +- [Releasing](./docs/releasing.md) explains mise-owned release and registry publication. diff --git a/docs/adapters.md b/docs/adapters.md index 70701ce..edad4af 100644 --- a/docs/adapters.md +++ b/docs/adapters.md @@ -1,131 +1,135 @@ -Adapter guide -============= +# Drivers and adapters -An adapter translates the package's canonical virtual filesystem operations into one concrete backend. The filesystem facade -owns filesystem semantics. The adapter owns backend mechanics. +## Purpose -That distinction lets the required backend contract stay small: +A driver owns backend-native storage. An adapter translates that driver into the small canonical filesystem primitive +contract. Keeping those roles separate lets applications use a driver directly, inspect real provider limits, and +measure adapter/facade overhead independently. ```text -stat read one entry's metadata -readFile read one file or range -writeFile commit one materialized write -readDir lazily list direct children -createDir create exactly one directory -remove remove one file or empty directory +backend/native API + | + driver + | + adapter + | +FileSystemType ``` -The facade builds parent creation, recursive walking, recursive copy/remove, OPFS-shaped handles, write-command staging, -coordination, and normalized errors on top. A backend can add native operations when it can do better than the facade fallback. +Use the convenience adapters for normal application code. Use explicit drivers when you need backend planning, physical +metrics, provider-specific operations, or a custom translation. + +## First-party matrix + +| Storage | Driver | Adapter | Native family | +| ------------------------ | --------------------- | ---------------------- | ------------- | +| browser OPFS | `driver/opfs` | `adapter/opfs` | file | +| Node filesystem | `driver/node` | `adapter/node` | file | +| Deno filesystem | `driver/deno` | `adapter/deno` | file | +| Bun filesystem | `driver/bun` | `adapter/bun` | file | +| memory | `driver/memory` | `adapter/memory` | record | +| Deno KV | `driver/deno-kv` | `adapter/deno-kv` | record | +| localStorage | `driver/localstorage` | `adapter/localstorage` | record | +| IndexedDB | `driver/indexeddb` | `adapter/indexeddb` | record | +| Cache Storage | `driver/cache` | `adapter/cache` | record | +| SQLite rows | `driver/sqlite` | `adapter/sqlite` | record | +| unstorage Storage | `driver/unstorage` | `adapter/unstorage` | record | +| RxDB collection | `driver/rxdb` | `adapter/rxdb` | record | +| db0 Database | `driver/db0` | `adapter/db0` | record | +| Drizzle database + table | `driver/drizzle` | `adapter/drizzle` | record | +| S3 | `driver/s3` | `adapter/s3` | object | +| Azure Blob | `driver/azure` | `adapter/azure` | object | + +Reusable family translators: ```text -openReadStream native streaming read -writeStream native streaming for declared write modes -copy native/server-side file copy -move native rename/move -openWritableFile long-lived asynchronous positional writes -openSyncFile synchronous random access +driver/file -> adapter/file +driver/record -> adapter/record +driver/object -> adapter/object ``` -`AdapterCapabilitiesType` must describe these native paths truthfully. `streamWriteModes` is a list rather than one boolean -because replacement, append, and update can have different backend costs. `nativeCopy` is separate from `nativeMove` because -object stores often copy efficiently but cannot rename an object atomically. +## File drivers -Adapters can also expose `limits` and `partition`. Limits are hard facts known by the configured backend, such as maximum file, -value, key, part, batch, or concurrency sizes. Missing fields mean unknown, not unlimited. Partition describes a durable physical -layout used when one logical file spans multiple provider values. These fields feed `FileSystemType.inspect()` and `plan()` but -do not change the required adapter method set. +`FileDriverType` preserves real file-like operations. Required primitives are metadata, materialized read/write, +direct-child listing, one-directory creation, and single-entry removal. Optional direct operations include streams, +copy, move, positional files, and synchronous random access. -Route-changing optimizations live on the facade, not inside capability flags. `optimizations.streamRead`, `streamWrite`, -`rangeRead`, `nativeCopy`, and `nativeMove` can force the safe fallback for differential testing or application policy. An adapter -should therefore implement the best native route it can and let the caller decide whether to use it. - -Use `createFileSystem()` to put the public API over any adapter: +A third-party file driver can be created with `defineFileDriver()`: ```ts import { createFileSystem } from "@okikio/opfs"; +import { createFileAdapter } from "@okikio/opfs/adapter/file"; +import { defineFileDriver } from "@okikio/opfs/driver/file"; -const fileSystem = createFileSystem(adapter, { - coordination: "auto", - maxBufferedWriteBytes: 64 * 1024 * 1024, +const driver = defineFileDriver(backend, { + name: "my-files", + capabilities: { + read: true, + write: true, + streamRead: true, + streamWriteModes: ["replace"], + rangeRead: true, + nativeCopy: false, + nativeMove: false, + positionalWrite: false, + syncAccess: false, + }, }); + +const fileSystem = createFileSystem(createFileAdapter(driver)); ``` -The first-party adapters cover three different storage shapes -------------------------------------------------------------- - -Native filesystems expose files and directories directly. Record stores expose values keyed by logical identity. Object stores -expose whole-object replacement, ranges, prefixes, and provider-side copy. Keeping those shapes separate is what prevents one -"universal" adapter from hiding important performance and consistency behavior. - -| Public subpath | Backend | Translation layer | -| --- | --- | --- | -| `adapter/opfs` | browser OPFS | native filesystem | -| `adapter/node` | Node `fs` | native filesystem | -| `adapter/deno` | Deno filesystem | native filesystem | -| `adapter/bun` | Bun + Node-compatible fs | native filesystem | -| `adapter/memory` | in-memory map | records | -| `adapter/record` | custom value/document store | records | -| `adapter/localstorage` | Web Storage | records | -| `adapter/indexeddb` | IndexedDB | records | -| `adapter/cache` | Cache Storage | records | -| `adapter/deno-kv` | Deno KV | records | -| `adapter/sqlite` | connected SQLite | records through db0-compatible SQL | -| `adapter/unstorage` | unstorage `Storage` | records | -| `adapter/rxdb` | RxDB `RxCollection` | records | -| `adapter/db0` | db0 `Database` | records | -| `adapter/drizzle` | Drizzle database + table | records | -| `adapter/object` | custom object store | objects | -| `adapter/s3` | direct S3/S3-compatible client | objects | -| `adapter/azure` | direct Azure Blob client | objects | - -Native browser and host filesystems ------------------------------------ - -`openFileSystem()` is the convenience path for native browser OPFS: +The backend must implement every capability it advertises. The adapter does not fabricate a native method from a flag. -```ts -import { openFileSystem } from "@okikio/opfs"; +### Browser OPFS -const fileSystem = await openFileSystem(); -``` +`createOpfsDriver(root)` owns native browser handles. `createOpfsAdapter(driver)` is the explicit translation. The +convenience `openFileSystem()` acquires `navigator.storage.getDirectory()`, creates the OPFS driver and adapter, then +creates the facade. -The explicit form is useful when the caller already owns the native root: +The driver retains the native root for advanced browser interop. Sync access is advertised only when the actual file +handle exposes the required method in the current realm. -```ts -import { createFileSystem } from "@okikio/opfs"; -import { createOpfsAdapter } from "@okikio/opfs/adapter/opfs"; +### Node -const root = await navigator.storage.getDirectory(); -const fileSystem = createFileSystem(createOpfsAdapter(root)); -``` +`createNodeDriver({ root })` maps virtual `/` below one host directory. The host-path mapper rejects escape from that +root. -The OPFS adapter reports synchronous access only when an actual file handle exposes `createSyncAccessHandle()`. The package -does not infer the feature from a browser name or from "worker" alone. +Node exposes: -Node, Deno, and Bun map virtual `/` below one configured host directory: +- materialized and streaming reads; +- ranged reads; +- materialized and streaming writes; +- native file copy; +- native rename/move; +- asynchronous positional files; +- synchronous random access and flush. -```ts -import { createNodeAdapter } from "@okikio/opfs/adapter/node"; +Node built-ins resolve only when the explicit Node driver is created/imported. The root module does not import Node +runtime code. -const adapter = createNodeAdapter({ root: "./data" }); -``` +### Deno + +`createDenoDriver({ root })` uses Deno file APIs for persistence and `@std/path` only for the host-root mapper. It +supports the same major file routes as the Node driver where Deno provides the native primitive. -The host path mapper resolves the configured root once and rejects every virtual path whose resolved host path would leave that -root. Host adapters expose ranges, streams, native copy, native move, and synchronous random access when the underlying runtime -provides them. +The driver requires filesystem permissions chosen by the host application. It does not request broad permissions itself. -Bun uses Bun's file APIs where they provide a direct read/write path and uses Bun's Node-compatible filesystem surface for the -operations whose exact semantics already live there. Importing the Bun adapter does not require the `Bun` global until adapter -creation. +### Bun -Record stores start small and can add byte lanes ------------------------------------------------ +`createBunDriver({ root })` uses `Bun.file()` and `Bun.write()` where they improve the direct read/replace path, then +delegates operations that need stronger host-filesystem semantics to the Node-compatible file driver. -The required `RecordStoreType` stays intentionally small: +The Bun global is resolved lazily during driver creation. Importing the module in Node or Deno does not require Bun. + +## Record drivers + +`RecordDriverType` is the native contract for value/document/database persistence. + +Required logical operations: ```ts -interface RecordStoreType { +interface RecordBackendType { get(path): Promise; set(record): Promise; delete(path): Promise; @@ -133,284 +137,251 @@ interface RecordStoreType { } ``` -This complete-record path is enough for memory, Web Storage, RxDB, unstorage, and SQL-backed integrations. File records use -base64 because the same durable shape must round-trip through JSON-oriented stores. Base64 is a compatibility format, not a claim -that every record backend is suitable for large binaries. - -A store with a more capable physical layout can add optional lanes without implementing the filesystem facade again: +Optional byte lanes can avoid the generic base64 fallback: ```text -stat metadata without file body -readFile direct/range byte read -openReadStream backpressure-preserving logical stream -writeFile selected direct materialized modes -writeStream selected direct stream modes +stat +readFile +openReadStream +writeFile +writeStream ``` -`RecordStoreCapabilitiesType` declares `rangeRead`, `streamRead`, `writeModes`, and `streamWriteModes`. The record adapter turns -only those declared lanes into native adapter capabilities. If a lane is absent, the complete-record implementation remains the -fallback. This is the extension point for third-party KV/document stores that can do better than one large JSON-shaped record. +Record capability metadata also identifies: -`createMemoryAdapter()` uses the complete-record path for deterministic tests and temporary data. +```text +replacement atomic | best-effort +binary native binary storage available +transactions driver has provider transaction behavior relevant to its writes +``` -`createLocalStorageAdapter(storage)` accepts an injected Web Storage object. Web Storage is synchronous, quota-limited, and -string-only underneath the adapter. It does not claim a portable maximum item size or native streaming. +A custom record driver uses `defineRecordDriver()` and then `createRecordAdapter()`. -`openIndexedDbAdapter()` can open its own IndexedDB database, while `createIndexedDbAdapter(database)` can borrow an existing -one. The store is keyed by canonical `path` and indexed by `parent` so direct directory listing stays indexed. Ownership remains -with the caller unless the adapter option explicitly transfers it. +The generic record adapter does not claim native streaming. If the driver does not provide `writeStream()`, the facade +can materialize an input only under `maxBufferedWriteBytes`. -`createCacheAdapter(cache)` stores records in an injected `Cache` using synthetic request URLs. No network request is made. Cache -Storage quota, eviction, and persistence policy remain browser decisions. +### Memory -Deno KV uses an explicit partition layout ------------------------------------------ +The memory driver is deterministic and dependency-free. It is useful for tests, examples, and temporary state. It is not +durable storage. -`createDenoKvAdapter(kv)` accepts an already-open Deno KV database. Deno KV has a 2 KiB serialized key limit and a 64 KiB -serialized value limit, so treating one filesystem file as one KV value would create a small and surprising file ceiling. The -default adapter policy is `partition: "auto"`. +### Deno KV + +The Deno KV driver has a provider-aware partition layout. Important policy options are: ```text -logical entry key - [prefix, "entry", parentPath, name] +partition auto | always | never +partBytes decoded bytes per raw part +inlineBytes decoded body budget for one inline record +maxParts logical-file part ceiling +concurrency bounded physical part I/O +``` -list one parent - prefix [prefix, "entry", parentPath] - -> direct children only +`partBytes` and `inlineBytes` are intentionally smaller than Deno KV's serialized value ceiling. The provider limit +applies after serialization, so accepting the full provider number as decoded application bytes would be unsafe. -small file - entry -> normal FileRecord +The driver planner also evaluates the concrete path against a conservative serialized-key estimate before provider I/O. -large file - [prefix, "part", canonicalPath, generation, 0] - [prefix, "part", canonicalPath, generation, 1] - ... - entry -> manifest committed last -``` +`DenoKvDriverType.collect()` performs explicit maintenance for crash-left physical generations. It scans only the +private part namespace, retains the published generation, ignores recent unpublished generations for a one-hour grace +period by default, and stops after the caller's deletion budget. Ordinary reads and writes never start this scan +implicitly. -Default decoded sizes are 32 KiB inline and 48 KiB per raw binary part. `maxParts` defaults to 10,000 and part I/O concurrency -to 8. Callers can set `partition: "never" | "auto" | "always"`, `inlineBytes`, `partBytes`, `maxParts`, and `concurrency`. The -adapter exposes these as inspectable limits/partition policy. +### localStorage -Manifest-last publication is the visibility rule. A reader sees the previous complete generation until all new parts exist and -the new manifest is stored. A process crash before the manifest commit can leave unreachable new-generation parts. That is a -storage leak, not a partially visible logical file. The adapter does not currently run a global orphan scavenger because doing so -would require a separate ownership/retention policy. +The localStorage driver maps canonical records into a private key prefix. It inherits Web Storage's synchronous +underlying API, but the package presents the normal asynchronous driver contract to keep the storage stack composable. -Large-file hot paths avoid generic reconstruction: +Applications should treat browser quota as dynamic. The driver does not invent a stable quota number. -- `stat()` reads the entry/manifest only. -- `list()` uses the direct-parent key prefix, so it reads direct-child metadata only and never scans descendant entry keys or body parts. -- range reads fetch only overlapping parts. -- stream reads load one physical part at a time under consumer backpressure. -- materialized append/update builds a new generation part-by-part and never joins the previous large file into one record. -- streamed replacement writes parts with bounded concurrency and publishes the manifest last. +### IndexedDB -Append/update still copy the untouched logical bytes into a new immutable generation because Deno KV has no provider-side range -copy primitive. The copy is bounded by `partBytes` and `concurrency`; the tradeoff is provider I/O proportional to the resulting -file size rather than JavaScript memory proportional to that size. Streamed append/update is not advertised as a direct lane, -so an incoming stream must still fit under the facade `maxBufferedWriteBytes` ceiling before this bounded patch path runs. +The IndexedDB driver borrows or owns an injected database according to options. It uses an object store and a parent +index for direct-child listing. The application remains responsible for database versioning/upgrades outside the driver +unless ownership is explicitly transferred. -`partition: "never"` disables the partitioned streaming write lane. A large streamed write then follows the facade's normal -bounded materialization rule and fails `too-large` once it exceeds `maxBufferedWriteBytes`. This gives applications a deliberate -way to reject the changed durable layout. Deno KV remains an unstable Deno API, so real integration tests run with -`--unstable-kv`. +### Cache Storage -`createSqliteAdapter(database)` accepts a small connected SQLite statement interface. It deliberately reuses the same SQL record -mapping used by the SQLite branch of `createDb0Adapter()` instead of creating a second schema and upsert implementation. +The Cache driver stores records under private request URLs. Cache Storage is a record/value persistence mechanism here, +not an HTTP cache policy abstraction. The driver only interprets entries in its private namespace. -Existing ecosystem adapters stay above the abstraction the application already owns: +### unstorage -```text -unstorage Storage -> RecordStoreType -RxDB RxCollection -> RecordStoreType -db0 Database -> RecordStoreType -Drizzle DB + table -> RecordStoreType -``` +`createUnstorageDriver(storage)` consumes the high-level unstorage `Storage` object. This deliberately sits above +whichever unstorage provider driver the application selected. -Bridge descriptors group these forward adapters with reverse drivers when a real reverse contract exists. They do not fabricate -a reverse direction for RxDB, db0, or Drizzle. See [ecosystems.md](./ecosystems.md). +Use `bridge/unstorage` for the opposite direction, where an existing `FileSystemType` must satisfy unstorage's Driver +contract. -Object stores keep object-store semantics visible -------------------------------------------------- +### RxDB -`ObjectStoreType` is the common client contract for S3, Azure Blob, and custom object storage. It models the operations those -systems actually have: +`createRxDbDriver(collection)` targets `RxCollection`, not `RxStorage`. RxDB keeps responsibility for its selected +RxStorage, conflict mechanics, wrappers, replication, multi-instance behavior, and licensing. -```text -HEAD exact key -GET full object or range -PUT replacement -DELETE exact key -LIST prefix + delimiter -COPY inside provider, when supported -``` +The package exports `RxDbRecordJsonSchema` for the collection used by this integration. `path` is the primary key and +`parent` is indexed for direct-child listing. -Its metadata includes byte size, media type, last modification time, ETag, provider version identity, and user metadata. Its -capabilities state whether range read, streaming read, streaming replacement, provider-side copy, and conditional writes are -really available. +### db0 -`createObjectAdapter()` maps that model to filesystem paths. A normal file maps to one object key. An empty directory maps to a -trailing-slash marker with private metadata, and prefix listing recognizes both those markers and foreign provider prefixes. +`createDb0Driver(database)` targets the db0 `Database` contract and its reported dialect. The current SQL branches are +SQLite, libSQL, PostgreSQL, and MySQL. -```text -/photos -> photos/ -/photos/a.jpg -> photos/a.jpg -/photos/2026/b.jpg -> photos/2026/b.jpg -``` +The driver owns its filesystem-record table only when configured to initialize it. The injected database is borrowed +unless `disposeDatabase` is true. -A raw object namespace can physically contain both `mixed` and `mixed/child`. The filesystem view resolves the exact `mixed` -object as the file because `stat("/mixed")` does the same. That rule keeps read, stat, and write behavior internally consistent -when foreign object layouts do not obey filesystem restrictions. +### Drizzle -Replacement can stream when the provider supports it. Append and update cannot normally mutate object bytes in place, so the -adapter performs: +`createDrizzleDriver({ database, table })` accepts a caller-owned connected Drizzle database and a caller-owned table +shape. Drizzle is not one SQL dialect, so this generic driver does not own DDL or migrations. + +Required logical columns are: ```text -HEAD current object - | - v -GET current bytes - | - v -apply append/update in memory - | - v -conditional PUT replacement +path +parent +name +kind +data +size +lastModified +mediaType ``` -When conditional writes are enabled and an existing object does not return an ETag, the adapter fails rather than quietly -performing an unsafe read-modify-write. When the file is being created through append/update, it uses create-only semantics where -the provider exposes them. +The portable replacement route is delete then insert, so the generic driver reports best-effort replacement rather than +claiming cross-process atomicity. A dialect-specific future driver can expose stronger transaction/upsert/binary +behavior. -The S3 client implements the protocol directly ------------------------------------------------ +### SQLite rows -`createS3Client()` uses Web Fetch, Web Crypto, `@std/encoding`, and `@std/xml`. It does not depend on the AWS SDK. +`createSqliteDriver(database)` stores filesystem rows inside an already connected SQLite database. This is the +**database-backed filesystem** direction. -```ts -import { createS3Client } from "@okikio/opfs/s3"; -import { createS3Adapter } from "@okikio/opfs/adapter/s3"; - -const client = createS3Client({ - endpoint: "https://s3.us-east-1.amazonaws.com", - bucket: "example", - region: "us-east-1", - credentials: async () => await credentials.get(), - addressing: "path", - concurrency: 4, -}); +It is not a SQLite VFS and does not make SQLite store its database file on `FileSystemType`. See `ecosystems.md` for +that opposite direction. -const adapter = createS3Adapter(client, { prefix: "app" }); -``` - -Signature Version 4 includes the request authority in canonical headers. Browser Fetch forbids application code from setting -the `Host` header, so the client signs `url.host` while leaving actual Host or `:authority` transmission to Fetch. Static browser -credentials are usually a security mistake; browser deployments should use appropriately scoped short-lived credentials or a -trusted service design. +## Object drivers -A streamed replacement uses multipart upload with bounded part concurrency: +`ObjectDriverType` preserves object storage concepts required for efficient translation: ```text -ReadableStream - | - v -fixed-size chunker - | - +--> UploadPart 1 --+ - +--> UploadPart 2 --+--> CompleteMultipartUpload - +--> UploadPart N --+ - | - failed operation - | - `--> wait active parts -> AbortMultipartUpload +stat object +get bytes/range +put bytes/stream +list prefix +remove object +native copy when available ``` -The client applies `If-Match` and `If-None-Match` to multipart completion, which is the commit operation that current S3 exposes -for these preconditions. It waits for already-started parts before aborting so a late part cannot arrive after the abort request. +An object driver can also report object-specific capability details, provider limits, continuation behavior, +partition/upload policy, and physical metrics. -S3 has two failure cases that are easy to miss. `CompleteMultipartUpload` can return HTTP 200 and then stream an XML error, and -`CopyObject` can also return an embedded error in HTTP 200. The client parses and rejects both bodies. +### S3 -`CopyObject` has a 5 GB source limit. Larger copies use `UploadPartCopy` ranges into a multipart destination. The copy part size -increases when necessary to stay within S3's 10,000-part limit. Source bytes stay inside the object provider rather than crossing -JavaScript memory or network twice. +The S3 client owns REST, SigV4, request policy, multipart operations, copy, listing, and protocol errors. -The filesystem copy contract intentionally preserves file bytes, media type, and user metadata where the client can do so. It -does not claim to clone every S3 control-plane property such as ACLs, tags, object-lock state, or every checksum policy. Use the -low-level signed `client.request()` API when the S3 object itself, rather than its filesystem view, is the thing being managed. +`createS3Driver(client)` adds backend capability/limit/optimization inspection. `createS3Adapter(driver)` translates +object keys and directory prefixes into filesystem primitives. -S3-compatible does not mean behavior-identical. Configure the client from the selected provider's current contract: +The client optimizations include independently controllable delayed multipart promotion and derived signing-key caching. +See `s3.md`. -- Cloudflare R2 commonly uses the `auto` region and has provider-specific supported/unsupported S3 operations. -- DigitalOcean Spaces implements a compatible subset rather than every AWS S3 feature. -- Google Cloud Storage's XML multipart API documents different precondition behavior; disable `conditionalWrite` when the - selected path does not provide the safety contract expected by the object adapter. -- Other compatible providers should be treated the same way: verify endpoint, signing region, addressing, copy, conditional - requests, multipart limits, checksums, and error behavior before enabling a capability flag. +### Azure Blob -The Azure Blob client keeps Azure's own model --------------------------------------------- +The Azure client owns Blob REST, authentication, block upload, server-side copy, listing, and provider errors. -`createAzureClient()` also uses Web Fetch and `@std/xml`, but it does not force Azure Blob through an S3-shaped client. +`createAzureDriver(client)` adds backend inspection. `createAzureAdapter(driver)` supplies filesystem translation. Block +upload and server copy are independently disableable. See `azure.md`. + +## Adapter contract + +`AdapterType` always references the driver it translates: ```ts -import { createAzureClient } from "@okikio/opfs/azure"; -import { createAzureAdapter } from "@okikio/opfs/adapter/azure"; +interface AdapterType { + readonly name: string; + readonly driver: DriverType; + readonly capabilities: AdapterCapabilitiesType; + // filesystem primitives... +} +``` -const client = createAzureClient({ - endpoint: "https://account.blob.core.windows.net", - container: "example", - credential: { kind: "sas", token }, -}); +`AdapterCapabilitiesType` describes native adapter routes, not every public filesystem operation. -const adapter = createAzureAdapter(client, { prefix: "app" }); +```text +read +write +streamRead +streamWriteModes +rangeRead +nativeCopy +nativeMove +positionalWrite +syncAccess ``` -Credentials can be SAS, refreshable bearer tokens, or a custom header callback. The service version is explicit and defaults to -the version pinned by this package. Block-size limits are selected from that service version rather than one timeless constant. +The facade can still emulate operations. `FileSystemType.inspect().support` is the authority for the effective route +after adapter capabilities and facade optimization policy are composed. -Streamed replacements use Put Block followed by Put Block List. Azure has no abort call for uncommitted blocks, so a failure -waits for already-started requests, leaves the old committed blob untouched, and allows Azure to garbage-collect the uncommitted -blocks later. +Adapters may retain compact `limits` or `partition` summaries for translation diagnostics. Detailed backend limits and +their provenance live on the driver. -Copy Blob From URL has a smaller synchronous copy limit. Larger files use Put Block From URL ranges followed by Put Block List. -This keeps large copies server-side without hiding a size cliff behind `nativeCopy: true`. +## Cancellation -Custom adapters must preserve the same invariants -------------------------------------------------- +Every async backend method that accepts `AbortSignal` checks it before expensive work and between bounded chunks. A +failed or aborted stream write cancels the upstream producer when practical. -Use `defineAdapter()` for a backend that already exposes filesystem-like primitives: +Provider cleanup can outlive the caller signal. Protocol drivers/clients use a separate bounded cleanup signal when an +already aborted caller signal would make cleanup impossible. -```ts -import { defineAdapter } from "@okikio/opfs/adapter"; +## Ownership -export const adapter = defineAdapter({ - name: "provider", - capabilities: { - read: true, - write: true, - streamRead: false, - streamWriteModes: [], - rangeRead: false, - nativeCopy: false, - nativeMove: false, - positionalWrite: false, - syncAccess: false, - }, - async stat(path) { /* ... */ }, - async readFile(path, options) { /* ... */ }, - async writeFile(path, bytes, options) { /* ... */ }, - async *readDir(path, options) { /* direct children only */ }, - async createDir(path, options) { /* parent already exists */ }, - async remove(path, options) { /* file or empty directory */ }, -}); -``` +Injected resources are borrowed by default. + +Examples of explicit ownership transfer: -Every adapter receives canonical virtual paths. It must respect requested ranges and write modes. It must yield direct children, -not recursive descendants, from `readDir()`. It must not configure logging, inspect process environment, or open unrelated -resources during module evaluation. +```text +disposeDatabase +disposeStorage +disposeDriver +disposeAdapter +``` -Injected resources are borrowed by default. Transfer ownership only through an explicit option such as `disposeDatabase`, -`disposeStore`, or `disposeAdapter`. A filesystem closing must never surprise another subsystem by disposing infrastructure that -it still owns. +A convenience adapter that creates a driver internally transfers ownership of that newly created driver to the adapter. +A caller that creates a driver explicitly can choose whether the adapter should dispose it. A driver exposes backend +disposal only when its own construction options transferred backend ownership; disposing an adapter therefore cannot +close a resource that the driver only borrowed. + +`driver.inspect().ownership` reports that relationship as `none`, `borrowed`, or `owned`. `driver.inspect().provides` +reports the stable backend operations or capabilities available on the configured driver. Higher layers can therefore +explain backend ownership and breadth without inferring either from adapter flags. + +A configured record driver can also be read-only. In that mode `driver.capabilities.write` is false, write primitives +are not exposed, and direct `set()`/`delete()` calls fail before backend mutation. The adapter reflects the same state +instead of relying on adapter-only policy. + +## Import safety + +Runtime and provider code stays behind explicit subpaths. Importing the package root does not: + +- import Node/Bun/Deno-only modules; +- resolve credentials; +- open a database; +- connect to a network endpoint; +- configure global logging; +- start worker/process resources. + +## Extension checklist + +Before adding a driver/adapter, verify: + +1. The driver is independently meaningful without `FileSystemType`. +2. Provider requirements and known limits are structured and attributed. +3. Unknown limits stay unknown rather than being treated as unlimited. +4. Observable optimizations can be disabled. +5. The driver planner can reject known bad inputs before I/O. +6. Every advertised direct operation has a real implementation. +7. Large work has explicit byte/part/concurrency/retry bounds. +8. Resource ownership is explicit. +9. The adapter contains translation, not duplicated provider behavior. +10. Tests exercise the driver directly and through the adapter/facade. +11. Benchmarks include the backend/client baseline and each added layer. diff --git a/docs/api.md b/docs/api.md index 62a4832..2c4fd67 100644 --- a/docs/api.md +++ b/docs/api.md @@ -1,626 +1,780 @@ -Public API guide -================ +# Public API guide -This guide is organized by developer task. Exact low-level schemas and types are also available through the explicit package subpaths. +## Purpose -Open or create a filesystem ---------------------------- +The public API is organized by layer. Normal application code can use the root filesystem facade and convenience +adapters. Storage libraries and infrastructure code can use the explicit client, driver, adapter, bridge, and +integration subpaths. + +```text +client -> driver -> adapter -> FileSystemType -> bridge +``` + +No public constructor requires a global registry. + +## Filesystem root ### `openFileSystem(options?)` -Opens the browser's native Origin Private File System and returns `FileSystemType`. +Opens the current browser realm's native Origin Private File System and returns `FileSystemType`. ```ts import { openFileSystem } from "@okikio/opfs"; -const fileSystem = await openFileSystem(); +await using fileSystem = await openFileSystem(); +await fileSystem.writeFile("/state.json", "{}", { parents: true }); ``` -Use this only when native browser OPFS is the chosen backend. Server runtimes should create a runtime adapter and pass it to `createFileSystem()`. +This convenience path composes native OPFS root -> OPFS driver -> OPFS adapter -> filesystem facade. ### `createFileSystem(adapter, options?)` Creates the adapter-independent facade. ```ts -const fileSystem = createFileSystem(adapter, { - coordination: "auto", - lockPrefix: "my-app:filesystem", - maxBufferedWriteBytes: 64 * 1024 * 1024, - metrics: "basic", - optimizations: { nativeCopy: true }, - disposeAdapter: false, -}); +import { createFileSystem } from "@okikio/opfs"; +import { createNodeAdapter } from "@okikio/opfs/adapter/node"; + +const fileSystem = createFileSystem( + createNodeAdapter({ root: "./data" }), + { + coordination: "local", + maxBufferedWriteBytes: 64 * 1024 * 1024, + metrics: "basic", + disposeAdapter: true, + }, +); ``` -`FileSystemOptionsType`: +Important `FileSystemOptionsType` fields: -- `coordination`: `auto`, `web-locks`, `local`, or `none`. -- `lockPrefix`: stable lock namespace used for cooperating filesystem facades. -- `maxBufferedWriteBytes`: maximum stream/file size the facade may materialize for a fallback route. -- `metrics`: `none`, `basic`, or `timing`. `none` is intended for baseline overhead measurements. -- `optimizations`: partial override for `streamRead`, `streamWrite`, `rangeRead`, `nativeCopy`, and `nativeMove`. Every route defaults to enabled. -- `disposeAdapter`: transfers adapter disposal ownership to the facade when true. +- `coordination`: `auto | web-locks | local | none`; +- `lockPrefix`: stable namespace for cooperating facade locks; +- `maxBufferedWriteBytes`: maximum facade-owned materialization for fallback routes; +- `metrics`: `none | basic | timing`; +- `optimizations`: independent facade route switches; +- `disposeAdapter`: transfer adapter ownership to the facade. -`coordination` is runtime-validated by `CoordinationModeSchema`. +## Filesystem methods +`FileSystemType` exposes the portable API: -Inspect and plan before I/O ---------------------------- +```text +getDirectoryHandle +getFileHandle +getFile +stat +exists +mkdir +ensureDir +ensureFile +readFile +readText +openReadStream +writeFile +readDir +walk +copy +move +remove +emptyDir +openWritableFile +openSyncFile +inspect +plan +getMetrics +close +root +``` + +All path methods use the canonical virtual namespace. Public input is normalized before adapter/driver calls. -### `inspect()` +## Read APIs -Returns a synchronous `InspectionType` for the configured filesystem stack. It contains: +### `readFile(path, options?)` + +Returns `Uint8Array`. + +Options include: ```text -adapter -native capabilities -effective support -known hard limits -optional partition layout -resolved optimization policy -maxBufferedWriteBytes -metrics mode and current metrics snapshot +at zero-based offset +length maximum bytes after at +signal cancellation ``` -Effective support uses `native | emulated | partitioned | unsupported`. Native capability flags never include facade emulation. -This makes a caller able to distinguish a fast provider range read from a full-read-and-slice fallback, or a Deno KV partitioned -stream from a complete-record materialization. +A driver/adapter with native range support can avoid materializing the complete file. Otherwise the facade can emulate +the range and reports that route through `inspect()`/`plan()`. -### `plan(input)` +### `readText(path, options?)` -Creates a deterministic `PlanType` without touching the backend. Supported operations are `read`, `write`, `copy`, and `move`. -For writes, `size` is the resulting logical file size and is used for `maxFileBytes` and partition-count checks. `inputBytes` is -the number of bytes supplied by the current write and is used for facade stream-buffer admission. For replace, omitted -`inputBytes` falls back to `size` because those values are normally equal. Append/update callers should provide both when known. +Reads file bytes and decodes text. UTF-8 is the default encoding. -```ts -const plan = fileSystem.plan({ - operation: "write", - source: "stream", - mode: "replace", - size: 200 * 1024 * 1024, - inputBytes: 200 * 1024 * 1024, -}); +### `openReadStream(path, options?)` -if (!plan.supported) throw new Error(plan.reasons.join(" ")); -``` +Returns `ReadableStream`. -An unknown limit remains unknown. The planner does not invent provider guarantees. For an unknown-size emulated stream, it warns -that `maxBufferedWriteBytes` remains the runtime admission limit. +When a native stream exists and `optimizations.streamRead` is enabled, the facade forwards it. Otherwise a readable +backend can be adapted to a materialized stream. `inspect().support.streamRead` identifies the effective route. -### `getMetrics()` +## Write APIs -Returns a detached `MetricsType`. `basic` tracks counts, failures, bytes, native/emulated/partitioned route counts, current -facade-owned buffered bytes, and peak buffered bytes. `timing` additionally records total and maximum durations. `none` keeps the -hot-path collector inactive. +### `writeFile(path, data, options?)` -Optimization controls can force a fallback for differential tests or application policy. A disabled optimization is never -relabelled native. If the fallback cannot satisfy the request within `maxBufferedWriteBytes`, runtime execution and `plan()` both -fail with the same `too-large`/unsupported condition when size is known. +Accepted input: -Path API --------- +```text +string +Blob +ArrayBuffer +ArrayBufferView +ReadableStream +AsyncIterable +``` -### `getDirectoryHandle(path, options?)` +Write modes: -Returns a package `DirectoryHandleType` for one directory. +```text +replace +append +update +``` -Options: +Important options: -- `create`: create exactly that directory when absent. -- `recursive`: create missing ancestors as well. -- `signal`: abort before commit. +```text +at +truncate +parents +mediaType +signal +``` -The virtual root `/` always exists. +A streaming source uses a driver-native stream route only when the selected write mode supports it and the route is +enabled. Otherwise the facade materializes the stream under `maxBufferedWriteBytes`. Crossing that limit cancels the +producer when possible and fails with `too-large`. -### `getFileHandle(path, options?)` +## Directory and tree APIs -Returns a package `FileHandleType`. +### `readDir(path, options?)` -Options: +Returns a lazy direct-child iterator. -- `create`: create the file when absent. -- `parents`: create missing parent directories. -- `signal`: abort before commit. +### `walk(path, options?)` -A read-only lookup never creates a file or directory. +Returns a lazy recursive iterator. Options control depth, root inclusion, file inclusion, directory inclusion, and +cancellation. The facade does not collect the complete tree first. -### `getFile(path, options?)` +### `copy(source, destination, options?)` -Returns a `File` snapshot. Changes written later are not reflected in the already-returned File object. +Copies one file or directory tree. The facade uses native/server-side copy when the adapter provides it and the route is +enabled. Otherwise it composes read/write work with bounded concurrency. -### `stat(path, options?)` +### `move(source, destination, options?)` -Returns: +Uses native move when available. The fallback is copy then remove and is not atomic. `plan()` returns a structured +warning for that route. -```ts -type StatType = FileStatType | DirectoryStatType; -``` +### `remove(path, options?)` -File stat includes canonical path, name, size, last-modified milliseconds, and media type. Directory stat includes canonical path, name, and last-modified when the adapter provides it. +Removes one entry or recursively removes descendants when requested. The virtual root cannot be removed. -### `exists(path, options?)` +### `emptyDir(path?, options?)` -Returns an advisory boolean. `kind` can restrict the answer to `file` or `directory`. +Removes children while keeping the directory. Root is the default. -Do not use `exists()` as a substitute for operation error handling. Another context can mutate the backend after the check. +## OPFS-shaped handles -### `mkdir(path, options?)` +Every facade exposes `root: DirectoryHandleType`. -Creates one directory. `recursive: true` creates missing ancestors. +`DirectoryHandleType`: -### `ensureDir(path, options?)` +```text +kind +name +path +getDirectoryHandle +getFileHandle +removeEntry +resolve +entries +keys +values +isSameEntry +Symbol.asyncIterator +``` -Ensures a directory and its parents exist. A file at the same path produces `type-mismatch`. +`FileHandleType`: -### `ensureFile(path, options?)` +```text +kind +name +path +getFile +createWritable +createSyncAccessHandle +isSameEntry +``` -Ensures an empty file exists and creates its parent directories. +These are package facades, not native browser handle instances. `path` is a package-specific canonical virtual path. -Read APIs ---------- +## Writable files -### `readFile(path, options?)` +`createWritable()` returns `WritableFileStreamType` with OPFS-style commands: -Returns `Uint8Array`. +```ts +await writable.write(bytes); +await writable.write({ type: "write", position: 10, data: bytes }); +await writable.write({ type: "seek", position: 20 }); +await writable.write({ type: "truncate", size: 100 }); +await writable.close(); +``` -Options: +The staged image commits on close and is discarded on abort. Large sequential writes should prefer `writeFile()` because +that path can select a native streaming driver route. -- `at`: zero-based byte offset. -- `length`: maximum bytes after `at`. -- `signal`: cancellation. +`openWritableFile()` exposes a direct asynchronous positional resource only when the adapter reports that capability. -### `readText(path, options?)` +## Synchronous files -Reads bytes and decodes them. `encoding` defaults to UTF-8. +`openSyncFile(path, options?)` returns `SyncFileType` only when the selected adapter exposes native synchronous random +access. -### `openReadStream(path, options?)` +Operations: -Returns `ReadableStream`. +```text +read +write +writeAll +getSize +truncate +flush +close +``` -If the adapter provides native stream reads, the facade forwards them. Otherwise it creates a stream from the adapter's materialized read result. Cancellation remains connected after the stream opens. +The facade keeps the path lock for the complete resource lifetime. `writeAll()` handles partial native writes. -Write API ---------- +## Inspection -### `writeFile(path, data, options?)` +### `fileSystem.inspect()` -Accepted `WriteDataType` values: +Returns `InspectionType`: -```text -string -Blob -ArrayBuffer -ArrayBufferView -ReadableStream -AsyncIterable +```ts +const inspection = fileSystem.inspect(); + +inspection.driver; +inspection.adapter; +inspection.support; +inspection.optimizations; +inspection.maxBufferedWriteBytes; +inspection.metricsMode; +inspection.metrics; +inspection.driverMetrics; ``` -Options: +`inspection.driver` contains: -- `mode`: `replace` (default), `append`, or `update`. -- `at`: starting byte offset for update mode. -- `truncate`: truncate at the final cursor. -- `parents`: create missing parents. -- `mediaType`: metadata for record/native adapters that can preserve it. -- `signal`: cancellation. +- `provides`: stable operation/capability identifiers exposed by this configured driver; +- `ownership`: `none`, `borrowed`, or `owned` for the long-lived backend resource; +- backend-native requirements and their current availability; +- provenance-aware provider, implementation, user, and probe limits; +- independently controllable driver optimization state. -The mode is runtime-validated by `WriteModeSchema`. +`provides` is intentionally an open string vocabulary. A third-party driver can add a provider-specific operation +without waiting for a core enum revision. The stable core driver families still expose typed operational methods +separately. -A non-streaming adapter buffers stream input up to `maxBufferedWriteBytes`. Crossing the limit cancels the producer and throws `too-large`. +`inspection.adapter` contains the translation layer's native route flags and compact translation summaries. -Long-lived positional output ----------------------------- +`inspection.support` contains effective routes after facade fallback and optimization policy: -### `openWritableFile(path, options?)` +```text +native +emulated +partitioned +unsupported +``` -Returns `WritableFileType` only when `adapter.capabilities.positionalWrite` is true. The facade does not emulate this operation with repeated `writeFile(..., { mode: "update" })` calls because a record-backed adapter can otherwise rematerialize the complete file for every chunk. +`inspection.metrics` is logical facade work. `inspection.driverMetrics`, when present, is physical backend/protocol +work. -Options: +## Planning -- `create`: create an empty file when it is absent. -- `parents`: create missing parent directories when `create` is true. -- `signal`: cancel ordinary work before commit. Cleanup through `close()` or `abort()` still releases the owned resource after cancellation. +### `fileSystem.plan(input)` -The returned resource owns the file mutation lock for its complete lifetime. The create/check/open sequence occurs under that same lock, so another mutation cannot enter between file creation and adapter open. +Preflights a concrete request without touching storage. + +Supported high-level operations are currently: + +```text +read +write +copy +move +``` + +Example: ```ts -const file = await fileSystem.openWritableFile("/media/output.mp4", { - create: true, - parents: true, +const plan = fileSystem.plan({ + operation: "write", + path: "/archive.bin", + source: "stream", + size: 800 * 1024 * 1024, + inputBytes: 800 * 1024 * 1024, + mode: "replace", }); -try { - await file.write(header, { at: 0 }); - await file.write(chunk, { at: chunkOffset }); - await file.flush(); - await file.close(); -} catch (error) { - await file.abort(error); - throw error; +if (!plan.supported) { + console.log(plan.problems); + console.log(plan.actions); } ``` -`write()` is positional, `truncate()` changes byte length, and `flush()` requests backend durability without closing. `close()` and `abort()` are idempotent terminal operations. A browser OPFS writable can discard its staged image on abort. Host filesystems generally cannot roll back bytes already written, so an application that needs publish-on-success semantics should write a staging path and move it after close. +`PlanType` includes: -Record/database adapters report `positionalWrite: false`. They remain valid for ordinary materialized writes and bounded stream buffering, but they are not presented as a large-file positional output path. +```text +operation +supported +support +driver +bufferBytes +partBytes +parts +problems[] +actions[] +``` -Directory iteration -------------------- +Problems have a stable code, layer, severity, message, and optional referenced limit. Actions have a stable kind and +optional code/detail. -### `readDir(path, options?)` +## Driver API -Lazy direct-child iterator. +`@okikio/opfs/driver` exports the generic driver definition model: -```ts -for await (const entry of fileSystem.readDir("/projects")) { - console.log(entry.kind, entry.name, entry.path); -} -``` +- `ProblemLayerSchema` / `ProblemLayerType`; +- `ProblemSeveritySchema` / `ProblemSeverityType`; +- `ActionKindSchema` / `ActionKindType`; +- `ProblemSchema` / `ProblemType`; +- `ActionSchema` / `ActionType`; +- `DriverOperationSchema` / `DriverOperationType`; +- `DriverPlanInputSchema` / `DriverPlanInputType`; +- `DriverPlanSchema` / `DriverPlanType`; +- `DriverInspectionSchema` / `DriverInspectionType`; +- `DriverType`; +- `DefineDriverOptionsType`; +- `defineDriver()`. -### `walk(path, options?)` +`defineDriver()` validates definition metadata. Concrete storage should normally use one of the family contracts below. -Lazy recursive iterator. +## File-driver API -Options: +`@okikio/opfs/driver/file` exports: -- `maxDepth`: maximum depth below the requested path. -- `includeRoot`: include the requested path itself before descendants. -- `includeFiles`: yield files. Defaults to true. -- `includeDirectories`: yield directories. Defaults to true. -- `signal`: cancel traversal between yielded entries. +- `FileDriverCapabilitiesSchema` / `FileDriverCapabilitiesType`; +- `FileDriverSignalOptionsType`; +- `FileDriverReadOptionsType`; +- `FileDriverWriteOptionsType`; +- `FileDriverCopyOptionsType`; +- `FileDriverMoveOptionsType`; +- `FileDriverDirectoryEntryType`; +- `FileDriverFileStatType`; +- `FileDriverDirectoryStatType`; +- `FileDriverStatType`; +- `FileDriverWritableFileType`; +- `FileDriverSyncFileType`; +- `FileDriverType`; +- `FileBackendType`; +- `DefineFileDriverOptionsType`; +- `defineFileDriver()`. -The iterator does not eagerly collect the entire tree. +The associated adapter is `createFileAdapter(driver)` from `@okikio/opfs/adapter/file`. -Structural operations ---------------------- +## Record-driver API -### `copy(source, destination, options?)` +`@okikio/opfs/driver/record` exports: -Copies one file or tree. Directory file bodies use bounded `concurrency`, default 4. +- `RecordListType`; +- `RecordReplacementSchema` / `RecordReplacementType`; +- `RecordDriverCapabilitiesSchema` / `RecordDriverCapabilitiesType`; +- `RecordBackendType`; +- `RecordDriverType`; +- `DefineRecordDriverOptionsType`; +- `defineRecordDriver()`. -`overwrite: true` replaces the destination tree instead of merging stale entries into it. +The associated adapter is `createRecordAdapter(driver)` from `@okikio/opfs/adapter/record`. -Source and destination cannot be the same path or ancestors of each other. +## Object-driver API -### `move(source, destination, options?)` +`@okikio/opfs/driver/object` exports: -Uses adapter-native move when `nativeMove` is true. Otherwise calls copy then remove. The fallback is not atomic. +- `ObjectDriverCapabilitiesSchema` / `ObjectDriverCapabilitiesType`; +- `ObjectStatType`; +- `ObjectEntryType`; +- `ObjectListType`; +- `ObjectGetOptionsType`; +- `ObjectPutOptionsType`; +- `ObjectCopyOptionsType`; +- `ObjectListOptionsType`; +- `ObjectBackendType`; +- `ObjectDriverType`; +- `DefineObjectDriverOptionsType`; +- `defineObjectDriver()`. -### `remove(path, options?)` +The associated adapter is `createObjectAdapter(driver)` from `@okikio/opfs/adapter/object`. -Removes one file or empty directory. `recursive: true` removes descendants first. +## Adapter API -The virtual root cannot be removed. +`@okikio/opfs/adapter` exports: -### `emptyDir(path?, options?)` +- `AdapterType`; +- `FileSystemOptionsType`; +- `defineAdapter()`. -Removes children while retaining the directory. `path` defaults to `/`. Child removals use bounded concurrency. +Every `AdapterType` has a `driver` reference. Adapter capability flags describe direct translation routes. They do not +replace `driver.inspect()` or `FileSystemType.inspect()`. -Synchronous file API --------------------- +## First-party file constructors -### `openSyncFile(path, options?)` +### Browser OPFS -Returns `SyncFileType` only when `adapter.capabilities.syncAccess` is true. +```ts +import { createOpfsDriver } from "@okikio/opfs/driver/opfs"; +import { createOpfsAdapter } from "@okikio/opfs/adapter/opfs"; -Options: +const root = await navigator.storage.getDirectory(); +const driver = createOpfsDriver(root); +const adapter = createOpfsAdapter(root); // convenience path creates its own driver +``` -- `create` -- `parents` -- `signal` +`createOpfsDriver(root)` returns `OpfsDriverType`. `createOpfsAdapter(root)` returns `OpfsAdapterType` and retains +`nativeRoot`. -The resource owns its native file and the facade path lock for the complete lifetime. +### Node ```ts -const file = await fileSystem.openSyncFile("/db.sqlite", { - create: true, - parents: true, -}); +createNodeDriver({ root: "./data" }); +createNodeAdapter({ root: "./data" }); +``` -try { - file.writeAll(bytes, { at: 0 }); - file.flush(); -} finally { - file.close(); -} +Types: `NodeDriverOptionsType`, `NodeAdapterOptionsType`. + +### Deno + +```ts +createDenoDriver({ root: "./data" }); +createDenoAdapter({ root: "./data" }); ``` -`SyncFileType` operations: +Types: `DenoDriverOptionsType`, `DenoAdapterOptionsType`. -```text -read -write -writeAll -getSize -truncate -flush -close +### Bun + +```ts +createBunDriver({ root: "./data" }); +createBunAdapter({ root: "./data" }); ``` -`writeAll()` loops over partial native writes. +Types: `BunDriverOptionsType`, `BunAdapterOptionsType`. -OPFS-shaped handle API ----------------------- +## First-party record constructors -Every `FileSystemType` has `root: DirectoryHandleType`. +### Memory -### Directory handle +```ts +createMemoryDriver(); +createMemoryAdapter(); +``` -```text -kind -name -path -getDirectoryHandle() -getFileHandle() -removeEntry() -resolve() -entries() -keys() -values() -isSameEntry() -[Symbol.asyncIterator]() +### Deno KV + +```ts +const driver = createDenoKvDriver(database, { + partition: "auto", + partBytes: 48 * 1024, + inlineBytes: 32 * 1024, + maxParts: 10_000, + concurrency: 8, +}); ``` -### File handle +Exports include: ```text -kind -name -path -getFile() -createWritable() -createSyncAccessHandle() -isSameEntry() +DENO_KV_MAX_KEY_BYTES +DENO_KV_MAX_VALUE_BYTES +DENO_KV_MAX_ATOMIC_BYTES +DENO_KV_SAFE_PART_BYTES +DENO_KV_SAFE_INLINE_BYTES +DENO_KV_DEFAULT_PART_BYTES +DENO_KV_DEFAULT_INLINE_BYTES +DENO_KV_DEFAULT_MAX_PARTS +DENO_KV_DEFAULT_CONCURRENCY +DENO_KV_DEFAULT_COLLECT_AGE_MS +DENO_KV_DEFAULT_COLLECT_DELETES +DenoKvEntryType +DenoKvType +DenoKvDriverOptionsType +DenoKvCollectOptionsType +DenoKvCollectResultType +DenoKvDriverType ``` -These are package facades, not native browser handle instances. Their `path` property is package-specific. +The specialized `DenoKvDriverType` also exposes `collect(options?)` for bounded, age-gated reclamation of unreachable +physical parts. The adapter path exports the same provider constants and maintenance types plus `createDenoKvAdapter()` +and `DenoKvAdapterOptionsType`. -### `createWritable()` +### localStorage -Returns `WritableFileStreamType`. The staged image commits on close and discards on abort. +`driver/localstorage` exports `LocalStorageType`, `LocalStorageDriverOptionsType`, and `createLocalStorageDriver()`. +`adapter/localstorage` exports `createLocalStorageAdapter()`. -Supported write commands: +### IndexedDB -```ts -await writable.write(data); -await writable.write({ type: "write", position: 10, data }); -await writable.write({ type: "seek", position: 20 }); -await writable.write({ type: "truncate", size: 100 }); -``` +`driver/indexeddb` exports `IndexedDbDriverOptionsType`, `IndexedDbOpenOptionsType`, and `createIndexedDbDriver()`. +`adapter/indexeddb` exports `createIndexedDbAdapter()`. -Blob also has a `type` property, so the implementation identifies a command only when `type` is exactly `write`, `seek`, or `truncate`. +### Cache Storage -Adapter and storage API ------------------------ +`driver/cache` exports `CacheDriverOptionsType` and `createCacheDriver()`. `adapter/cache` exports +`createCacheAdapter()`. -`@okikio/opfs/adapter` exports the backend contract consumed by `createFileSystem()`. The required primitive set stays small: +### unstorage -```text -stat -readFile -writeFile -readDir -createDir -remove -``` +`driver/unstorage` exports `UnstorageStorageType`, `UnstorageDriverOptionsType`, and `createUnstorageDriver()`. +`adapter/unstorage` exports `createUnstorageAdapter()`. + +### RxDB + +`driver/rxdb` exports the structural collection/document/query contracts, `RxDbRecordJsonSchema`, and +`createRxDbDriver()`. `adapter/rxdb` exports `createRxDbAdapter()`. + +### db0 -Optional methods expose stronger native paths: +`driver/db0` exports `Db0PrimitiveType`, statement/database contracts, `Db0DriverOptionsType`, and `createDb0Driver()`. +The constructor is asynchronous because table initialization can perform database I/O. + +`adapter/db0` exports `createDb0Adapter()`. + +### Drizzle + +`driver/drizzle` exports: ```text -openReadStream -> capabilities.streamRead -writeStream -> mode is present in capabilities.streamWriteModes -copy -> capabilities.nativeCopy -move -> capabilities.nativeMove -openWritableFile -> capabilities.positionalWrite -openSyncFile -> capabilities.syncAccess +DrizzleTableType +DrizzleRowType +DrizzleDriverOptionsType +createDrizzleDriver ``` -The exported contract includes `AdapterCopyOptionsType` as well as the signal, read, write, move, stat, directory-entry, -writable-file, sync-file, adapter, and filesystem option types. `defineAdapter()` validates the adapter name and capability -record. It does not add a global registry or change the adapter. +`adapter/drizzle` exports `createDrizzleAdapter()`. -The capability record describes what the backend performs natively. For example, an object adapter can expose -`streamWriteModes: ["replace"]` because replacement can stream to multipart/block upload while append and update still need a -read-modify-write cycle. The facade can emulate operations, but it does not relabel an emulation as native support. +### SQLite -### Record storage +`driver/sqlite` exports: -`@okikio/opfs/adapter/record` is the common translation point for backends that naturally store values, documents, or SQL rows. -It exports `RecordStoreType`, `RecordAdapterOptionsType`, and `createRecordAdapter()`. +```text +SqliteStatementType +SqliteDatabaseType +SqliteDriverOptionsType +createSqliteDriver +``` -A record store always supports the complete logical `get/set/delete/list` contract. Simple stores can stop there. More capable -value stores can additionally expose `stat`, direct range `readFile`, `openReadStream`, selected direct materialized write modes, -and selected `writeStream` modes through `RecordStoreCapabilitiesType`. `createRecordAdapter()` translates only the declared -lanes into native adapter capabilities. +`adapter/sqlite` exports `createSqliteAdapter()`. -The portable complete-record fallback uses the versioned `RecordType` union and base64 file bytes so JSON, Web Storage, RxDB, -unstorage, and SQL text columns share one representation. A specialized store such as Deno KV can keep large body parts as raw -binary values and use metadata-only/range/stream/direct-write lanes so the generic base64 representation is not on its -large-file hot path. Deno KV declares materialized replace/append/update as direct store modes; append/update rebuild the next -immutable generation part-by-part instead of reconstructing the previous complete logical file. Its stream lane remains -replace-only, so streamed append/update can still require facade input buffering. +## Object clients and drivers -### Object storage +### S3 client -`@okikio/opfs/adapter/object` exports the provider-neutral object-storage layer: +`@okikio/opfs/s3` exports: -- `ObjectCapabilitiesSchema` / `ObjectCapabilitiesType` -- `ObjectStatType` and `ObjectEntryType` -- object GET, PUT, COPY, and LIST option types -- `ObjectStoreType` -- `ObjectAdapterOptionsType` -- `createObjectAdapter()` +- `S3AddressingSchema` / `S3AddressingType`; +- `S3CredentialsSchema` / `S3CredentialsType`; +- `S3CredentialSourceType`; +- `S3_LIMITS`; +- `S3ClientOptionsType`; +- `S3RequestOptionsType`; +- `S3CompleteOptionsType`; +- `S3UploadType`; +- `S3PartType`; +- `S3Error`; +- `S3ClientType`; +- `createS3Client()`. -`ObjectStoreType` preserves object concepts such as ETags, provider version IDs, user metadata, prefix listing, conditional -writes, and server-side copy. `createObjectAdapter()` then maps that model into files and directories. Empty directories use -trailing-slash marker objects, while ordinary prefix listing also recognizes directories created outside this library. +Important client optimization options: -The direct S3 API lives at `@okikio/opfs/s3`: +```text +delayedMultipart +signingKeyCache +``` + +The driver paths are: ```ts -import { createS3Client } from "@okikio/opfs/s3"; -import { createS3Adapter } from "@okikio/opfs/adapter/s3"; - -const client = createS3Client({ - endpoint, - bucket, - region, - credentials, -}); -const adapter = createS3Adapter(client); +createS3Driver(options); +createS3DriverFromClient(client); ``` -`S3ClientType` extends `ObjectStoreType` and also exposes signed `request()`, `createUpload()`, `uploadPart()`, -`completeUpload()`, and `abortUpload()`. The lower-level request method is the deliberate escape hatch for provider-specific -S3 features that do not belong in the filesystem API. - -The Azure counterpart lives at `@okikio/opfs/azure` and `@okikio/opfs/adapter/azure`: +The convenience adapter is: ```ts -import { createAzureClient } from "@okikio/opfs/azure"; -import { createAzureAdapter } from "@okikio/opfs/adapter/azure"; - -const client = createAzureClient({ endpoint, container, credential }); -const adapter = createAzureAdapter(client); +createS3Adapter(client, options?); ``` -`AzureClientType` retains the REST request escape hatch, provider request IDs, range access, block upload, and server-side copy. +For explicit layering, use `createObjectAdapter(createS3DriverFromClient(client))`. -The object-store interfaces are intentionally smaller than either provider protocol. See [S3 client protocol](./s3.md) and -[Azure Blob client protocol](./azure.md) for signing, version gates, multipart/block lifecycles, limits, provider failures, and -known unsupported operations. +### Azure Blob client -### Reverse key-value APIs +`@okikio/opfs/azure` exports: -`@okikio/opfs/driver/kv` exports `createKeyValueDriver()`. It maps colon-delimited keys onto private directories so both `foo` -and `foo:bar` can exist at the same time: +- `AZURE_STORAGE_VERSION`; +- `AzureStorageVersionSchema` / `AzureStorageVersionType`; +- `AzureCredentialType`; +- `AzureClientOptionsType`; +- `AzureRequestOptionsType`; +- `AzureClientType`; +- `AzureError`; +- `AZURE_LIMITS`; +- `createAzureClient()`. + +Important optimization options: ```text -foo -> /key-foo/value -foo:bar -> /key-foo/key-bar/value +blockUpload +serverCopy ``` -The driver supports string/raw get and set, existence, metadata, hierarchical key enumeration, clear, and explicit filesystem -ownership transfer. It also exposes `inspect()`, `plan()`, and `getMetrics()`. Those methods delegate to the backing -`FileSystemType`, so a reverse ecosystem consumer sees the same effective routes, limits, partition policy, buffer ceiling, and -observed metrics instead of receiving a second approximation of storage capability. +The driver paths are: -`@okikio/opfs/driver/unstorage` is a thin translation over this generic driver. It supplies unstorage method names and -`maxDepth` behavior while reusing the same collision-safe filesystem mapping. +```ts +createAzureDriver(options); +createAzureDriverFromClient(client); +``` -### Schemas +The convenience adapter is `createAzureAdapter(client, options?)`. -`@okikio/opfs/schema` exports the executable project data contracts: +## Bridge API -- `PathSchema` / `PathType` -- `AdapterNameSchema` / `AdapterNameType` -- `EntryKindSchema` / `EntryKindType` -- `OpfsContextSchema` / `OpfsContextType` -- `CoordinationModeSchema` / `CoordinationModeType` -- `WriteModeSchema` / `WriteModeType` -- `AdapterCapabilitiesSchema` / `AdapterCapabilitiesType` -- `ErrorCodeSchema` / `ErrorCodeType` -- `RecordVersionSchema` / `RecordVersionType` -- `DirectoryRecordSchema` / `DirectoryRecordType` -- `FileRecordSchema` / `FileRecordType` -- `RecordSchema` / `RecordType` -- `Db0DialectSchema` / `Db0DialectType` -- `SqlIdentifierSchema` / `SqlIdentifierType` +`@okikio/opfs/bridge/kv` exports: -The package exports Zod 4 schemas directly. Zod 4 implements Standard Schema, so a consumer that accepts that interface can use -the same schema value without a second OPFS-owned wrapper contract. +- `KeyValueMetaType`; +- `KeyValueBridgeOptionsType`; +- `KeyValueBridgeType`; +- `createKeyValueBridge()`. -### Bridge descriptors +The bridge includes `inspect()`, `plan()`, and `getMetrics()` so consumers can reason about the storage stack beneath +the KV projection. -`@okikio/opfs/bridge` groups the adapter and reverse-driver directions for an ecosystem. It does not replace either primitive. +`@okikio/opfs/bridge/unstorage` exports: -- `UnstorageBridge`: both directions. -- `RxDbBridge`: collection to OPFS only. -- `Db0Bridge`: database to OPFS only. -- `DrizzleBridge`: database/table to OPFS only. -- `KeyValueBridge`: OPFS to generic asynchronous key-value only. +- `UnstorageBridgeMetaType`; +- `UnstorageBridgeTransactionOptionsType`; +- `UnstorageBridgeType`; +- `UnstorageBridgeOptionsType`; +- `createUnstorageBridge()`. -`defineBridge()` validates that `directions.toOpfs/fromOpfs` agree with the constructors. Every unsupported direction must state a -reason. This is intentional for ecosystems where the reverse shape would require query, conflict, transaction, synchronous, or -other semantics a filesystem does not own. +## Integration API -Path utility API ----------------- +`@okikio/opfs/integration/definition` exports: -`@okikio/opfs/path` exposes the canonical virtual-path model used by adapters: +- `IntegrationDirectionSchema` / `IntegrationDirectionType`; +- `IntegrationDirectionsSchema` / `IntegrationDirectionsType`; +- `IntegrationType`; +- `defineIntegration()`. -- `ROOT_PATH`: the canonical `/` root. -- `normalizePath(path)`: resolves `.`, `..`, duplicate separators, and relative input while rejecting root escape, backslashes, and NUL. -- `splitPath(path)`: returns canonical path segments without `/`. -- `joinPath(...parts)`: joins inputs and returns a canonical `PathType`. -- `dirname(path)`: returns the canonical parent path. -- `basename(path)`: returns the final name. Root returns an empty string. -- `isAncestorPath(ancestor, path)`: tests strict ancestry after normalization. -- `validateName(name)`: validates one direct File System API child name. -- `PathType`: validated canonical virtual path type. +`@okikio/opfs/integration` exports current first-party direction definitions: -Use the high-level filesystem methods for normal application work. These helpers are primarily for adapters, drivers, and code that persists canonical paths. +```text +UnstorageIntegration +RxDbIntegration +Db0Integration +DrizzleIntegration +DrizzleIntegrationSourceType +``` -Error API ---------- +Direction metadata never makes an unsupported reverse contract executable. -The root module exports `FileSystemError`, `getErrorName()`, `getErrorMessage()`, and `toFileSystemError()`. +## Schema and path API -`FileSystemError` carries: +`@okikio/opfs/schema` owns project serializable schemas and their inferred types. Important groups include: ```text -code stable ErrorCodeType -operation filesystem operation that failed -path canonical path when one exists -cause original runtime/provider failure when retained +PathSchema / PathType +WriteModeSchema / WriteModeType +CoordinationModeSchema / CoordinationModeType +AdapterCapabilitiesSchema / AdapterCapabilitiesType +SupportModeSchema / SupportModeType +MetricsModeSchema / MetricsModeType +PartitionModeSchema / PartitionModeType +RecordSchema / RecordType +DriverKindSchema / DriverKindType +LimitKindSchema / LimitKindType +LimitSourceSchema / LimitSourceType +LimitUnitSchema / LimitUnitType +LimitSchema / LimitType +RequirementStateSchema / RequirementStateType +RequirementSchema / RequirementType +DriverOptimizationSchema / DriverOptimizationType ``` -`toFileSystemError()` maps known DOMException names and server error codes such as `ENOENT`, `EEXIST`, and quota/permission failures into the stable package categories. Unknown provider failures remain `unknown` and retain the original cause. +`@okikio/opfs/path` exposes canonical virtual-path helpers: -`getErrorName()` and `getErrorMessage()` are safe extraction helpers for diagnostics where the caught value is `unknown`. - -Browser capability APIs ------------------------ +```text +ROOT_PATH +normalizePath +splitPath +joinPath +dirname +basename +isAncestorPath +validateName +PathType +``` -### `probeOpfs()` +## Error API -Returns a non-throwing `OpfsCapabilitiesType` report with: +The root module exports: -- execution context; -- root availability and normalized root error; -- embedded/same-origin-top facts when observable; -- Web Locks availability; -- sync access exposure; -- storage estimate when available; -- persistence status when available. +```text +FileSystemError +getErrorName +getErrorMessage +toFileSystemError +``` -It does not report `isIncognito` or `isPrivate`. +`FileSystemError` carries stable code, operation, optional canonical path, and original cause. -### `getOpfsContext()` +## Browser capability API -Classifies the current browser execution context as window, dedicated worker, shared worker, service worker, generic worker, or unknown. +The root exports `probeOpfs()` and `getOpfsContext()`. -### iframe subpath +`probeOpfs()` returns a non-throwing current-realm report. It probes actual APIs/policy rather than maintaining a +browser-brand table. `@okikio/opfs/iframe` exports: -- `supportsUnpartitionedOpfsRequest()` -- `requestUnpartitionedFileSystem()` +```text +supportsUnpartitionedOpfsRequest +requestUnpartitionedFileSystem +``` -The request is explicit because browser permission/user-activation requirements must remain under application control. +The application remains responsible for the permission/user-activation flow. -Lifecycle ---------- +## Metrics API -`FileSystemType` implements `AsyncDisposable`. +`@okikio/opfs/metrics` exports logical facade metrics plus `DriverMetricsType`. -```ts -await fileSystem.close(); -``` +Logical facade metrics count operations, failures, logical bytes, native/emulated/partitioned routes, buffering, and +optional timing. + +Driver metrics are physical and provider-specific enough to include request/retry/part/physical-byte/cleanup information +without forcing that data into logical filesystem counters. + +## Lifecycle -or with supported explicit resource management syntax: +`FileSystemType` implements async disposal. Closing is idempotent. The adapter is disposed only when ownership was +transferred. Each adapter/driver/bridge has its own explicit ownership option for injected resources. ```ts await using fileSystem = createFileSystem(adapter, { @@ -628,4 +782,5 @@ await using fileSystem = createFileSystem(adapter, { }); ``` -Closing is idempotent. The adapter is closed only when ownership was explicitly transferred. +Cancellation and disposal remain distinct. A caller can cancel one operation without implicitly disposing a shared +storage resource. diff --git a/docs/azure.md b/docs/azure.md index 465cab5..3bf6d92 100644 --- a/docs/azure.md +++ b/docs/azure.md @@ -1,23 +1,22 @@ -Azure Blob client protocol guide -================================ +# Azure Blob client protocol guide -Purpose -------- +## Purpose -This document defines the Azure Blob Storage REST contract implemented by -`@okikio/opfs/azure`. It is intended for maintainers changing authentication, -service-version behavior, block upload, copy, conditional replacement, listing, -or Azurite interoperability. +This document defines the Azure Blob Storage REST contract implemented by `@okikio/opfs/azure`. It is intended for +maintainers changing authentication, service-version behavior, block upload, copy, conditional replacement, listing, or +Azurite interoperability. -The client uses the Blob REST API directly rather than wrapping the Azure SDK. -That keeps the dependency graph small, but it also means this repository owns -the protocol work it chooses to implement. The source and tests must therefore +The client uses the Blob REST API directly rather than wrapping the Azure SDK. That keeps the dependency graph small, +but it also means this repository owns the protocol work it chooses to implement. The source and tests must therefore make the exact REST contract explicit. ```text -ObjectStoreType / AzureClientType - | - v +AzureClientType + | + +--> direct protocol use + | + `--> Azure object driver -> object adapter -> FileSystemType + Blob REST request construction | +--> SAS query authorization @@ -32,27 +31,22 @@ Web Fetch `--> Azurite ``` -Web Crypto owns HMAC-SHA256 for Shared Key. `@std/encoding` owns Base64, -`@std/async/pool` owns bounded block concurrency, and `@std/xml` owns list and -block-list documents. +Web Crypto owns HMAC-SHA256 for Shared Key. `@std/encoding` owns Base64, `@std/async/pool` owns bounded block +concurrency, and `@std/xml` owns list and block-list documents. This guide uses these evidence classes: - - **Implemented** means current source contains the behavior. - - **Protocol** means current Microsoft REST documentation defines the behavior. - - **Emulator** means Azurite reproduces enough of the contract for local - integration tests but is not treated as complete Azure parity. +- **Implemented** means current source contains the behavior. +- **Protocol** means current Microsoft REST documentation defines the behavior. +- **Emulator** means Azurite reproduces enough of the contract for local integration tests but is not treated as + complete Azure parity. -The current implementation was reviewed against Microsoft Learn and current -Azurite documentation on August 14, 2026. +The current implementation was reviewed against Microsoft Learn and current Azurite documentation on August 14, 2026. +## The client targets one container -The client targets one container --------------------------------- - -`createAzureClient()` binds one endpoint and one container. Blob keys supplied -to `head()`, `get()`, `put()`, `delete()`, `copy()`, and `list()` are relative to -that container. +`createAzureClient()` binds one endpoint and one container. Blob keys supplied to `head()`, `get()`, `put()`, +`delete()`, `copy()`, and `list()` are relative to that container. A normal cloud endpoint looks like: @@ -78,36 +72,59 @@ which becomes: http://127.0.0.1:10000/devstoreaccount1/container/path/to/blob ``` -This difference matters to Shared Key canonicalization. Microsoft documents -that the emulator account segment appears once in the URL path and is prefixed -again by the signing account name. The implementation derives that duplicated +This difference matters to Shared Key canonicalization. Microsoft documents that the emulator account segment appears +once in the URL path and is prefixed again by the signing account name. The implementation derives that duplicated canonical-resource form from the URL rather than hard-coding an Azurite branch. -`AZURE_STORAGE_VERSION` defaults to `2026-04-06`. A caller can select another -service version when it needs the size or authentication behavior of an older -REST contract. +`AZURE_STORAGE_VERSION` defaults to `2026-04-06`. A caller can select another service version when it needs the size or +authentication behavior of an older REST contract. + +The driver remains a separate public layer: + +```ts +import { createAzureClient } from "@okikio/opfs/azure"; +import { createAzureDriverFromClient } from "@okikio/opfs/driver/azure"; +import { createObjectAdapter } from "@okikio/opfs/adapter/object"; + +const client = createAzureClient(options); +const driver = createAzureDriverFromClient(client); +const adapter = createObjectAdapter(driver); +``` + +The driver adds provider requirements, limits, generic optimization metadata, deterministic planning, and physical +request metrics. The adapter translates Blob keys/prefixes into the filesystem primitive contract. + +## Optimizations are independently controllable + +`blockUpload` : Defaults to true. Large/streamed complete replacements can stage blocks with bounded concurrency and +commit a block list. When disabled, the client does not advertise native stream-write/multipart behavior. Materialized +values above the single Put Blob ceiling fail rather than silently changing physical upload strategy. +`serverCopy` : Defaults to true when the selected authorization strategy can support the source and destination +semantics used by the client. When disabled, the client does not advertise provider-side copy and the object +adapter/facade can select an honest fallback. -Authorization is explicit -------------------------- +Both switches are exposed by the client and through `driver.inspect().optimizations`. A future cache, prefetch, +block-selection, or copy optimization that changes observable behavior must be inspectable and independently +disableable. + +## Authorization is explicit `AzureCredentialType` supports four strategies: -| Kind | Wire mechanism | Intended use | -| ---- | -------------- | ------------ | -| `sas` | SAS fields remain in the request query | Browser/server delegated access | -| `bearer` | `Authorization: Bearer ...` | Microsoft Entra access token | -| `shared-key` | Canonical Shared Key HMAC-SHA256 | Trusted server and Azurite | -| `headers` | Caller returns authorization headers | Provider/host integration not otherwise modeled | +| Kind | Wire mechanism | Intended use | +| ------------ | -------------------------------------- | ----------------------------------------------- | +| `sas` | SAS fields remain in the request query | Browser/server delegated access | +| `bearer` | `Authorization: Bearer ...` | Microsoft Entra access token | +| `shared-key` | Canonical Shared Key HMAC-SHA256 | Trusted server and Azurite | +| `headers` | Caller returns authorization headers | Provider/host integration not otherwise modeled | -The client does not read environment variables. Credentials are supplied by -the caller, and bearer tokens can be refresh functions resolved immediately -before the request. +The client does not read environment variables. Credentials are supplied by the caller, and bearer tokens can be refresh +functions resolved immediately before the request. -Shared Key credentials contain an account name and Base64 account key. The -account key is a root-level storage credential. It should not be embedded in an -untrusted browser bundle. Browser applications normally use a scoped SAS or a -Microsoft Entra flow with suitable permissions. +Shared Key credentials contain an account name and Base64 account key. The account key is a root-level storage +credential. It should not be embedded in an untrusted browser bundle. Browser applications normally use a scoped SAS or +a Microsoft Entra flow with suitable permissions. ### Shared Key string to sign @@ -130,30 +147,26 @@ CanonicalizedHeaders CanonicalizedResource ``` -The signer supports the augmented Blob Shared Key format from service version -`2009-09-19` onward. Earlier versions are rejected rather than being signed -with modern rules that only look plausible. +The signer supports the augmented Blob Shared Key format from service version `2009-09-19` onward. Earlier versions are +rejected rather than being signed with modern rules that only look plausible. Two canonicalization rules change with the selected service version: - - `2014-02-14` and earlier sign a zero byte `Content-Length` as the literal - `0`. Later versions contribute an empty line for the same header. - - Versions before `2016-05-31` omit empty `x-ms-*` headers from - `CanonicalizedHeaders`. Version `2016-05-31` and later retain them as - `name:\n`. +- `2014-02-14` and earlier sign a zero byte `Content-Length` as the literal `0`. Later versions contribute an empty line + for the same header. +- Versions before `2016-05-31` omit empty `x-ms-*` headers from `CanonicalizedHeaders`. Version `2016-05-31` and later + retain them as `name:\n`. Canonical `x-ms-*` headers that participate in the selected version are: -1. converted to lowercase names; -2. normalized by collapsing linear whitespace outside quoted strings while - preserving whitespace inside quoted strings; -3. sorted by code-unit order; -4. emitted as `name:value\n`. +1. converted to lowercase names; +2. normalized by collapsing linear whitespace outside quoted strings while preserving whitespace inside quoted strings; +3. sorted by code-unit order; +4. emitted as `name:value\n`. -The quoted-string rule is significant for metadata and other extension headers. -For example, `alpha beta` canonicalizes to `alpha beta`, while the two spaces -inside `alpha "beta gamma"` remain two spaces. Collapsing the quoted value -would sign different bytes from the value Azure receives. +The quoted-string rule is significant for metadata and other extension headers. For example, `alpha beta` +canonicalizes to `alpha beta`, while the two spaces inside `alpha "beta gamma"` remain two spaces. Collapsing the +quoted value would sign different bytes from the value Azure receives. The canonical resource starts with: @@ -161,8 +174,7 @@ The canonical resource starts with: /account-name/request-path ``` -then appends lowercase query names in sorted order. Repeated values are sorted -and joined with commas. +then appends lowercase query names in sorted order. Repeated values are sorted and joined with commas. The signature is: @@ -176,39 +188,33 @@ and the HTTP header is: Authorization: SharedKey account-name:signature ``` -The deterministic unit suite signs Azurite requests with the documented -`devstoreaccount1` key and a fixed timestamp. It freezes exact signatures on -both sides of the `2014-02-14` zero-length change, verifies the `2016-05-31` -empty-header change, and rejects Shared Key versions older than `2009-09-19`. -This makes the tests independent from the implementation clock and catches -service-version canonicalization drift. +The deterministic unit suite signs Azurite requests with the documented `devstoreaccount1` key and a fixed timestamp. It +freezes exact signatures on both sides of the `2014-02-14` zero-length change, verifies the `2016-05-31` empty-header +change, and rejects Shared Key versions older than `2009-09-19`. This makes the tests independent from the +implementation clock and catches service-version canonicalization drift. -A low-level streamed body using Shared Key must provide `content-length` because -the signer cannot know the stream length without consuming it. High-level -`put()` avoids this problem by splitting the stream into known-size block +A low-level streamed body using Shared Key must provide `content-length` because the signer cannot know the stream +length without consuming it. High-level `put()` avoids this problem by splitting the stream into known-size block requests. +## The service version controls write limits -The service version controls write limits ------------------------------------------ - -Azure Blob limits changed across REST service versions. The client resolves -the relevant limit from the selected version instead of assuming the newest -size everywhere. +Azure Blob limits changed across REST service versions. The client resolves the relevant limit from the selected version +instead of assuming the newest size everywhere. `AZURE_LIMITS` records the values used by planning: -| Operation/era | Client limit | -| ------------- | ------------ | -| Maximum committed blocks | 50,000 | -| Maximum uncommitted blocks | 100,000 | -| `Copy Blob From URL` synchronous copy | 256 MiB | -| Old `Put Block` | 4 MiB | -| 2016-05-31 through 2019-era `Put Block` | 100 MiB | -| Current `Put Block` | 4,000 MiB | -| Old `Put Blob` | 64 MiB | -| 2016-05-31 through 2019-era `Put Blob` | 256 MiB | -| Current `Put Blob` | 5,000 MiB | +| Operation/era | Client limit | +| --------------------------------------- | ------------ | +| Maximum committed blocks | 50,000 | +| Maximum uncommitted blocks | 100,000 | +| `Copy Blob From URL` synchronous copy | 256 MiB | +| Old `Put Block` | 4 MiB | +| 2016-05-31 through 2019-era `Put Block` | 100 MiB | +| Current `Put Block` | 4,000 MiB | +| Old `Put Blob` | 64 MiB | +| 2016-05-31 through 2019-era `Put Blob` | 256 MiB | +| Current `Put Blob` | 5,000 MiB | The implementation selects: @@ -224,27 +230,20 @@ Put Blob limit older -> 64 MiB ``` -`Put Block From URL` uses the version table published on the current REST page: -4,000 MiB from version `2020-04-08` onward and 100 MiB before that point. -Microsoft's same page currently contains a contradictory prose sentence that -still says the operation is limited to 100 MiB. The version table is also -consistent with the service's modern block-blob capacity model, so the client -follows the table. This contradiction is recorded as an upstream documentation -risk rather than hidden. The Docker suite uses small ranges and therefore does -not prove the 4,000 MiB ceiling; an opt-in real Azure test must protect that -limit before it is treated as independently verified. - -`blockSize` defaults to 8 MiB. The constructor rejects a configured block size -above the selected service-version limit. A known body can require a larger -block size to remain within 50,000 committed blocks; the planner chooses the -larger legal size and rejects an impossible request before starting the commit. +`Put Block From URL` uses the version table published on the current REST page: 4,000 MiB from version `2020-04-08` +onward and 100 MiB before that point. Microsoft's same page currently contains a contradictory prose sentence that still +says the operation is limited to 100 MiB. The version table is also consistent with the service's modern block-blob +capacity model, so the client follows the table. This contradiction is recorded as an upstream documentation risk rather +than hidden. The Docker suite uses small ranges and therefore does not prove the 4,000 MiB ceiling; an opt-in real Azure +test must protect that limit before it is treated as independently verified. +`blockSize` defaults to 8 MiB. The constructor rejects a configured block size above the selected service-version limit. +A known body can require a larger block size to remain within 50,000 committed blocks; the planner chooses the larger +legal size and rejects an impossible request before starting the commit. -High-level upload has two paths ------------------------------- +## High-level upload has two paths -A materialized `Uint8Array` at or below the selected `Put Blob` limit uses one -`Put Blob` request with: +A materialized `Uint8Array` at or below the selected `Put Blob` limit uses one `Put Blob` request with: ```text x-ms-blob-type: BlockBlob @@ -270,29 +269,24 @@ Put Block List Get Blob Properties ``` -Block IDs are deterministic Base64 values derived from zero-padded sequential -numbers. Every request in one upload therefore has a stable order and the -commit document can list exactly the intended blocks. +Block IDs are deterministic Base64 values derived from zero-padded sequential numbers. Every request in one upload +therefore has a stable order and the commit document can list exactly the intended blocks. -`Put Block` requests intentionally do not receive destination `If-Match` or -`If-None-Match`. Uncommitted blocks are not yet the authoritative destination -blob. The precondition and final metadata belong on `Put Block List`, which is -the operation that commits the new block blob. +`Put Block` requests intentionally do not receive destination `If-Match` or `If-None-Match`. Uncommitted blocks are not +yet the authoritative destination blob. The precondition and final metadata belong on `Put Block List`, which is the +operation that commits the new block blob. -The XML block list is generated through `@std/xml/stringify`, so block IDs are -serialized by a real XML implementation rather than hand-escaped text. +The XML block list is generated through `@std/xml/stringify`, so block IDs are serialized by a real XML implementation +rather than hand-escaped text. -When `ObjectPutOptionsType.size` is supplied, the final streamed byte count must -match. A mismatch rejects the operation before final commit. +When `ObjectPutOptionsType.size` is supplied, the final streamed byte count must match. A mismatch rejects the operation +before final commit. -Unlike S3 multipart uploads, Azure uncommitted blocks do not have a separate -abort REST operation. Failed uploads can leave uncommitted blocks until Azure -cleans them up according to service policy. Documentation and tests therefore -must not describe stream cancellation as an atomic remote rollback. +Unlike S3 multipart uploads, Azure uncommitted blocks do not have a separate abort REST operation. Failed uploads can +leave uncommitted blocks until Azure cleans them up according to service policy. Documentation and tests therefore must +not describe stream cancellation as an atomic remote rollback. - -Range reads use the Blob range contract ---------------------------------------- +## Range reads use the Blob range contract `get()` maps package range fields to: @@ -300,23 +294,19 @@ Range reads use the Blob range contract x-ms-range: bytes=start-end ``` -`at` is the zero-based first byte. `length` controls the inclusive final byte. -When no length is supplied, the range remains open-ended. - -The client reports `rangeRead: true` because Azure Blob Storage can satisfy the -range at the provider rather than materializing the complete object in the -library first. +`at` is the zero-based first byte. `length` controls the inclusive final byte. When no length is supplied, the range +remains open-ended. +The client reports `rangeRead: true` because Azure Blob Storage can satisfy the range at the provider rather than +materializing the complete object in the library first. -Server-side copy preserves the Azure model ------------------------------------------- +## Server-side copy preserves the Azure model -Azure has more than one URL-based copy primitive. The client selects between -them rather than presenting one fictitious universal copy call. +Azure has more than one URL-based copy primitive. The client selects between them rather than presenting one fictitious +universal copy call. -For a source up to 256 MiB, `copy()` uses synchronous `Copy Blob From URL`. -For a larger source, it performs ranged `Put Block From URL` operations and then -commits them with `Put Block List`. +For a source up to 256 MiB, `copy()` uses synchronous `Copy Blob From URL`. For a larger source, it performs ranged +`Put Block From URL` operations and then commits them with `Put Block List`. ```text Get Blob Properties(source) @@ -331,55 +321,45 @@ Get Blob Properties(source) Put Block List ``` -The large-copy block size is increased when required to remain at or below -50,000 committed blocks. It is also constrained by the selected REST version's -`Put Block From URL` range limit. +The large-copy block size is increased when required to remain at or below 50,000 committed blocks. It is also +constrained by the selected REST version's `Put Block From URL` range limit. -The copy feature is advertised only when the selected service version supports -URL-copy operations and the configured credential type lets the client derive a -source authorization strategy. +The copy feature is advertised only when the selected service version supports URL-copy operations and the configured +credential type lets the client derive a source authorization strategy. For same-client copies: - - SAS includes its authorization on the generated source URL; - - Shared Key can authorize the destination and a same-account source; - - bearer credentials can use `x-ms-copy-source-authorization` from service - version `2020-10-02` onward; - - custom header authorization is not assumed to work for the source, so the - portable `copy` capability is disabled. +- SAS includes its authorization on the generated source URL; +- Shared Key can authorize the destination and a same-account source; +- bearer credentials can use `x-ms-copy-source-authorization` from service version `2020-10-02` onward; +- custom header authorization is not assumed to work for the source, so the portable `copy` capability is disabled. -Cross-account copy has additional source-authorization requirements. A Shared -Key for the destination account cannot sign a different account's source. Use -a source SAS or a suitable bearer/source authorization design rather than +Cross-account copy has additional source-authorization requirements. A Shared Key for the destination account cannot +sign a different account's source. Use a source SAS or a suitable bearer/source authorization design rather than assuming one account key grants cross-account access. -Source conditions map to the `x-ms-source-*` condition family where the selected -operation supports them. Destination `If-Match` / `If-None-Match` apply to the -single synchronous copy or to the final block-list commit for multipart copy. +Source conditions map to the `x-ms-source-*` condition family where the selected operation supports them. Destination +`If-Match` / `If-None-Match` apply to the single synchronous copy or to the final block-list commit for multipart copy. - -Listing is container pagination, not a directory API ------------------------------------------------------ +## Listing is container pagination, not a directory API `list()` requests the container with `restype=container&comp=list` and maps: | Package field | Azure query field | | ------------- | ----------------- | -| `prefix` | `prefix` | -| `delimiter` | `delimiter` | -| `limit` | `maxresults` | -| `cursor` | `marker` | - -`Blob` entries become `ObjectEntryType`. `BlobPrefix` entries become child -prefixes. `NextMarker` is returned as the next cursor. +| `prefix` | `prefix` | +| `delimiter` | `delimiter` | +| `limit` | `maxresults` | +| `cursor` | `marker` | -The filesystem adapter interprets directory markers and provider prefixes. The -Azure client itself retains object/blob terminology because Azure has no native -filesystem directory in the Blob service contract used here. +`Blob` entries become `ObjectEntryType`. `BlobPrefix` entries become child prefixes. `NextMarker` is returned as the +next cursor. +The Azure object driver and filesystem adapter interpret directory markers and provider prefixes. The Azure client +itself retains object/blob terminology because Azure has no native filesystem directory in the Blob service contract +used here. -Errors retain Azure request evidence ------------------------------------- +## Errors retain Azure request evidence `AzureError` keeps: @@ -390,100 +370,81 @@ x-ms-request-id when present original Response ``` -Azure can return XML or provider-specific text. The error parser uses structured -XML when available and retains the response even when a field is missing. - -The client has a configurable transport retry policy built on `@std/async/retry`. Client options control retry count, exponential -delay, jitter, and an optional per-attempt timeout. The policy retries 408, 429, 5xx, and transport failures for replayable -requests. Authorization is rebuilt on every attempt, which matters for refreshable bearer/custom credentials and Shared Key -dates. Redirects are manual so authorization is not silently carried to another authority. +Azure can return XML or provider-specific text. The error parser uses structured XML when available and retains the +response even when a field is missing. -A one-shot `ReadableStream` receives one attempt. The low-level `request()` API also accepts `retry: false` because replayability -does not prove that a provider-specific operation is safe to repeat. `request: { retries: 0 }` disables automatic retry for the -client. Provider-specific `Retry-After` interpretation is not yet modeled. +The client has a configurable transport retry policy built on `@std/async/retry`. Client options control retry count, +exponential delay, jitter, and an optional per-attempt timeout. The policy retries 408, 429, 5xx, and transport failures +for replayable requests. Authorization is rebuilt on every attempt, which matters for refreshable bearer/custom +credentials and Shared Key dates. Redirects are manual so authorization is not silently carried to another authority. -`getMetrics()` returns request, retry, terminal-failure, response, and optional Fetch-duration counters. `metrics: "none"` is the -baseline benchmark setting; `basic` counts; `timing` adds monotonic duration. +A one-shot `ReadableStream` receives one attempt. The low-level `request()` API also accepts `retry: false` because +replayability does not prove that a provider-specific operation is safe to repeat. `request: { retries: 0 }` disables +automatic retry for the client. Provider-specific `Retry-After` interpretation is not yet modeled. -`AbortSignal` reaches every Fetch operation. Cancellation ends local admission and HTTP work where Fetch can abort it. It does -not guarantee that Azure failed to accept a request before the signal reached the network stack. +`getMetrics()` returns request, retry, terminal-failure, response, and optional Fetch-duration counters. +`metrics: "none"` is the baseline benchmark setting; `basic` counts; `timing` adds monotonic duration. +`AbortSignal` reaches every Fetch operation. Cancellation ends local admission and HTTP work where Fetch can abort it. +It does not guarantee that Azure failed to accept a request before the signal reached the network stack. -Azurite is an integration target, not the specification -------------------------------------------------------- +## Azurite is an integration target, not the specification -`tests/provider/fixture.ts` uses the official `@testcontainers/azurite` module -with the pinned Azurite image and the well-known development account. -Testcontainers chooses a free mapped host port, so the concrete endpoint changes -per run while the account identity remains `devstoreaccount1`. +`tests/provider/fixture.ts` uses the official `@testcontainers/azurite` module with the pinned Azurite image and the +well-known development account. Testcontainers chooses a free mapped host port, so the concrete endpoint changes per run +while the account identity remains `devstoreaccount1`. -The test uses Shared Key, creates a disposable logical Azure container through the client's -signed low-level request, then exercises PUT, HEAD, range GET, conditional -create, block upload, server-side copy, list, delete, and the object-store -filesystem adapter. +The test uses Shared Key, creates a disposable logical Azure container through the client's signed low-level request, +then exercises PUT, HEAD, range GET, conditional create, block upload, server-side copy, list, delete, the Azure driver, +and the object filesystem adapter. -Azurite is intentionally treated as an emulator. Its documentation states that -it provides best-effort Azure Storage compatibility and can differ from the -cloud service. A green Azurite suite therefore proves real HTTP/authentication +Azurite is intentionally treated as an emulator. Its documentation states that it provides best-effort Azure Storage +compatibility and can differ from the cloud service. A green Azurite suite therefore proves real HTTP/authentication interoperability, not complete conformance with every Azure version or feature. -The deterministic unit suite remains responsible for exact Shared Key string -construction, version-dependent limits, condition placement, and source bearer -version gates. - -Before release, an opt-in real Azure Blob test should run against a disposable -container with short-lived CI credentials when organizational secret policy -permits it. - +The deterministic unit suite remains responsible for exact Shared Key string construction, version-dependent limits, +condition placement, and source bearer version gates. -Known non-goals ---------------- +Before release, an opt-in real Azure Blob test should run against a disposable container with short-lived CI credentials +when organizational secret policy permits it. -The current direct client does not claim full Azure Storage coverage. Important -features outside this focused contract include: +## Known non-goals - - hierarchical namespace/Data Lake Gen2 filesystem semantics; - - append blobs and page blobs; - - leases as a first-class high-level API; - - snapshots/version-ID aware filesystem paths; - - customer-provided encryption-key convenience APIs; - - immutability policies and legal holds; - - blob index tags; - - asynchronous `Copy Blob` polling workflows; - - batch operations; - - account/container administration beyond low-level requests; - - adaptive provider throttling beyond the configured exponential retry policy, including provider-specific `Retry-After` scheduling; - - Microsoft Entra token acquisition itself. +The current direct client does not claim full Azure Storage coverage. Important features outside this focused contract +include: -`request()` can reach an unmodeled REST operation when a caller supplies the -correct method, query, headers, and body. A feature should become a typed public -operation only when its ownership, failure behavior, version gates, and tests -are explicit. +- hierarchical namespace/Data Lake Gen2 filesystem semantics; +- append blobs and page blobs; +- leases as a first-class high-level API; +- snapshots/version-ID aware filesystem paths; +- customer-provided encryption-key convenience APIs; +- immutability policies and legal holds; +- blob index tags; +- asynchronous `Copy Blob` polling workflows; +- batch operations; +- account/container administration beyond low-level requests; +- adaptive provider throttling beyond the configured exponential retry policy, including provider-specific `Retry-After` + scheduling; +- Microsoft Entra token acquisition itself. +`request()` can reach an unmodeled REST operation when a caller supplies the correct method, query, headers, and body. A +feature should become a typed public operation only when its ownership, failure behavior, version gates, and tests are +explicit. -Primary specification sources ------------------------------ +## Primary specification sources Review these Microsoft sources before changing protocol behavior: - - Shared Key authorization: - https://learn.microsoft.com/rest/api/storageservices/authorize-with-shared-key - - Versioning for Azure Storage services: - https://learn.microsoft.com/rest/api/storageservices/versioning-for-the-azure-storage-services - - Put Blob: - https://learn.microsoft.com/rest/api/storageservices/put-blob - - Put Block: - https://learn.microsoft.com/rest/api/storageservices/put-block - - Put Block List: - https://learn.microsoft.com/rest/api/storageservices/put-block-list - - Put Block From URL: - https://learn.microsoft.com/rest/api/storageservices/put-block-from-url - - Copy Blob From URL: - https://learn.microsoft.com/rest/api/storageservices/copy-blob-from-url - - List Blobs: - https://learn.microsoft.com/rest/api/storageservices/list-blobs - - Azurite: - https://learn.microsoft.com/azure/storage/common/storage-use-azurite - -The REST documentation is authoritative for Azure. Azurite source and behavior -are integration evidence for the emulator only. +- Shared Key authorization: https://learn.microsoft.com/rest/api/storageservices/authorize-with-shared-key +- Versioning for Azure Storage services: + https://learn.microsoft.com/rest/api/storageservices/versioning-for-the-azure-storage-services +- Put Blob: https://learn.microsoft.com/rest/api/storageservices/put-blob +- Put Block: https://learn.microsoft.com/rest/api/storageservices/put-block +- Put Block List: https://learn.microsoft.com/rest/api/storageservices/put-block-list +- Put Block From URL: https://learn.microsoft.com/rest/api/storageservices/put-block-from-url +- Copy Blob From URL: https://learn.microsoft.com/rest/api/storageservices/copy-blob-from-url +- List Blobs: https://learn.microsoft.com/rest/api/storageservices/list-blobs +- Azurite: https://learn.microsoft.com/azure/storage/common/storage-use-azurite + +The REST documentation is authoritative for Azure. Azurite source and behavior are integration evidence for the emulator +only. diff --git a/docs/design.md b/docs/design.md index 4cee4b1..0d47c9a 100644 --- a/docs/design.md +++ b/docs/design.md @@ -1,397 +1,747 @@ -Architecture and invariants -=========================== +# Architecture and invariants -The architecture starts from one rule: +## Purpose -> The filesystem facade owns filesystem semantics. An adapter owns the mechanics of one backend. +`@okikio/opfs` is a storage programming model with an OPFS-shaped filesystem frontend. The package supports storage +systems that have very different native contracts. Browser OPFS exposes file and directory handles. Node, Deno, and Bun +expose host file APIs. Deno KV and IndexedDB expose values and transactions. S3 and Azure Blob expose object protocols. +Drizzle, db0, RxDB, and unstorage sit above their own storage engines. -That rule matters because OPFS, Node files, Deno KV, SQLite, S3, and Azure Blob do not have the same native operations. A useful -portable library must make common application behavior consistent without hiding those differences from performance-sensitive -or correctness-sensitive code. +The architecture keeps those differences visible while giving applications one portable filesystem API where that API +can be implemented correctly. -The complete data path is: +The defining path is: ```text - application - | - +---------------+---------------+ - | | - v v - path API OPFS-shaped handles - readFile / writeFile / walk DirectoryHandle / FileHandle - | | - +---------------+---------------+ - | - v - FileSystemType - | - normalize paths / normalize failures / cancellation - parent creation / recursive operations / staging - file locks / tree locks / sync-file lock lifetime - stream selection / bounded materialization / ownership - | - v - AdapterType - +----------------+----------------+ - | | | - v v v - native filesystem record storage object storage - OPFS/Node/Deno KV/doc/SQL S3/Azure Blob - | | | - | RecordStoreType ObjectStoreType - | | | - v v v - native bytes versioned row object key +native API / ecosystem + | + v + client optional protocol client + | + v + driver backend-native persistence + | + v + adapter driver -> canonical filesystem primitives + | + v + FileSystemType portable OPFS-shaped behavior + | + v + bridge FileSystemType -> real ecosystem contract ``` -The frontend therefore has one filesystem contract, but the adapter capability record still tells the truth about how that -backend gets the work done. +`integration` definitions are separate metadata that describe which of the two directions exist. They are not executable +bridges. -The adapter contract stays deliberately small ---------------------------------------------- +## Layer ownership -Every adapter implements six primitives: `stat`, `readFile`, `writeFile`, `readDir`, `createDir`, and `remove`. Everything else -is an optional acceleration or stronger native lifecycle. +### Client -This avoids a common adapter failure mode where every backend reimplements recursive copy, walk, parent creation, handles, -locking, and error normalization separately. If those policies live in every adapter, semantics drift as soon as one backend gets -a bug fix that the others do not. +A client owns a wire protocol when the protocol is useful independently of the filesystem abstraction. -Optional operations exist only when the backend can perform them natively: +Current examples: ```text -streamRead -> openReadStream -write mode -> writeStream when mode is in streamWriteModes -nativeCopy -> copy -nativeMove -> move -positionalWrite -> openWritableFile -syncAccess -> openSyncFile +src/s3.ts S3 REST + SigV4 + multipart + request policy +src/azure.ts Azure Blob REST + authentication + block upload + request policy ``` -`streamWriteModes` is intentionally a list. A local file can stream replace, append, and update. An object store can usually -stream a complete replacement but cannot append bytes to an existing object in place. One `streamWrite: true` flag would hide -that difference and make callers reason from a capability that was too broad. +A client can expose protocol operations that do not belong in a filesystem. For example, an S3 client can retain ETags, +conditional requests, upload IDs, provider request IDs, presigned requests, object metadata, and multipart controls. + +Node, Deno, Bun, OPFS, IndexedDB, and localStorage do not need a package-owned protocol client. They begin at the driver +layer. -The adapter can additionally publish hard `limits` and a durable `partition` description. Those are facts about the configured -backend, not policy guesses. `FileSystemType.inspect()` combines them with resolved optimization controls and effective facade -support. `plan()` uses the same information before I/O, so runtime execution and preflight selection share one route model. +### Driver -Route-changing optimizations are facade policy: +A driver owns one configured backend's storage mechanics. It must remain useful without `FileSystemType`. + +A driver owns: + +- backend-native operations; +- required resources and current availability facts; +- provider hard limits; +- implementation safety limits; +- caller-selected policy limits; +- dynamic limits that still require a probe; +- driver-specific optimization switches; +- deterministic preflight planning; +- physical backend metrics when available; +- disposal of resources whose ownership was explicitly transferred. + +A driver does not own recursive filesystem semantics merely because the adapter above it needs them. + +The three reusable driver families are: ```text -streamRead -streamWrite -rangeRead -nativeCopy -nativeMove +FileDriverType + OPFS / Node / Deno / Bun + +RecordDriverType + memory / Deno KV / localStorage / IndexedDB / Cache + SQLite rows / db0 / Drizzle / RxDB / unstorage + +ObjectDriverType + S3 / Azure Blob / custom object storage ``` -Each defaults to enabled and can be disabled independently. This is deliberately different from capability detection. The adapter -should expose the strongest implementation it has; the caller can force the safe fallback for differential tests, observability, -provider workarounds, or policy. A disabled route is never relabelled native. +These families preserve stronger native concepts. They are not a forced lowest-common-denominator interface. + +### Adapter + +An adapter is deliberately smaller. It translates one driver into the filesystem primitive set consumed by +`FileSystemType`. -`nativeCopy` is also separate. A provider-side S3 or Azure copy can move terabytes without transferring the source through this -process, even though the same provider has no filesystem rename. The facade checks native copy before opening a source stream. -That ordering is a performance invariant, not an implementation detail: +Required adapter primitives: ```text -correct selection -filesystem.copy() - | - +-- native copy available -> adapter.copy() - | - `-- no native copy -------> open source stream -> transfer +stat +readFile +writeFile +readDir +createDir +remove +``` + +Optional direct routes: -incorrect selection -open source stream -> discover native copy -> source GET was already wasted +```text +openReadStream +writeStream +copy +move +openWritableFile +openSyncFile ``` -Paths are virtual identities, not host paths -------------------------------------------- +The adapter reports whether those direct routes are native. It does not say that a portable operation is unavailable +merely because the facade can emulate it. -Every adapter receives a canonical `PathType`: +This distinction is important: ```text -/ -/a -/a/b.txt +adapter.nativeMove = false +FileSystemType.move() can still exist as copy + remove ``` -The adapter seam rejects or never receives forms such as: +The first value describes the translation layer. The second describes the effective public route. + +### FileSystemType + +`FileSystemType` owns portable filesystem behavior: + +- canonical virtual paths; +- parent creation; +- OPFS-shaped file and directory handles; +- recursive walk, copy, move, remove, and empty-directory operations; +- staged writable-file semantics; +- synchronization and lock lifetime; +- bounded stream materialization when a backend cannot stream; +- normalized filesystem failures; +- logical metrics; +- adapter/facade optimization switches; +- composition of driver preflight with adapter and facade policy. + +The facade must not claim that an emulated route has the atomicity, consistency, or memory behavior of a native route. + +### Bridge + +A bridge starts from an existing `FileSystemType` and implements another ecosystem's real contract. + +Current bridges are: ```text -a/b -/a/ -/a//b -/a/./b -/a/../b -/a\b +bridge/kv hierarchical asynchronous key/value view +bridge/unstorage unstorage Driver-shaped view ``` -Public methods accept more convenient input and call `normalizePath()` first. This lets callers write ordinary path-like input -without making every adapter repeat normalization rules. +A bridge is valid only when the filesystem can satisfy the ecosystem contract. A filesystem cannot become a SQL engine +by renaming methods. A real RxDB reverse bridge would need to implement the complete `RxStorage` semantics, including +conflict, query, checkpoint, change-stream, cleanup, and lifecycle behavior. + +### Integration definition + +`integration/definition` stores import-safe direction metadata: + +```text +toOpfs ecosystem/native resource -> driver/adapter -> FileSystemType +fromOpfs FileSystemType -> ecosystem bridge +``` -The host filesystem adapters map this virtual namespace under one configured host root. A virtual path cannot escape that root -after host resolution. The virtual namespace does not expose symbolic-link identity, permission bits, or arbitrary host paths -as part of the portable contract. +An unsupported direction must state why it is unsupported. A definition never substitutes for the missing executable +contract. + +## Driver definitions and third-party extension + +`defineDriver()` is the smallest third-party extension seam. It validates structured driver metadata without registering +global state. + +```ts +import { defineDriver } from "@okikio/opfs/driver"; + +const driver = defineDriver({ + name: "example", + kind: "record", + provides: ["get", "set", "delete", "list"], + ownership: "borrowed", + requirements: [ + { code: "database", state: "available" }, + ], + limits: [ + { + code: "value-bytes", + kind: "hard", + source: "provider", + unit: "bytes", + value: 64 * 1024, + }, + ], + optimizations: [ + { + code: "partition", + enabled: true, + changesBehavior: true, + disableable: true, + }, + ], +}); +``` -Record stores and object stores need different translation layers ------------------------------------------------------------------ +`provides` records stable backend capability names for inspection. It is an open vocabulary so a provider-specific +driver can report operations beyond the three core driver families. `ownership` reports the long-lived backend resource +relationship: -A value store naturally answers "what value is stored at this key?" It does not naturally answer filesystem questions such as -"what are the direct children of this directory?" `RecordStoreType` supplies the reusable record translation for that family. +```text +none no disposable external backend resource is owned by this driver +borrowed the caller retains ownership of the injected backend resource +owned the driver owns the backend resource and can release it +``` -The complete persisted record has a canonical `path` plus a separate `parent`. Direct directory listing can therefore use an -index or prefix query over `parent` instead of scanning and parsing every path. File bytes are base64 so the fallback shape can -survive JSON, Web Storage, document databases, and SQL text. The extra storage and encoding work is accepted only for this -complete-record path. Native file and object adapters do not use that representation. +This report is separate from method presence. Typed file, record, and object driver contracts remain the operational +API. -The record contract also has optional byte lanes. A store can provide metadata-only stat, direct ranges, streaming reads, direct -materialized writes for selected modes, or streaming writes for selected modes. The generic record adapter advertises only the -lanes the store declares. Deno KV uses these lanes so a partitioned file is not reconstructed into one base64 record for stat, -listing, range reads, streaming reads, materialized append/update, or streamed replacement. Its append/update lane constructs a -new immutable generation one part at a time. This still performs provider I/O for untouched bytes, but it keeps JavaScript memory -bounded by the configured part/concurrency policy. Simpler record stores keep the small complete-record contract. +A concrete storage implementation should normally use `defineFileDriver()`, `defineRecordDriver()`, or +`defineObjectDriver()` so its operational contract is type-checked as well as its metadata. -An object store has a different strength: large objects, byte ranges, prefix listing, whole-object replacement, conditional -requests, and provider-side copy. `ObjectStoreType` preserves those concepts before `createObjectAdapter()` translates them into -filesystem operations. +No process-global driver or adapter registry is required. A package can export a definition and normal constructors. The +application chooses and composes them explicitly. -Files map directly to object keys. Empty directories need a marker object because a pure prefix does not exist until at least one -child exists: +## Requirements describe availability + +A requirement is structured data with a stable `code` and one state: ```text -/photos -> photos/ marker -/photos/a.jpg -> photos/a.jpg -/photos/2026/b.jpg -> photos/2026/b.jpg +available +missing +unknown ``` -The adapter also accepts implicit directories inferred from foreign prefixes. This matters when the bucket/container is not -created exclusively by this library. +A requirement can describe facts such as: + +```text +Deno KV database supplied +IndexedDB exposed in the current realm +S3 credentials resolved +browser OPFS root acquired +transaction capability supplied by a database integration +``` -An object namespace can contain both `mixed` and `mixed/child`. A real filesystem cannot. The filesystem view resolves an exact -`mixed` object as the file, because exact `stat()` already does that. Reads and writes follow the same rule. This creates one -stable interpretation for a foreign namespace instead of making `stat()` and `writeFile()` disagree. +A driver should not run hidden provider I/O from `inspect()` merely to turn every unknown into a known value. Dynamic +facts can remain unknown until the caller runs an explicit probe or performs the operation. -Streaming stays native only when the backend really streams ------------------------------------------------------------ +## Limits have provenance -`writeFile()` accepts strings, Blob, ArrayBuffer, typed-array views, ReadableStream, and AsyncIterable input. +A numeric limit is not meaningful unless the caller can tell where it came from. -When the selected adapter supports native streaming for the requested write mode, the facade forwards a byte stream directly. -When it does not, the facade collects the stream below `maxBufferedWriteBytes` and then calls the materialized adapter write. -Crossing the limit cancels the producer and returns `too-large`. +`LimitType` records: ```text -ReadableStream - | - +-- native mode supported ------> adapter.writeStream() - | - `-- no direct adapter stream lane - | - v - bounded collector - | | - | +-- over limit -> cancel producer -> too-large - v - Uint8Array - | - v - adapter.writeFile() +code +kind hard | policy | dynamic +source provider | implementation | user | probe +unit bytes | count | milliseconds | operations +value optional for a dynamic unknown ``` -This makes memory behavior visible. A simple record-backed adapter can accept streamed input through the public API while still -reporting an emulated stream route because the complete record is materialized before storage. A specialized record store can -report a partitioned stream lane when its own physical layout preserves backpressure. Deno KV does exactly that for -replacement streams when partitioning is enabled. +Examples: -Partitioning is not hidden as an optimization. It changes durable physical layout, so the adapter publishes `mode`, part size, -threshold, maximum parts, and layout identity. Deno KV exposes `never | auto | always`. Its parts are written under a new -generation and the manifest is committed last. A pre-manifest crash can leak unreachable parts but cannot publish a partial new -logical file. +```text +Deno KV serialized value ceiling + kind: hard + source: provider -Multipart and block-upload clients use `@std/async/pool` for bounded request admission. The surrounding client still owns the -provider lifecycle: S3 waits for already-started part requests before it sends AbortMultipartUpload, while Azure documents that -uncommitted blocks have no equivalent abort operation. The pool limits concurrent work; it does not become authority for remote -commit, cleanup, or the terminal provider failure. +Deno KV conservative inline decoded-body budget + kind: policy + source: implementation/user -HTTP retry policy is separate from body replayability. Direct clients rebuild authorization on every retry and use exponential -backoff with jitter, but a mechanically replayable request can still be semantically non-idempotent. S3 multipart initiation and -completion therefore disable automatic request retry. Uploaded parts use stable part numbers and can use the normal retry policy. -The low-level S3/Azure request APIs expose `retry: false` so a caller can make the same decision for provider-specific operations. -One-shot `ReadableStream` request bodies are never retried automatically. +maxParts chosen by the application + kind: policy + source: user -Object append and update are optimistic read-modify-write --------------------------------------------------------- +available browser storage quota + kind: dynamic + source: probe +``` -Object stores do not expose a portable in-place byte update. Append and update therefore use the current object as the starting -image, modify that image, and replace the object. +Missing limits mean unknown. They never mean unlimited. -Without a precondition, two writers can both read version A and then publish different replacements; the later one silently -loses the earlier write. When the object client advertises conditional writes, the adapter uses the current ETag as `If-Match`. -A concurrent change then fails the replacement instead of becoming silent data loss. +## Optimizations are inspectable policy + +Every optimization that can change observable behavior must be independently disableable. + +Observable behavior includes more than returned bytes. It includes: + +- request count; +- failure timing; +- storage layout; +- atomicity or visibility points; +- consistency/caching behavior; +- provider-side resource lifetime; +- retry/cancellation timing; +- memory use when the alternate route has different materialization behavior. + +`DriverOptimizationSchema` enforces the critical invariant: + +> `changesBehavior: true` requires `disableable: true`. + +Current examples include: ```text -writer A: HEAD ETag=A -> GET A ---------> PUT if-match A -> succeeds -writer B: HEAD ETag=A -> GET A -------------------------> PUT if-match A -> fails +S3 delayed multipart promotion +S3 derived signing-key cache +Azure block upload +Azure server-side copy +Deno KV partition layout ``` -If a provider claims conditional writes but does not return an ETag for an existing object, the adapter refuses append/update. -That is safer than publishing an unconditional write while the capability record says optimistic protection exists. +The facade has its own route switches for streaming, ranges, native copy, and native move. Driver and facade switches +remain separate because they control different layers. -The provider client can disable `conditionalWrite` when a compatible protocol implementation does not support the required -precondition. The library does not choose provider behavior from a provider-name table. +## Planning is deterministic -Copy and move preserve their real commit behavior -------------------------------------------------- +A driver planner accepts the concrete operation shape: -Native host filesystems use their copy and rename operations when available. Object stores use provider-side copy when the -client can do it. The facade removes/rejects the destination according to its own overwrite contract before invoking native copy, -so the adapter does not have to invent another overwrite policy. +```text +operation +canonical path +canonical destination when relevant +logical size when known +input bytes when different from final size +bytes vs stream source +write mode +range flag +``` -When no native copy exists, file bytes move through a stream when both ends support streaming or through bounded materialization -otherwise. +`plan()` performs no storage or network I/O. It returns: -A native move can be atomic or near-atomic according to the host/provider operation. The portable fallback is explicitly: +```text +supported +support native | partitioned | unsupported at the driver layer +partBytes +parts +problems[] +actions[] +``` + +Problems and actions are structured. Their human messages are not the identity used by application logic. + +The facade planner then adds adapter and filesystem facts. One final plan can therefore explain all of these at once: ```text -copy source -> destination - | - +-- copy failed -> source remains - | - `-- copy succeeded -> remove source +driver: Deno KV key exceeds a provider serialized-key ceiling +adapter: no native range route +filesystem: requested stream would exceed maxBufferedWriteBytes +``` + +The caller can distinguish each cause and choose a concrete action. + +## File drivers + +A file driver is closest to native OPFS semantics. Node, Deno, Bun, and browser OPFS can expose direct ranges, streams, +native copy/move, asynchronous positional files, or synchronous random access when the runtime supports them. + +The adapter above a file driver is intentionally close to delegation: + +```text +Node file APIs + | +Node file driver + | +file adapter + | +FileSystemType +``` + +The host-path mapper lives with drivers. It maps virtual `/` below one configured host directory and rejects escape from +that host root. + +## Record drivers + +Record storage naturally addresses complete values or documents instead of files. `RecordDriverType` therefore defines +logical record operations plus optional stronger byte lanes. + +Portable record methods: + +```text +get(path) +set(record) +delete(path) +list(parent) +``` + +Optional stronger methods: + +```text +stat(path) metadata without body reconstruction +readFile(path, range) direct byte/range access +openReadStream(path) direct streaming +writeFile(path, bytes) direct write modes +writeStream(path, stream) direct streaming write modes +``` + +The driver declares replacement semantics, binary support, and transaction availability separately. This lets a SQLite +or Deno KV driver preserve stronger behavior without pretending localStorage has it. + +The generic record format remains a portable fallback. It uses base64 file bodies because JSON/document/text-column +stores can all preserve that representation. A specialized driver is free to use native BLOB/byte storage internally and +expose the same logical record contract above it. + +## Object drivers + +Object storage preserves object semantics before the adapter translates them into files/directories. + +An object driver can retain: + +- ranged GET; +- conditional writes; +- validators/ETags; +- provider object versions; +- metadata; +- native/server-side copy; +- multipart/block upload; +- continuation tokens; +- provider request metrics. + +The filesystem adapter does not remove these concepts from the driver. It uses the subset required to provide canonical +filesystem primitives. + +## Partitioning belongs to drivers + +Partitioning changes physical storage layout, so it belongs at the backend driver layer. + +Examples: + +```text +Deno KV one logical file -> manifest + value parts +S3 one object upload -> multipart upload parts +Azure Blob one blob upload -> blocks + committed block list +SQL possible future file row -> part rows / BLOB segments +``` + +These systems have different visibility, cleanup, atomicity, and retry rules. A universal facade chunker would hide +those provider-specific guarantees. + +A partitioning strategy should describe: + +```text +whether it changes durable layout +its activation policy +part size +part count ceiling +visibility/commit point +cleanup behavior +streaming capability +memory behavior +whether callers can disable it +``` + +## Deno KV reference layout + +Deno KV demonstrates the full model. + +The provider documents serialized key/value limits. The driver also chooses smaller decoded-body budgets because a raw +byte count is not equal to serialized value size. + +The large-file layout uses an immutable generation and manifest-last publication: + +```text +old manifest -> old generation + +write new part 0 +write new part 1 +... +write new part N + | + v +write new manifest logical visibility point + | + v +remove old reachable generation ``` -That fallback is not atomic. A failure after copy and before remove can leave both entries. The API documents this instead of -claiming POSIX rename semantics on every backend. +If part writing fails, the new manifest is not published. The previous generation remains visible. The driver +best-effort removes parts from the failed generation. -Before copy or move, source and destination are checked for overlap. The library never removes an overwrite destination that is -an ancestor or descendant of the source. +A process crash before publication can still leave unreachable physical parts. That is storage leakage, not a partially +visible logical file. `DenoKvDriverType.collect()` exposes explicit, age-gated reclamation. The default one-hour grace +period avoids ordinary collection racing a recent unpublished generation, and `maxDeletes` bounds one maintenance pass. +Background deletion is not hidden inside ordinary reads or writes. -Coordination protects cooperating callers, not the whole storage system ------------------------------------------------------------------------- +The Deno KV planner also estimates physical tuple size from the concrete logical path. A file can be small enough to fit +by byte count while its physical key is too large. The planner reports that condition before provider I/O. -There are two classes of mutation. +## Filesystem path invariant -A file mutation acquires a shared tree lock plus an exclusive lock for that canonical file path: +Every adapter and driver path that participates in the filesystem seam is canonical: + +```text +/ +/a +/a/b.txt +``` + +The public facade can accept normalizable input, but `normalizePath()` runs before backend calls. Root escape, +backslashes, and NUL are rejected. + +The virtual path namespace is not an operating-system path namespace. Host file drivers map the canonical path below one +configured host root. + +## Streaming and memory invariant + +Large file size must not automatically become JavaScript heap size. + +A native streaming route is used only when the selected driver and adapter expose it and the corresponding optimization +is enabled. Otherwise the facade can materialize an input only up to `maxBufferedWriteBytes`. + +```text +stream + | + +-- native driver route --------------------> bounded backend streaming + | + `-- facade fallback -> bounded collector + | + +-- under limit -> materialized adapter write + `-- over limit -> cancel producer + too-large +``` + +The capability report distinguishes those routes. It does not label a buffered fallback as native streaming. + +## Writable-file invariant + +OPFS-shaped `createWritable()` stages a logical file image and commits on close. Abort discards the staged image. + +This is useful compatibility behavior, not the preferred large sequential write path. A caller that can use +`writeFile()` gives the facade a chance to select a true streaming adapter route. + +## Synchronous-file lifetime + +A synchronous file has two coupled resources: + +```text +facade path lock <------ same lifetime ------> driver sync file + | | + +---------------- close() ----------------+ +``` + +The path lock must remain held for the native file lifetime. `writeAll()` repeats partial writes until the complete +input is written or the backend reports no progress. + +## Coordination invariant + +The facade coordinates callers that use the same library lock namespace. + +File mutation: ```text shared tree lock | -exclusive /a/file lock +exclusive file-path lock | -write or sync-file lifetime +write / writable file / sync file lifetime ``` -A structural mutation such as recursive copy, move, remove, or empty-directory work acquires the exclusive tree lock: +Structural mutation: ```text exclusive tree lock | -structural mutation +copy / move / recursive remove / emptyDir ``` -Independent files can therefore make progress concurrently while a tree mutation cannot race an active library file mutation. -The in-realm lock implementation queues new readers behind an already-waiting writer so a busy read/write workload does not -starve structural work. +`local` coordination only spans one JavaScript realm. `web-locks` can coordinate cooperating browser realms that share +the lock namespace. `none` performs no library coordination. + +Database or distributed applications that require cross-process serialization must use the database/provider's real +transaction, lease, advisory-lock, or equivalent primitive. A local facade lock cannot provide that guarantee. -`coordination: "web-locks"` uses the browser Web Locks API. `auto` uses Web Locks when exposed and falls back to in-realm FIFO -coordination. `local` is one-realm coordination only. `none` preserves cancellation and adapter semantics but makes the caller -responsible for concurrency. +## Copy and move invariant -None of these modes becomes a distributed lock. Separate Node processes, browser profiles, hosts, or independent applications -need provider/database coordination when same-path atomicity matters across those processes. +Native copy and native move are separate capabilities. -Synchronous and asynchronous writable resources own locks for their complete lifetime --------------------------------------------------------------------------------------- +A native copy can avoid routing bytes through JavaScript. Object stores commonly provide this even when they cannot +provide rename semantics. -A synchronous file is not one short method call. It owns both the adapter file resource and the facade path lock until close: +When native move is absent: ```text -facade path lock <-------- same lifetime --------> adapter sync file - | | - +------------------- close() ------------------+ +source -> copy -> destination + | + `---------- remove source after successful copy ``` -This prevents an asynchronous write through the same facade from entering while synchronous random access is active. -`writeAll()` loops over partial native writes until the complete input is written or the backend reports no progress. +This fallback is not atomic. A failure after copy can leave both paths. Inspection and planning identify the route as +emulated. + +Source/destination overlap is checked before recursive structural work. An overwrite cannot delete an ancestor or +descendant that contains the source. -The OPFS-shaped `createWritable()` facade stages one file image and commits it on close. Abort discards the staged image. This is -useful for compatibility with File System API write commands, including seek and truncate. It is not the recommended path for -very large sequential files because the staged image is materialized. `FileSystemType.writeFile()` can use an adapter's native -streaming path instead. +## Database topology invariant -Integration direction is explicit ---------------------------------- +Two database directions must remain distinct. -An adapter is `ecosystem -> OPFS`. A driver is `OPFS -> ecosystem`. A bridge is only a descriptor that groups those directions; -it does not add a third translation layer to each operation. +Database-backed filesystem: ```text -ecosystem/client -> adapter -> FileSystemType -> driver -> ecosystem contract +Drizzle/db0/RxDB/SQLite database + | + record driver + | + record adapter + | + FileSystemType ``` -Some ecosystems are genuinely bidirectional. unstorage has a storage contract that can be consumed as a record backend and a -driver contract that can be implemented over `FileSystemType`. RxDB, db0, and Drizzle are not symmetric: their reverse -contracts require query, conflict, dialect, schema, or change-stream semantics a filesystem does not own. `defineBridge()` -therefore requires an explicit reason for unsupported directions instead of encouraging a false adapter. +SQLite database stored on OPFS: + +```text +application + | + Drizzle + | +SQLite engine + | +SQLite VFS + | +FileSystemType / native OPFS +``` -Cancellation and disposal are different operations --------------------------------------------------- +The current `driver/sqlite` and `adapter/sqlite` implement the first topology. They do not implement a SQLite VFS. A +future VFS must implement the SQLite engine's real file/VFS contract. -An `AbortSignal` asks active work to stop before a commit when possible. Closing a filesystem ends ownership of the facade. -Closing the facade does not dispose the adapter unless `disposeAdapter: true` transferred that ownership. +## Resource ownership -The same rule continues below the adapter: +Injected resources are borrowed by default. ```text -caller creates database/client/cache - | - +--> adapter borrows it - | | - | +--> filesystem closes - | `--> resource remains open - | - `--> caller still owns resource +caller creates database/client/filesystem + | + +--> driver/adapter/bridge borrows it + | + `--> caller remains owner ``` -An adapter option such as `disposeDatabase`, `disposeStore`, or another explicit ownership flag changes that lifecycle. The -option exists because connection pools, RxDB collections, unstorage instances, object clients, and caches are commonly shared by -more than one subsystem. +Ownership transfers only through an explicit option such as: -Errors normalize the portable category without erasing the provider cause -------------------------------------------------------------------------- +```text +disposeDatabase +disposeStorage +disposeDriver +disposeAdapter +disposeFileSystem +``` + +Disposal is idempotent at the owning layer where the public contract promises idempotency. A library must not close a +shared connection or filesystem merely because a facade closes. A driver only exposes backend disposal when its +construction options transferred ownership, so higher layers cannot accidentally dispose a borrowed database or storage +instance. + +Read-only policy is also a driver property for record backends. A read-only driver reports `write: false`, omits +optional write primitives, and rejects direct mutations before backend I/O. Adapters preserve that state rather than +inventing a second write policy that can disagree with the driver. + +## Cancellation invariant + +Cancellation asks active work to stop. Disposal releases owned resources. They are different operations. + +Long-running drivers check the signal before expensive work and between bounded chunks. When the facade aborts a stream +write, it cancels the producer when practical so upstream work does not continue after the file operation has become +terminal. + +Provider cleanup can need a separate bounded signal. For example, canceling an S3 multipart write must not use the +already aborted caller signal for the `AbortMultipartUpload` cleanup request. + +## Error invariant + +Backends fail with different error types. The facade normalizes known filesystem conditions to `FileSystemError` codes +while retaining the original cause. -Browsers use DOMException names. Node/Deno/Bun expose host error codes. Databases and cloud providers have their own errors. -`toFileSystemError()` maps known failures to stable package categories while retaining the original `cause`. +```text +DOMException / Node error code / provider error + | + v + FileSystemError + code + operation + path + cause +``` + +Protocol clients keep their own rich errors where provider-specific data matters. Translation into a filesystem error +happens at the storage/filesystem layer, not by deleting provider information at the client. + +## Metrics are layered + +Logical and physical work are not the same metric. + +`MetricsType` belongs to `FileSystemType` and records logical operations, logical bytes, route selection, facade +buffering, and optional facade timing. + +`DriverMetricsType` belongs to a driver and can record physical work such as: + +```text +provider requests +retries +responses/failures +logical payload bytes +physical bytes +parts/blocks/chunks +peak active provider work +backend duration +cleanup duration +``` -S3 and Azure clients also retain provider request identities on their own errors. Those IDs matter when a service returns an -unexpected result and the provider support logs are the only authoritative trace. +The benchmark matrix should compare each layer independently rather than attributing every cost to the facade. -Import safety follows the package graph ---------------------------------------- +## Import-safety invariant -The root package exports the portable facade, native browser OPFS convenience path, schemas, errors, handles, and capability -probes. It does not export every adapter from the root. +The root package is browser-safe. Runtime/provider code remains on explicit subpaths. ```text -@okikio/opfs browser-safe core + native OPFS -@okikio/opfs/adapter/node node:fs imports -@okikio/opfs/adapter/deno Deno runtime APIs -@okikio/opfs/adapter/bun Bun + Node-compatible APIs -@okikio/opfs/s3 Web Fetch/Crypto S3 client -@okikio/opfs/azure Web Fetch Azure Blob client -@okikio/opfs/adapter/drizzle optional drizzle-orm peer +@okikio/opfs +@okikio/opfs/driver/node +@okikio/opfs/driver/deno +@okikio/opfs/driver/bun +@okikio/opfs/driver/s3 +@okikio/opfs/driver/azure +@okikio/opfs/adapter/* +@okikio/opfs/bridge/* ``` -Importing a module does not read environment variables, configure logs, connect to providers, start workers, or mutate a global -adapter registry. +Importing a module does not connect to storage, read environment variables, start a worker, configure global logging, or +mutate a process registry. -Schemas are executable contracts, not duplicated type declarations -------------------------------------------------------------------- +## Review rules -Project-owned structural values use Zod schemas and inferred TypeScript output types. Public schema constants end in `Schema`. -Serializable project-owned types normally end in `Type`. +A storage change is not complete until these questions have concrete answers: -Zod 4 implements Standard Schema. The exported Zod value is therefore also the Standard Schema value. Creating a second OPFS -schema wrapper would add maintenance without adding a stronger contract. +1. Which layer owns the behavior? +2. Is the provider/native contract preserved below the adapter? +3. Are provider, implementation, user, and dynamic limits distinguishable? +4. Can an observable optimization be disabled? +5. Does planning use the actual path/size/source shape needed to detect known limits? +6. Is growing work bounded by bytes, parts, concurrency, retries, or time? +7. Who owns cancellation and who owns disposal? +8. Does an emulated route state its weaker atomicity, consistency, or memory behavior? +9. Does a reverse bridge implement the ecosystem's real contract? +10. Do tests and benchmarks exercise the layer being claimed rather than bypassing it? diff --git a/docs/ecosystems.md b/docs/ecosystems.md index 73fcfde..d3f95a3 100644 --- a/docs/ecosystems.md +++ b/docs/ecosystems.md @@ -1,158 +1,189 @@ -Ecosystem integrations -====================== +# Ecosystem integrations -`@okikio/opfs` integrates with another storage ecosystem at the highest stable abstraction the application already owns. It does -not duplicate every provider driver from that ecosystem. +## Purpose -That rule gives two complementary directions: +Storage ecosystems can connect to `@okikio/opfs` in two different directions. The package names those directions +explicitly and does not force symmetry where the upstream contract cannot be implemented correctly. ```text -existing storage resource existing OPFS filesystem - | | - v v - adapter reverse driver - | | - v v - FileSystemType KV / unstorage API +ecosystem/native resource FileSystemType + | | + driver bridge + | | + adapter v + | ecosystem API + v + FileSystemType ``` -The forward path lets filesystem-shaped application code use another storage system. The reverse path lets another ecosystem -consume any backend already reachable through `FileSystemType`. +`integration` definitions describe the two directions. They are metadata, not bridges. -Bridge descriptors make asymmetry part of the contract ------------------------------------------------- +## Direction inventory -`@okikio/opfs/bridge` groups the existing forward adapter and reverse driver for an ecosystem. It does not force every -integration to be symmetric. +| Integration | ecosystem -> OPFS | OPFS -> ecosystem | Reverse status | +| ----------- | ----------------- | ----------------- | -------------------------------------------------------------------------- | +| unstorage | yes | yes | `bridge/unstorage` implements a real Driver-shaped contract | +| RxDB | yes | no | a complete reverse path must implement `RxStorage` | +| db0 | yes | no | a filesystem is not a SQL database/query engine | +| Drizzle | yes | no | a filesystem is not a Drizzle dialect/schema/query engine | +| generic KV | n/a | yes | `bridge/kv` exposes the small contract the filesystem can actually satisfy | -| Bridge | ecosystem -> OPFS | OPFS -> ecosystem | Why the reverse side is absent when unsupported | -| --- | --- | --- | --- | -| `UnstorageBridge` | yes | yes | both stable shapes exist | -| `RxDbBridge` | yes | no | `RxStorage` also owns queries, conflicts, change streams, cleanup, and storage-instance semantics | -| `Db0Bridge` | yes | no | a filesystem is not a SQL query/dialect engine | -| `DrizzleBridge` | yes | no | a filesystem does not own Drizzle schema, dialect, or query-builder behavior | -| `KeyValueBridge` | no | yes | the generic reverse KV shape does not define persistence semantics needed to build an adapter | +`defineIntegration()` validates that declared support and constructors agree. An unsupported direction must have a +reason. -`defineBridge()` validates that direction declarations agree with real constructors. An unsupported direction must include a -reason. Third-party integrations can therefore publish capability honestly without inventing a method that only works for a -small subset of the upstream contract. +## unstorage -Unstorage works in both directions without a provider explosion ---------------------------------------------------------------- +### unstorage as the backend -The forward adapter accepts the high-level unstorage `Storage` contract: +The forward direction accepts unstorage's high-level `Storage` object: -```ts -import { createStorage } from "unstorage"; -import memoryDriver from "unstorage/drivers/memory"; -import { createFileSystem } from "@okikio/opfs"; -import { createUnstorageAdapter } from "@okikio/opfs/adapter/unstorage"; - -const storage = createStorage({ driver: memoryDriver() }); -const fileSystem = createFileSystem(createUnstorageAdapter(storage)); +```text +unstorage Storage + | +createUnstorageDriver + | +createRecordAdapter + | +FileSystemType ``` -Unstorage remains responsible for its selected driver, mounts, provider SDKs, retry behavior, and provider-specific limits. The -OPFS adapter only uses the stable high-level operations it needs to persist records. +This means the package does not reimplement unstorage's provider catalogue. Memory, filesystem, Redis, S3, Azure, +Cloudflare, Deno KV, IndexedDB, db0, or another upstream driver remains unstorage's responsibility. + +The OPFS record driver uses a private prefix and reversible key encoding. Read-only storage can be declared read-only at +the adapter layer. -This is intentionally broader than maintaining separate OPFS adapters for every unstorage driver. Current unstorage drivers span -browser storage, Cloudflare, Azure, S3, Deno KV, filesystem, Redis, databases, blobs, HTTP, and other providers. Duplicating that -catalog here would create a second compatibility matrix that would drift from upstream. +### FileSystemType as an unstorage Driver -The reverse direction is more powerful after the generic key-value driver: +The reverse direction is a real bridge: ```text -unstorage Storage - | - v -@okikio/opfs unstorage Driver - | - v -KeyValueDriverType - | - v +unstorage + | +createUnstorageBridge(fileSystem) + | FileSystemType - | - +-- native OPFS - +-- Node / Deno / Bun - +-- S3 / Azure Blob - +-- localStorage / IndexedDB / Cache - +-- Deno KV / SQLite - +-- RxDB / db0 / Drizzle / unstorage - `-- custom adapter ``` -`createKeyValueDriver()` owns the collision-safe filesystem mapping. `createUnstorageDriver()` only translates that contract to -unstorage's driver method names and `maxDepth` flag. Both reverse views retain `inspect()`, `plan()`, and `getMetrics()` from the -backing filesystem. An ecosystem caller can therefore reject a value above `maxFileBytes`, see when a streamed write would -buffer or partition, disable a native route on the filesystem, and observe the same counters without a second capability table. +The bridge supports values, raw bytes, metadata, keys, clear, and disposal behavior used by the stable unstorage Driver +surface implemented here. -The extra key directory is required because a KV store can contain both `foo` and `foo:bar`: +A private directory+leaf layout is necessary because unstorage can contain both: ```text -foo -> /key-foo/value -foo:bar -> /key-foo/key-bar/value +foo +foo:bar ``` -A naive `:` to `/` conversion would try to make `/foo` both a file and a directory. The private `value` leaf removes that -conflict while reversible segment encoding keeps `%`, `~`, spaces, slashes inside a key segment, and other URI-sensitive text -distinct. +A normal filesystem cannot make `/foo` both a file and a directory, so the bridge stores each exact value in a private +`value` leaf below its encoded hierarchy directory. -RxDB stays above RxStorage --------------------------- +Literal percent/tilde and separator-like characters are encoded reversibly so logical keys cannot collide. -RxDB already defines `RxStorage` as its storage-engine abstraction. An RxCollection adds document behavior, indexes, conflict -handling, and the selected RxStorage implementation. +## RxDB -`createRxDbAdapter()` therefore accepts an existing collection instead of implementing another RxStorage engine. The exported -`RxDbRecordJsonSchema` uses canonical `path` as the primary key and indexes `parent` for direct-child listing. +The forward driver targets an injected `RxCollection`: ```text -RxDB application +chosen RxStorage | - v -RxCollection + RxCollection | - +--> selected RxStorage + createRxDbDriver | - v -createRxDbAdapter() + record adapter | - v + FileSystemType +``` + +RxDB remains responsible for the selected `RxStorage`, document revision/conflict behavior, wrappers, replication, +multi-instance coordination, and licensing. + +`RxDbRecordJsonSchema` defines the collection shape expected by this integration: + +```text +path primary key +parent indexed direct-parent path +name +kind +data +size +lastModified +mediaType +``` + +The path fields have an explicit length ceiling because RxDB indexed string fields need a declared maximum length in the +supported schema shape. The driver rejects an oversized path before it asks the collection to query or write it. + +### Why there is no reverse RxDB bridge yet + +RxDB's storage-engine contract is `RxStorage`, not a key/value object. A real implementation must own semantics such as: + +- bulk writes with per-document conflict results; +- prepared Mango queries and counts; +- attachments; +- changed-document checkpoints; +- change streams; +- cleanup of deleted documents; +- multi-instance behavior where applicable; +- close and remove lifecycle. + +`FileSystemType` does not provide those semantics automatically. A future `bridge/rxdb` should therefore be a +substantial RxStorage implementation with its own tests and performance model, not a wrapper that renames filesystem +methods. + +## db0 + +The db0 direction begins with an already connected `Database`: + +```text +db0 connector + | + Database + dialect + | +createDb0Driver + | +record adapter + | FileSystemType ``` -This keeps the adapter compatible with the collection regardless of whether the application selected memory, IndexedDB, OPFS, -filesystem, SQLite, remote, worker, or another RxStorage family. RxDB keeps ownership of replication, multi-instance behavior, -licensing, storage wrappers, and conflicts. +The current dialect branches are: -db0 and direct SQLite share the same SQL record model ------------------------------------------------------- +```text +sqlite +libsql +postgresql +mysql +``` -`createDb0Adapter()` targets db0's high-level `Database` contract and its reported dialect. The SQL generation currently covers -SQLite, libSQL, PostgreSQL, and MySQL branches. It depends on database behavior rather than connector names, so a new db0 -connector does not need a new OPFS adapter when it presents the same database contract. +The driver generates the table/CRUD SQL required for the selected dialect. It targets db0's portable database/statement +contract rather than connector names. -`createSqliteAdapter()` is the focused direct SQLite path for applications that already own a small connected statement API. It -reuses the SQLite branch of the same SQL record contract instead of maintaining a second table layout and upsert implementation. +The path primary key uses a SHA-256 identity while the original path is retained separately. That avoids assuming +arbitrary long text is a portable primary-key type across the supported SQL families. -The default SQL table is `opfs_entries`. The path identity and parent path are stored separately so direct-child listing can use -a provider-appropriate index. The db0 path uses portable parameter placeholders and lets the db0 connector translate them where -its dialect requires a different native parameter shape. +The table is initialized only when requested. The injected database remains caller-owned unless `disposeDatabase` is +true. -Drizzle keeps schema ownership with the application ---------------------------------------------------- +There is no reverse db0 bridge because a filesystem cannot implement arbitrary SQL parsing, query planning, +transactions, dialects, schema metadata, or connector behavior. -Drizzle spans several SQL dialects and runtime drivers. A universal OPFS-owned Drizzle table would either choose one dialect or -hide dialect-specific DDL details. +## Drizzle -`createDrizzleAdapter()` therefore receives: +Drizzle is also a forward record driver, but the schema stays caller-owned: -1. an already-connected Drizzle database; -2. a table built for that database dialect; -3. the required logical columns. +```text +Drizzle database + table + | + createDrizzleDriver + | + record adapter + | + FileSystemType +``` -Required logical fields are: +The caller provides columns for: ```text path @@ -165,102 +196,123 @@ lastModified mediaType ``` -`path` must be unique or primary. `size` and `lastModified` must round-trip JavaScript safe integers. The bridge uses the common -select/insert/delete builder shape and keeps Drizzle an optional peer dependency. +`path` must be unique or a primary key. `size` and `lastModified` must round-trip JavaScript safe integers. -Inside one `FileSystemType`, normal coordination serializes same-path mutations. Separate processes or hosts are not serialized -by an in-memory facade lock. A database-backed deployment that needs cross-process replacement atomicity must use transactions, -leases, advisory locks, or another mechanism provided by its actual database/driver. +The generic driver uses Drizzle's common CRUD surface. Replacement is delete then insert, so the driver reports +best-effort replacement. This route is serialized inside one cooperating `FileSystemType`, but it is not an atomic +cross-process database replacement. -S3-compatible providers are configured by capability, not brand guesses ------------------------------------------------------------------------- +A database-specific Drizzle driver can expose stronger upsert, transaction, binary, and partition behavior without +changing the portable generic integration. -The direct S3 client is meant to work with AWS S3 and compatible XML/SigV4 services, but "S3-compatible" is not a promise that -all operations, preconditions, limits, checksums, or control-plane features are identical. +### Drizzle-backed filesystem versus Drizzle over OPFS-backed SQLite -The client options deliberately separate the parameters that compatible services vary: +These architectures are opposite directions. + +**A. Database-backed filesystem** is implemented now: ```text -endpoint -bucket -region -addressing: path | virtual -headers -copy: boolean -conditionalWrite: boolean -partSize -copyPartSize -concurrency -credentials +Drizzle + | +database rows + | +record driver + | +FileSystemType ``` -The safe rule is to read the selected provider's current primary documentation and enable only the capabilities it actually -implements for the operations used by the filesystem. - -A few current examples show why this matters: - -| Provider family | Current nuance that affects this client | -| --- | --- | -| AWS S3 | baseline SigV4, multipart upload, CopyObject, UploadPartCopy, conditional completion | -| Cloudflare R2 | S3-compatible endpoint with its own supported-operation set; `auto` is a documented region value | -| DigitalOcean Spaces | implements a documented subset of the S3 API and its own published object/multipart limits | -| Google Cloud Storage XML API | S3-compatible multipart exists, but documented multipart precondition behavior differs | -| Backblaze B2 S3 API | S3-compatible surface with its own unsupported/changed AWS control-plane features | - -For a Google Cloud Storage XML multipart path that does not support the preconditions expected by optimistic object -read-modify-write, create the client with `conditionalWrite: false`. That does not make append/update magically atomic; it makes -the absence of that safety property explicit. - -Cloudflare R2 and other services can also disable `copy` if their selected endpoint/path does not provide the server-side copy -contract expected by the adapter. The filesystem then falls back to the normal streamed/materialized copy path instead of -calling a native capability that was never real. - -The S3 request escape hatch is intentional ------------------------------------------ - -A filesystem does not need to model every S3 object or bucket feature. The direct client therefore exposes signed -`request(options)` in addition to `ObjectStoreType`. - -Use the filesystem/object layer for portable file behavior. Use the lower-level request API when the application needs a -provider-specific control such as an object-lock header, tag operation, checksum policy, versioning call, or another S3 operation -whose semantics should not be flattened into a generic filesystem method. - -The same principle applies to Azure Blob ----------------------------------------- - -Azure Blob has enough differences that the package implements a native Azure REST client rather than translating Azure through -an S3 compatibility layer. - -`createAzureClient()` supports SAS, Microsoft Entra bearer tokens, Shared Key, and custom-header credentials. Its service version is explicit. Its streamed -writes use Azure block APIs, and its large server-side copies use Put Block From URL when synchronous Copy Blob From URL is too -small. - -The object adapter above Azure is still the same `createObjectAdapter()` used by S3. The provider client owns Azure-specific -HTTP mechanics; the object adapter owns the file/directory translation. - -Choose the integration that matches the resource you already own ------------------------------------------------------------------ - -| Existing application resource | Preferred integration | -| --- | --- | -| browser OPFS root | `createOpfsAdapter()` or `openFileSystem()` | -| host directory | Node, Deno, or Bun adapter | -| unstorage `Storage` | `createUnstorageAdapter()` | -| RxDB collection | `createRxDbAdapter()` | -| db0 `Database` | `createDb0Adapter()` | -| connected SQLite statement API | `createSqliteAdapter()` | -| Drizzle database + table | `createDrizzleAdapter()` | -| Deno KV database | `createDenoKvAdapter()` | -| localStorage/sessionStorage-like Web Storage | `createLocalStorageAdapter()` | -| IndexedDB | `createIndexedDbAdapter()` / `openIndexedDbAdapter()` | -| Cache Storage `Cache` | `createCacheAdapter()` | -| S3-compatible endpoint | `createS3Client()` + `createS3Adapter()` | -| Azure Blob container | `createAzureClient()` + `createAzureAdapter()` | -| custom KV/document layer | `createRecordAdapter()` | -| custom object storage | `createObjectAdapter()` | -| any `FileSystemType` needed as KV | `createKeyValueDriver()` | -| any `FileSystemType` needed by unstorage | `createUnstorageDriver()` | - -Adding an extra ecosystem layer only because this package already has an adapter for it usually makes the system harder to -reason about. Use the direct adapter/client/driver for the abstraction the application already owns, and use bridge descriptors -when code needs to inspect both directions as one integration. +**B. Drizzle over an OPFS-backed SQLite database** requires SQLite below Drizzle: + +```text +application + | +Drizzle ORM + | +SQLite engine + | +SQLite VFS + | +FileSystemType / native OPFS +``` + +The current `driver/sqlite` does not implement a VFS. It stores OPFS logical records inside SQLite rows. + +A future SQLite VFS should be designed against the SQLite engine's actual VFS/file contract. It can then use +`FileSystemType` where that contract can be mapped correctly, including sync-access and locking requirements. Drizzle +can sit above the SQLite engine normally. + +## SQLite + +The current SQLite driver is intentionally direct and small. It consumes an injected connected SQLite database that can +prepare and run statements. + +Use it when the application already has a SQLite database and wants to store a virtual filesystem in rows. + +Do not use it as evidence that arbitrary SQLite WASM engines can already store their database file on this package. That +second capability is future VFS work. + +## Generic key/value bridge + +`createKeyValueBridge(fileSystem)` is a reverse bridge with a deliberately small contract: + +```text +has +get / set +getRaw / setRaw +remove +meta +keys +clear +inspect +plan +getMetrics +dispose +``` + +It is useful for ecosystem adapters that need hierarchical string/raw values but do not need SQL, document queries, or +conflict semantics. + +It exposes the backing filesystem's inspection and plan results rather than inventing a separate capability system. + +## Direction metadata + +`integration/definition` exists so applications and third-party packages can reason about asymmetry without executing +storage constructors. + +```ts +import { defineIntegration } from "@okikio/opfs/integration/definition"; + +const integration = defineIntegration({ + name: "example", + directions: { + toOpfs: { supported: true }, + fromOpfs: { + supported: false, + reason: "The upstream reverse contract requires query semantics.", + }, + }, + toOpfs(source) { + return createExampleAdapter(source); + }, +}); +``` + +There is no global registration. Applications import the integration definitions they want to use. + +## Choosing an integration + +Start from the resource the application already owns. + +| Existing resource | Preferred path | +| --------------------------------------------- | ----------------------------------------- | +| unstorage `Storage` | `driver/unstorage` -> `adapter/unstorage` | +| RxDB `RxCollection` | `driver/rxdb` -> `adapter/rxdb` | +| db0 `Database` | `driver/db0` -> `adapter/db0` | +| Drizzle database + table | `driver/drizzle` -> `adapter/drizzle` | +| connected SQLite database | `driver/sqlite` -> `adapter/sqlite` | +| custom value/document storage | `driver/record` -> `adapter/record` | +| existing `FileSystemType` needed as KV | `bridge/kv` | +| existing `FileSystemType` needed by unstorage | `bridge/unstorage` | + +Do not wrap a resource through an unrelated ecosystem only to reach OPFS. Every additional abstraction adds semantics, +requirements, failure behavior, and measurable overhead. diff --git a/docs/environments.md b/docs/environments.md index 836078e..d78e71c 100644 --- a/docs/environments.md +++ b/docs/environments.md @@ -1,178 +1,193 @@ -Execution environments -====================== +# Execution environments -`@okikio/opfs` keeps the filesystem frontend separate from backend availability. The same core source can compile for Window, -workers, Deno, Bun, and Node while runtime-specific adapters remain on explicit subpaths. +## Purpose -The package does not maintain a browser-brand or runtime-brand behavior table. It asks the current realm or adapter what it can -actually do and preserves the resulting capability/failure information. +`@okikio/opfs` keeps the portable filesystem frontend separate from the runtime/backend driver. The root module remains +safe to import in Window, workers, Deno, Bun, and Node. Runtime-specific code stays on explicit driver/adapter subpaths. -Browser OPFS follows the storage key of the current realm ---------------------------------------------------------- +The package does not select behavior from a runtime-brand table. It probes actual browser capabilities and reads +configured driver capabilities/requirements. -Native browser OPFS requires `navigator.storage.getDirectory()`. +## Browser OPFS -In Window, use the asynchronous facade: +Native browser OPFS begins with `navigator.storage.getDirectory()`. + +The root convenience path is: ```ts -const fileSystem = await openFileSystem(); -await fileSystem.writeFile("/state.json", "{}", { parents: true }); -``` +import { openFileSystem } from "@okikio/opfs"; -Do not assume synchronous access from Window. `openSyncFile()` succeeds only when the actual native file handle exposes the sync -handle API and the adapter reports that capability. +await using fileSystem = await openFileSystem(); +``` -DedicatedWorker is the important worker case because browsers commonly expose synchronous OPFS access there. The library still -probes the handle instead of saying "DedicatedWorker means sync": +The explicit layers are: ```ts -const capabilities = await probeOpfs(); -if (capabilities.syncAccessHandleExposed) { - const file = await fileSystem.openSyncFile("/database.sqlite", { - create: true, - parents: true, - }); - try { - // synchronous random access - } finally { - file.close(); - } -} +import { createFileSystem } from "@okikio/opfs"; +import { createFileAdapter } from "@okikio/opfs/adapter/file"; +import { createOpfsDriver } from "@okikio/opfs/driver/opfs"; + +const root = await navigator.storage.getDirectory(); +const driver = createOpfsDriver(root); +const fileSystem = createFileSystem(createFileAdapter(driver)); ``` -SharedWorker and ServiceWorker use the same asynchronous filesystem APIs when storage is exposed. A ServiceWorker must keep the -browser event alive itself. The filesystem cannot call `event.waitUntil()` on behalf of the application. +### Window + +Use the asynchronous filesystem API. Do not assume sync access is available from Window. + +### DedicatedWorker + +Use asynchronous methods normally. `openSyncFile()` succeeds only when the actual OPFS file handle exposes synchronous +access in that realm. The library probes the method rather than inferring support from the worker type. + +### SharedWorker + +Use the same capability-driven approach. A SharedWorker does not automatically imply synchronous access. + +### ServiceWorker + +Use asynchronous methods and keep the event lifetime explicit: ```ts self.addEventListener("message", (event) => { - event.waitUntil(saveMessage(event.data)); + const work = (async () => { + const fileSystem = await openFileSystem(); + await fileSystem.writeFile("/events/latest.json", "{}", { parents: true }); + })(); + + event.waitUntil(work); }); ``` -Iframes need policy-aware tests, not a blanket promise ------------------------------------------------------ +The filesystem cannot extend a ServiceWorker event lifetime by itself. -A same-origin iframe normally observes storage under the same applicable storage key as its embedding context. +## Iframes and storage partitioning -A third-party iframe can be partitioned by browser privacy/storage policy. The package does not try to escape that policy -implicitly. The optional iframe API is separate because requesting unpartitioned storage, where the browser supports it, belongs -inside an explicit user-activation and permission flow. +A same-origin iframe opens storage for its current storage key normally. -An opaque sandbox can reject storage because it has no usable origin. `probeOpfs()` returns the actual root result and normalized -failure instead of inferring the outcome from the sandbox flag alone. +A third-party iframe can receive partitioned storage according to browser policy. Normal `openFileSystem()` does not +attempt to escape that policy. + +`@okikio/opfs/iframe` contains the explicit Storage Access API-related OPFS request helpers for browsers that expose +them: ```text -iframe starts - | - v -probe actual storage API - | - +-- root opens ------> use selected strategy - | - `-- root rejected ---> preserve normalized reason -> application fallback +supportsUnpartitionedOpfsRequest +requestUnpartitionedFileSystem ``` -Private browsing, packaged `file:` pages, enterprise browser policy, quota, and persistence can also change availability or -lifetime. The package deliberately does not fingerprint private mode or promise OPFS on `file:` URLs. Probe the realm you are -actually running in. +The application owns user activation and permission presentation. + +A sandboxed opaque-origin iframe can reject storage. Use `probeOpfs()` and the actual normalized error rather than +browser-name guessing. + +## Private browsing and quota + +Private/incognito modes can change availability, quota, persistence, and lifetime. The package does not fingerprint +private browsing. -Playwright tests the browser contexts directly ----------------------------------------------- +```text +probe actual storage capability + | + +-- available -> use selected route + `-- unavailable -> inspect problem/error and choose application fallback +``` -The canonical browser suite uses Playwright Test projects for Chromium, Firefox, and WebKit. The test matrix exercises: +Quota is a dynamic fact. A driver/facade should not advertise one fixed unlimited capacity when the browser/provider +does not guarantee it. -| Context or behavior | Chromium | Firefox | WebKit | -| --- | ---: | ---: | ---: | -| Window async OPFS | probe + execute | probe + execute | probe + execute | -| DedicatedWorker async | probe + execute | probe + execute | probe + execute | -| DedicatedWorker sync handle | probe + open | probe + open | probe + open | -| SharedWorker | probe + execute | probe + execute | probe + execute | -| ServiceWorker | black-box + instrumentation | black-box | black-box | -| same-origin iframe | probe + execute | probe + execute | probe + execute | -| cross-origin iframe | observe policy result | observe policy result | observe policy result | -| opaque sandbox | observe policy result | observe policy result | observe policy result | -| fresh context isolation | execute | execute | execute | -| persistent profile reopen | execute | execute | execute | -| abort before commit | execute | execute | execute | -| localStorage / IndexedDB / Cache adapters | execute | execute | execute | +## Web Locks -Playwright's deeper ServiceWorker inspection is Chromium-specific, so only that instrumentation is browser-specific. The actual -ServiceWorker OPFS behavior stays a black-box page-to-worker message test in every browser that exposes the API. +`coordination: "auto"` uses Web Locks when available and otherwise falls back to one-realm local FIFO locks. -Deno keeps the same library contracts with runtime-specific capabilities ------------------------------------------------------------------------- +`web-locks` can coordinate cooperating same-storage-key browser realms using the same lock names. `local` cannot +coordinate a separate tab/worker process. `none` disables library coordination. -Use `@okikio/opfs/adapter/deno` for a host directory. The runtime needs the filesystem permissions required by the configured -root. The adapter does not request permissions or inspect environment variables itself. +## Deno -Deno KV is a separate adapter because it is a record store, not a host filesystem: +Use the convenience adapter: ```ts -import { createFileSystem } from "@okikio/opfs"; -import { createDenoKvAdapter } from "@okikio/opfs/adapter/deno-kv"; +import { createDenoAdapter } from "@okikio/opfs/adapter/deno"; +``` + +or the explicit driver: -const kv = await Deno.openKv("./data.kv"); -const fileSystem = createFileSystem(createDenoKvAdapter(kv)); +```ts +import { createDenoDriver } from "@okikio/opfs/driver/deno"; ``` -Current Deno documentation still marks KV as unstable. Real Deno KV tests therefore use `--unstable-kv`. The production adapter -accepts a structural KV contract, so simply importing the module does not require a global `Deno` object. +The configured `root` becomes virtual `/`. The runtime needs the filesystem permissions selected by the host +application. + +Deno KV is a separate record driver under `driver/deno-kv`. It is not the host-filesystem driver. -Deno KV also has a much smaller physical value limit than an ordinary filesystem file. The adapter exposes that limit and a -partition policy through `inspect()`. With the default `partition: "auto"`, small materialized files stay inline while large -files and unknown-size replacement streams use a manifest plus raw byte parts. Callers that need a one-record layout can set -`partition: "never"`; the adapter then stops advertising its partitioned replacement-stream lane and large values fail -explicitly instead of changing layout. +## Node -Node and Bun use explicit host adapters ---------------------------------------- +Use `driver/node` and `adapter/node`. The driver uses `node:fs`/`node:fs/promises`/`node:stream` through the explicit +runtime subpath and maps virtual paths below one configured host root. -The Node adapter uses `node:fs` and `node:fs/promises`. It supports native streaming reads and writes, ranges, copy, rename, -synchronous random access, and flush. The configured host root is the only host directory intentionally exposed through the -virtual path namespace. +Node supports native streaming, ranges, copy, rename/move, positional writes, and synchronous random access where +implemented by the driver. -The Bun adapter uses Bun file APIs for the direct byte path and Bun's Node-compatible filesystem APIs for directory, update, -copy/move, and synchronous host-file behavior. The same public tests import `node:test`; Bun currently supports the in-process -`node:test` API when those files are run with `bun test`. +The package engine range starts at Node 22.18. A validation host below that version can provide supplemental evidence +but cannot stand in for the declared runtime matrix. -The Bun benchmark keeps two raw file-copy baselines: Node-compatible `copyFile()` and `Bun.write(destination, Bun.file(source))`. -The second path lets Bun select its file-backed Blob fast path. A Bun-only provider benchmark also compares Bun's native -`S3Client` with the AWS SDK baseline, this package's direct SigV4 client, the object adapter, and the filesystem facade. These -measurements are evidence for route selection; they do not make runtime brand part of the portable API contract. +## Bun -Electron can use the Node adapter in a trusted main-process layer. The OPFS-shaped API is not a reason to expose an arbitrary host -root directly to untrusted renderer content. Use an application-specific IPC/service contract or browser OPFS where that matches -the trust model. +Use `driver/bun` and `adapter/bun`. -Object clients are runtime-neutral Web clients, but deployment policy still matters ------------------------------------------------------------------------------------ +The driver resolves Bun lazily. It uses Bun's native file primitives for the paths where they provide a clear benefit +and the Node-compatible file driver for operations that require stronger host-filesystem behavior. -The S3 and Azure clients use Web Fetch, Web Crypto where signing is required, Web Streams, AbortSignal, and focused `@std/*` -packages for concurrency, byte assembly, stream limits, path mapping, and XML. They are -therefore usable across Deno, Bun, Node, browsers, and workers that expose those Web APIs. +Bun runtime tests are required before claiming Bun behavior complete. Structural TypeScript compatibility alone is not +runtime evidence. -That does not make every deployment equally appropriate. +## Electron -A browser calling S3 directly needs CORS rules that permit the required methods and headers. More importantly, long-lived cloud -storage secrets should not be shipped to untrusted browser code. Use short-lived scoped credentials or a trusted server design. +A trusted Electron main process can use the Node file driver. Do not expose an arbitrary host root directly to untrusted +renderer content merely because the public API looks like OPFS. A renderer should use a controlled IPC service or +browser OPFS according to the application's security model. -The same applies to Azure. SAS tokens and bearer tokens should be scoped to the actual client threat model. The library accepts a -refresh function so a long-lived process does not have to freeze one credential at client creation. +## Record/database environments -Provider endpoints can also have runtime-specific network rules. A Cloudflare Worker, browser extension, serverless host, or -corporate browser policy can allow or reject different origins. Those network policies are outside the filesystem abstraction. +localStorage, IndexedDB, Cache Storage, RxDB, unstorage, db0, Drizzle, and SQLite integrations work where their injected +upstream resource and the required Web primitives are available. -Coordination scope is part of the execution environment -------------------------------------------------------- +The driver describes backend requirements. The adapter/facade describes effective filesystem routes. Database-backed +record storage generally cannot claim native streaming unless the driver implements a dedicated byte lane. -`web-locks` can coordinate cooperating browser realms that share the relevant Web Locks namespace. `local` only coordinates one -JavaScript realm. Separate OS processes and hosts need backend-level coordination where same-path atomicity matters. +## Server coordination -Database/object adapters can use provider transactions, ETags, versions, advisory locks, or leases where the provider exposes -them. The facade does not pretend an in-memory lock became distributed merely because the persisted bytes live on a remote -service. +`coordination: "local"` is one JavaScript realm only. It does not serialize two Node/Deno/Bun processes or two machines. +When cross-process same-path atomicity is required, use the actual backend primitive: + +```text +database transaction +advisory lock +lease +provider conditional write +provider-specific lock/serialization +``` + +A facade-local lock cannot upgrade a best-effort database replacement into a distributed transaction. + +## Import safety + +The root package and provider-neutral core do not import runtime-specific implementations automatically. + +Use explicit subpaths for: + +```text +driver/node +driver/deno +driver/bun +driver/deno-kv +driver/s3 +driver/azure +adapter/* +``` -Provider protocol details are kept in [S3 client protocol](./s3.md), [Azure Blob client protocol](./azure.md), and the -[provider integration test guide](./providers.md). Shared Key is intended for trusted server/Azurite contexts because it exposes -the Azure storage account key to the runtime. +No driver reads environment variables or configures global application logging at import time. diff --git a/docs/providers.md b/docs/providers.md index e7d1d8d..f95b6bb 100644 --- a/docs/providers.md +++ b/docs/providers.md @@ -1,232 +1,233 @@ -Object-provider integration tests -================================= +# Provider and filesystem baseline tests -Purpose -------- +## Purpose -The direct S3 and Azure clients own HTTP protocol behavior that a pure mock cannot prove. This guide explains the local provider -fixtures used to test real signing, request routing, ranges, multipart/block state, copy, listing, and filesystem translation -without requiring cloud credentials for every maintainer run. +S3 and Azure Blob have two kinds of external baseline: -Provider containers supplement deterministic protocol tests. They do not replace the Amazon S3 or Azure Blob specifications. -They also do not become a production dependency or a storage abstraction. Testcontainers exists only in development tooling. +1. protocol/API clients such as the AWS SDK and Azure SDK; +2. filesystem clients such as AWS Mountpoint and Azure BlobFuse. -Testcontainers owns provider lifecycle --------------------------------------- +They answer different questions. SDK/provider tests validate object protocol behavior. Filesystem-client benchmarks +compare the performance and semantics of a mature provider-to-filesystem translation against this package's own +driver/adapter/facade stack. -`tests/provider/fixture.ts` uses Testcontainers for Node.js instead of a repository-owned Docker Compose and polling harness: +## Testcontainers owns disposable protocol providers + +`tests/provider/fixture.ts` starts local provider services with Testcontainers for Node.js. ```text -node:test - | - v -ProviderFixture - | - +--> Testcontainers GenericContainer - | | - | `--> SeaweedFS 4.41 - | S3-compatible API - | - `--> @testcontainers/azurite - | - `--> Azurite 3.36.0 - Azure Blob API +node:test / Mitata + | + ProviderFixture + | + +-- SeaweedFS + | S3-compatible endpoint + | + `-- Azurite + Azure Blob endpoint ``` -The images remain pinned: +The fixture uses mapped host ports and Testcontainers-owned readiness/cleanup. There is no repository-owned Docker +Compose lifecycle, fixed provider port, curl readiness loop, or shell trap for these services. + +The provider fixture is test infrastructure only. Public package code does not import Testcontainers. + +## SeaweedFS S3 target + +SeaweedFS is used as one independent S3-compatible implementation. The test config supplies endpoint, bucket, region, +and test credentials to: ```text -chrislusf/seaweedfs:4.41 -mcr.microsoft.com/azure-storage/azurite:3.36.0 +AWS SDK +project S3 client +project S3 driver +project S3 adapter +FileSystemType ``` -Testcontainers chooses free host ports and waits for the exposed service instead of requiring fixed `8333` and `10000` host -ports. The S3 fixture combines listening-port and HTTP readiness. The official Azurite module owns its emulator-specific startup -contract, uses in-memory persistence, skips only Azurite's API-version allow-list check, and exposes the mapped Blob endpoint. +Tests cover the implemented shared contract, including: -SeaweedFS still uses `GenericContainer` because the Testcontainers Node catalog does not provide a SeaweedFS module. The code -uses the official Azurite module because Testcontainers recommends a focused module when one exists instead of duplicating that -container's configuration in every project. +- signed put/head/get/delete; +- byte ranges; +- conditional create/replace where supported by the fixture; +- multipart streamed replacement; +- server-side copy; +- prefix/delimiter listing; +- filesystem directory/object translation. -`ProviderFixture` owns every container it starts. `close()` is idempotent and attempts to stop every owned container even if one -cleanup operation fails. Partial construction also stops SeaweedFS if Azurite cannot start. This keeps acquisition and cleanup in -one place instead of spreading lifecycle work across shell traps, readiness polling, and the test body. +SeaweedFS does not define Amazon S3 semantics. AWS-specific limits and canonical signing behavior remain covered by +deterministic protocol tests and AWS primary documentation. -Testcontainers is an interim compute layer ------------------------------------------- +## Azurite target -The project deliberately does not treat Testcontainers as the final runtime/provider abstraction. Testcontainers Node currently -centers Docker-compatible container runtimes. Its documentation covers Docker directly and documents setup/limitations for -Podman, Colima, and Rancher Desktop. +The official `@testcontainers/azurite` module supplies a Blob endpoint and test account credentials. -That is sufficient for the current local service fixtures. It is not the model for future Apple `container`, WSL containers, -microVMs, cloud VMs, Kubernetes, or other compute providers. A future environment/provider layer can replace how fixtures are -started while preserving this test contract: +The provider suite exercises: ```text -provider fixture - | - +--> endpoint - +--> credentials - +--> lifecycle ownership - `--> diagnostics - -protocol/client tests consume only those facts +Azure SDK +project Azure client +project Azure driver +project Azure adapter +FileSystemType ``` -The provider test therefore does not inspect Docker container IDs or Docker networks after startup. Those are fixture mechanics, -not S3/Azure test semantics. +Coverage includes: -Run the provider suite ----------------------- +- Shared Key authentication against the emulator; +- blob put/head/range/delete; +- block upload and block-list commit; +- server-side copy; +- container listing/prefix translation; +- filesystem translation. -The canonical command is: +Azurite is an emulator. A green Azurite result is interoperability evidence, not proof of every cloud Azure service +version or feature. -```sh -mise run test-providers +## Provider benchmark staircase + +`bench/provider.bench.ts` keeps every layer visible. + +S3: + +```text +AWS SDK + -> project S3 client + -> project S3 driver + -> project object adapter + -> facade metrics:none + -> facade metrics:basic ``` -The task is intentionally small: +Azure: ```text -mise - | - +--> deno ci - | - `--> node --test tests/provider.test.ts - | - `--> Testcontainers owns startup/readiness/cleanup +Azure SDK + -> project Azure client + -> project Azure driver + -> project object adapter + -> facade metrics:none + -> facade metrics:basic ``` -`node:test` remains the repository test runner. Testcontainers supplies resources to the test; it does not become a second test -framework. +These are separate benchmark samples, not one chain executed for each operation. The staircase wording means each result +adds one project layer over the same configured provider. -GitHub Actions calls the same mise task. The workflow installs mise, asks mise for the Deno and Node versions required by the -provider job, and then runs `mise run test-providers`. GitHub Actions owns only job topology, permissions, runner selection, -timeouts, and secrets. It does not duplicate provider startup commands. +`bench/bun-provider.bench.ts` adds Bun's native S3 client as another S3 baseline when Bun is available. -Playwright owns browser lifecycle separately --------------------------------------------- +Container startup and image pull are completed before measured Mitata cases begin. -Testcontainers and Playwright solve different lifecycle problems: +## Physical versus logical metrics -```text -node:test + Testcontainers - S3/Azure service interoperability +The provider client/driver can report physical counters such as requests, retries, responses, failures, and +multipart/block work. The facade reports logical filesystem operations and facade buffering. -Playwright Test - Chromium/Firefox/WebKit runtime interoperability - Window/Worker/ServiceWorker/iframe/storage lifecycle -``` +A benchmark should not infer provider requests from logical operations. One logical large write can become many physical +parts. + +## AWS Mountpoint baseline -The browser matrix stays under `tests/browser/`. Playwright owns browser installation, isolated `BrowserContext` instances, -persistent profiles, traces, retries, and Vite fixture-server lifecycle. Provider tests do not launch browsers, and browser tests -do not launch provider containers merely to share a framework. +AWS Mountpoint is a separate filesystem-client comparator. It translates file operations to S3 and deliberately supports +a subset of ordinary filesystem semantics. -Provider benchmarks keep startup outside timed work ----------------------------------------------------- +The package does not assume Mountpoint is a generic S3-compatible implementation. The benchmark target should use Amazon +S3 when that is the intended conformance/performance comparison. A custom endpoint can be used for local experimentation +when Mountpoint supports the selected endpoint configuration, but that does not turn the result into an AWS-supported +compatibility claim. -The same Testcontainers fixture owns provider startup for: +Mountpoint setup/mount lifecycle is external to the normal Testcontainers provider fixture because it is a host +FUSE/system client. After the mount exists, set: ```sh -mise run bench-providers +OPFS_MOUNTPOINT_S3_ROOT=/path/to/mount +mise run bench-filesystem-clients ``` -`bench/providers.ts` starts SeaweedFS and Azurite once, obtains their random host endpoints, and passes those endpoints to the -actual benchmark programs. Container startup, image pull, and readiness time therefore do not enter a Mitata sample. - -S3 is measured as: +The benchmark then measures: ```text -AWS SDK v3 baseline -Bun native S3Client baseline -@okikio/opfs direct SigV4 client -direct ObjectStore adapter -FileSystemType with metrics none -FileSystemType with metrics basic +raw Node fs against the mount +Node file driver against the mount +file adapter against the driver +FileSystemType against the adapter ``` -Azure is measured as: +This shows the project overhead when the provider-to-filesystem translation is owned by Mountpoint rather than by the +package's S3 client/driver. + +## Azure BlobFuse baseline + +BlobFuse is the analogous Azure Blob filesystem-client comparator. It also has caching/configuration semantics that can +change observable filesystem behavior and performance. + +Provision/mount BlobFuse externally, then run: + +```sh +OPFS_BLOBFUSE_ROOT=/path/to/mount +mise run bench-filesystem-clients +``` + +The same raw -> driver -> adapter -> facade staircase is measured. + +Caching mode and write mode must be recorded with benchmark results. A cached BlobFuse read is not semantically +equivalent to a direct uncached REST read simply because both return the same bytes in one sample. + +## Why filesystem-client mounts are not a default CI job + +Mountpoint and BlobFuse need operating-system packages, FUSE support, mount permissions, and cleanup. Those requirements +are materially different from a disposable HTTP test container. + +The repository therefore provides a canonical benchmark program and mise task, while the runner owns system-level mount +setup. A dedicated privileged benchmark runner can automate installation/mounting without making ordinary pull-request +CI depend on FUSE privileges. + +## Comparable operations only + +Filesystem clients do not necessarily implement all POSIX operations, and object stores do not naturally have POSIX +semantics. The benchmark therefore starts from supported operations rather than treating an unsupported operation as a +slow operation. + +For each baseline, record: ```text -@azure/storage-blob baseline -@okikio/opfs direct Azure REST client -direct ObjectStore adapter -FileSystemType with metrics none -FileSystemType with metrics basic +operation +native/support status +cache mode +write mode +payload size +concurrency +elapsed/throughput +request/physical metrics when available ``` -Small replacement and multipart/block cases remain separate. Different request plans must not be reported as facade overhead. -Loopback results are diagnostics about client/abstraction cost, not cloud-throughput claims. - -What the live tests prove -------------------------- - -The S3 path validates: - -- a real SigV4 HTTP request accepted by an independent S3-compatible server; -- PUT and HEAD; -- byte-range GET; -- create-only conditional replacement; -- multipart stream upload with a legal non-final part size; -- provider-side copy; -- prefix listing; -- delete cleanup; -- `ObjectStoreType -> AdapterType -> FileSystemType` translation. - -The Azure path validates: - -- Shared Key accepted by Azurite; -- explicit container creation; -- Put Blob and Get Blob Properties; -- byte-range GET; -- create-only conditional replacement; -- Put Block / Put Block List streaming upload; -- same-account server-side copy; -- prefix listing; -- delete cleanup; -- `ObjectStoreType -> AdapterType -> FileSystemType` translation. - -The provider receives the actual headers, query fields, bytes, XML, and signatures created by the library. - -What the live tests do not prove --------------------------------- - -A compatible server or emulator cannot prove every detail of a cloud service. The suite does not use it as the oracle for: - -- exact AWS canonical-request text; -- AWS-only HTTP-200 embedded error bodies; -- every S3 service limit; -- Azure Shared Key construction independently of Azurite; -- every historical Azure service-version size limit; -- cloud role/identity acquisition; -- region routing and redirects; -- provider throttling; -- cloud durability or consistency guarantees; -- billing, retention, encryption, replication, or versioning. - -Those cases belong to deterministic protocol tests where possible and opt-in real-cloud suites when a local provider cannot -represent the behavior faithfully. - -Why the matrix keeps both test styles -------------------------------------- - -| Test style | Strong at | Weak at | -| ---------- | --------- | ------- | -| Deterministic request test | Exact canonical text, headers, limits, branch selection | Real HTTP parser/auth integration | -| Testcontainers provider | Real socket/HTTP/auth/protocol interoperability | Complete cloud parity | -| Optional real cloud | Actual provider behavior | Cost, credentials, availability, reproducibility | - -A client change that affects signing, multipart/block state, copy, conditions, retries, cancellation, or provider errors should -update the deterministic test and the provider integration when the local implementation can represent that behavior. - -Future provider breadth ------------------------ - -The current fixture is deliberately small. Additional S3-compatible services should be added only when they exercise a materially -different contract, not to increase a provider count. Useful differences include addressing, copy/condition support, multipart -errors, non-AWS region behavior, and pagination. - -Network-fault fixtures are also a useful next layer. Testcontainers provides a Toxiproxy module, which can be used to prove -retry, timeout, cancellation, and cleanup behavior against a real socket path without adding unreliable sleeps to the tests. -That belongs in a focused failure suite rather than the normal happy-path provider test. +A performance comparison is valid only when the operation semantics being compared are close enough to answer the same +question. + +## Real cloud suites + +Local provider fixtures should be supplemented by opt-in cloud suites before high-confidence releases of protocol +changes. + +Amazon S3: + +- short-lived credentials; +- disposable bucket/prefix; +- explicit cleanup; +- same deterministic operation set used locally where service semantics match. + +Azure Blob: + +- short-lived identity/SAS credentials; +- disposable container/prefix; +- explicit cleanup; +- service-version coverage relevant to the changed code. + +Cloud credentials must never become required for ordinary portable tests. + +## Failure injection + +Network fault injection is still future work. Toxiproxy or another socket-level test service can add latency, reset, +timeout, retry, cancellation, and multipart/block cleanup scenarios around the real provider clients. + +It should remain test-fixture infrastructure. The public OPFS/client/driver architecture must not depend on a +fault-injection service. diff --git a/docs/releasing.md b/docs/releasing.md index fa75192..758d9ca 100644 --- a/docs/releasing.md +++ b/docs/releasing.md @@ -1,111 +1,185 @@ -Release process -=============== +# Release process -`@okikio/opfs` publishes one Git release to JSR and npm. Git history decides the version. `deno.json` owns the authored package graph and public exports. semantic-release does not commit generated version files back to `main`. +## Purpose -Conventional commits --------------------- +The repository has one command authority: mise. GitHub Actions decides when a release/publish operation should run and +supplies GitHub-specific permissions, refs, secrets, and outputs. The actual project commands live under `.mise/tasks/`. -Commits merged to `main` use Conventional Commits 1.0.0. Release-relevant examples are: +```text +GitHub Actions + | + jdx/mise-action + | + mise task + | +Deno / Node / Bun / npm / semantic-release +``` + +## Conventional commits + +Commits merged to `main` use Conventional Commits. ```text -fix: correct a public behavior -> patch -feat: add a compatible capability -> minor -feat!: replace a public contract -> major +fix: ... patch +feat: ... minor +feat!: ... major -BREAKING CHANGE: describe the consumer impact -> major +BREAKING CHANGE: ... ``` -`build`, `chore`, `ci`, `docs`, `refactor`, `style`, and `test` are valid types. They do not cause a release unless the analyzed commit carries a breaking change. Pull requests should normally squash to one conventional commit so `main` remains a clear release input. +Commit validation runs through: -The checked-in `0.0.1` is development metadata, not the first public release decision. semantic-release does not support selecting `0.0.1` as the initial stable version. With no earlier release tag, the normal first release on `main` is `1.0.0`. If the project is not ready for a stable `1.0.0`, configure a prerelease branch before enabling the release workflow. Registry commands receive the semantic-release version through `--set-version`, so they do not publish the development placeholder. +```sh +BASE_SHA=... HEAD_SHA=... mise run commits +``` -Dependency graph ----------------- +The mise task pins the commitlint CLI package used by CI. Workflow YAML does not own a second commitlint command. -Dependency changes are not complete until both package-manager views are reproducible. After changing `deno.json` or -`package.json`, update the Deno lockfile intentionally with: +## Quality before release + +The normal release evidence includes: + +```sh +mise run quality +mise run test-deno +mise run test-node +mise run test-bun +mise run test-browser +mise run test-providers +mise run verify-npm +``` + +CI also runs benchmark smoke jobs so gross performance/runtime regressions are visible before a semantic release can +run. + +The quality task includes strict checks, lint, documentation lint, formatting, stress, coverage, and package dry-runs. + +## Lockfiles + +Dependency changes are incomplete until both package-manager views are reproducible. + +After dependency/import changes, regenerate with the canonical managers: ```sh deno install --frozen=false pnpm install ``` -Review and commit both `deno.lock` and `pnpm-lock.yaml`. CI uses `deno ci`, which deliberately rejects a missing or stale Deno -lockfile instead of resolving an unseen dependency graph during the build. +Review and commit `deno.lock` and `pnpm-lock.yaml`. Do not hand-edit integrity/resolution data merely because one +validation host cannot access the registry. + +## Semantic release + +The release workflow starts only after the successful `CI` workflow for `main`. + +It checks out the tested commit, verifies that the checked-out `main` commit still equals the CI commit, then runs: + +```sh +mise run release +``` +The mise task invokes pinned semantic-release tooling. semantic-release determines the version from Git history, creates +the Git tag, and creates the GitHub release according to repository configuration. -Release flow ------------- +The repository uses tags shaped as: ```text -push to main - | - v -CI - | frozen dependencies, format, lint, docs, types, - | runtime tests, browser matrix, stress, package dry-runs - v -Release workflow - | - +-- verify main still equals the tested CI SHA - +-- semantic-release reads commits since opfs@ - +-- semantic-release creates the Git tag and GitHub Release - | - `-- only when a new tag appeared: - workflow_dispatch Publish Registries - | - +----> JSR - `----> npm +opfs@1.2.3 ``` -The registry workflow is dispatched explicitly instead of relying on the GitHub `release` event. A release created with the repository `GITHUB_TOKEN` does not normally trigger another workflow from that release event. GitHub does allow `workflow_dispatch` events created with `GITHUB_TOKEN`, so the release workflow can start the separate publisher without a long-lived release PAT. +## Registry publication uses the immutable tag -A rerun on a commit that already has an `opfs@` tag does not dispatch another automatic publication. Manual registry retries remain available through `Publish Registries` and require the existing immutable tag. +The publish workflow receives an existing release tag. It never publishes arbitrary `main` state. -semantic-release owns only version analysis, tag creation, release notes, and the GitHub Release. It does not publish either registry and it does not create a release commit. +```text +release workflow + | +created opfs@X.Y.Z + | +dispatch publish.yml(tag=opfs@X.Y.Z) + | +checkout exact tag + | +JSR and/or npm +``` -npm packaging -------------- +A partial registry failure can be retried against the same immutable tag by selecting one target. -JSR consumes the authored TypeScript package directly. npm receives generated JavaScript and declaration files. +## JSR -`.mise/tasks/npm` runs `deno pack --no-deno-shim --set-version`. Deno derives the npm graph and exports from the same `deno.json` used by JSR. Drizzle needs one npm-only correction: it is an optional integration, so the task removes `drizzle-orm` from generated normal dependencies and writes it as an optional peer before the final `npm pack`. +JSR publication runs: -The generated tarball is then installed into a clean consumer. The verifier checks exports and declarations, Node/Deno/Bun imports when those runtimes are present, browser bundling, and the optional Drizzle subpath after the peer is installed. +```sh +RELEASE_VERSION=X.Y.Z mise run publish-jsr +``` -Registry setup --------------- +The task performs a frozen Deno install, a JSR dry-run, then the actual publish with the release version supplied +explicitly. -Before the first release: +## npm -1. Create `@okikio/opfs` on JSR and link it to `okikio/opfs` for GitHub OIDC publication. -2. Bootstrap the first npm publication with a granular/automation token in `NPM_TOKEN` when trusted publisher settings are not yet available for the package. -3. After npm has the package, configure trusted publishing for `.github/workflows/publish.yml`. -4. Remove `NPM_TOKEN` when the bootstrap fallback is no longer wanted. +npm publication runs: -The publisher uses GitHub OIDC for JSR and for normal npm trusted publication. npm is upgraded to a current npm 11 release in the publish job before trusted publishing. +```sh +RELEASE_VERSION=X.Y.Z mise run publish-npm +``` -Partial failure ---------------- +The task: -If one registry succeeds and the other fails, do not create another version. Run `Publish Registries` manually with the same existing `opfs@` tag and select only the failed registry. +1. resolves bootstrap token versus trusted-publishing mode; +2. runs `mise run verify-npm`; +3. builds the npm package from the Deno package graph; +4. verifies the tarball from Node/Deno/Bun consumers; +5. publishes the exact generated tarball with provenance. -Local release checks --------------------- +`drizzle-orm` remains an optional peer in the generated npm package rather than being converted into a mandatory runtime +dependency. -A Deno-capable machine can run: +## Bootstrap versus trusted npm publishing -```sh -mise install -mise run check -mise run test -mise run bench -deno task release:check +The publish workflow can choose: + +```text +auto +trusted +token ``` -The GitHub release and registry workflows use mise for their Node, Deno, and Bun toolchain too. Release-only policy remains in -GitHub Actions: OIDC permissions, immutable tag resolution, semantic-release, registry authentication, and the final publish -commands are deployment concerns rather than reusable repository tasks. +`auto` checks whether `@okikio/opfs` already exists on npm through `mise run npm-exists`. + +- first publication: token mode; +- later publications: trusted mode when configured. + +The workflow owns secret selection. The npm publish command still lives in the mise task. + +## Package verification + +`mise run verify-npm` builds a test-version tarball and executes the package consumer verifier. The verifier must +exercise the published package, not source-tree relative imports. + +A release is not accepted because `npm pack` returned success. The generated manifest, exports, optional peers, and +runtime consumer entrypoints must also be checked. + +## Version metadata + +The checked-in development version is not the release decision. The immutable Git tag and semantic-release result define +the published version. Packaging receives that version explicitly instead of committing generated version edits back to +`main`. + +## Recovery rules + +If GitHub release creation succeeds but one registry publish fails: + +1. do not create a replacement tag; +2. do not rebuild from a newer `main` commit; +3. re-run `Publish Registries` with the same immutable tag; +4. select only the failed registry when appropriate. + +If validation fails before release creation, fix the source and produce a new commit. Do not weaken a quality task +solely to make the release workflow progress. + +## Trusted npm publishing -`deno task release:check` performs package dry-runs. It does not upload a release. +The publish job installs `npm:npm@11.18.0` through mise. npm trusted publishing requires npm CLI 11.5.1 or later and +Node 22.14.0 or later. The explicit npm pin prevents the release path from depending on the bundled npm version of the +selected Node release. diff --git a/docs/s3.md b/docs/s3.md index 349162d..8537658 100644 --- a/docs/s3.md +++ b/docs/s3.md @@ -1,25 +1,24 @@ -S3 client protocol guide -======================== +# S3 client protocol guide -Purpose -------- +## Purpose -This document defines the S3 protocol contract implemented by -`@okikio/opfs/s3`. It is written for maintainers who need to change signing, -upload, copy, listing, conditional-write, or compatibility behavior without -silently changing the public filesystem guarantees. +This document defines the S3 protocol contract implemented by `@okikio/opfs/s3`. It is written for maintainers who need +to change signing, upload, copy, listing, conditional-write, or compatibility behavior without silently changing the +public filesystem guarantees. -The client is intentionally not a replacement for the complete AWS SDK. It -implements the S3 REST operations needed by the object-store adapter and keeps a -low-level signed `request()` method for S3-compatible features that do not belong -in the portable filesystem API. +The client is intentionally not a replacement for the complete AWS SDK. It implements the S3 REST operations needed by +the S3 object driver and keeps a low-level signed `request()` method for S3-compatible features that do not belong in +the portable filesystem API. The implementation is direct: ```text -ObjectStoreType / S3ClientType - | - v +S3ClientType + | + +--> direct protocol use + | + `--> S3 object driver -> object adapter -> FileSystemType + request construction | v @@ -32,40 +31,35 @@ Web Fetch `--> S3-compatible endpoint ``` -The protocol code uses Web Crypto for SHA-256 and HMAC-SHA256. `@std/encoding` -owns hexadecimal encoding, `@std/async/pool` owns bounded multipart concurrency, -`@std/xml` owns XML parsing/serialization, and `@std/bytes` supports the shared -chunk layer. No AWS SDK package is imported. +The protocol code uses Web Crypto for SHA-256 and HMAC-SHA256. `@std/encoding` owns hexadecimal encoding, +`@std/async/pool` owns bounded multipart concurrency, `@std/xml` owns XML parsing/serialization, and `@std/bytes` +supports the shared chunk layer. No AWS SDK package is imported. This document distinguishes three evidence classes: - - **Implemented** means the current source contains the behavior. - - **Protocol** means the behavior is required or described by current AWS S3 - documentation. - - **Provider-dependent** means an S3-compatible service can intentionally - differ and the client exposes configuration for that difference. - -The current implementation was reviewed against AWS S3 REST documentation on -August 14, 2026. The authoritative source remains AWS documentation when this -file and the service specification disagree. +- **Implemented** means the current source contains the behavior. +- **Protocol** means the behavior is required or described by current AWS S3 documentation. +- **Provider-dependent** means an S3-compatible service can intentionally differ and the client exposes configuration + for that difference. +The current implementation was reviewed against AWS S3 REST documentation on August 14, 2026. The authoritative source +remains AWS documentation when this file and the service specification disagree. -The client owns a focused S3 contract -------------------------------------- +## The client owns a focused S3 contract -`S3ClientType` extends the package `ObjectStoreType`, so it provides the object -operations the filesystem adapter needs: +`S3ClientType` implements the object-backend operations consumed by `driver/s3`, while retaining S3-specific protocol +methods. The object operations used by the driver are: -| Object operation | S3 REST operation | Current behavior | -| ---------------- | ----------------- | ---------------- | -| `head()` | `HeadObject` | Exact-key metadata or `null` for `404` | -| `get()` | `GetObject` | Complete object or one `Range` | -| `put(Uint8Array)` | `PutObject` | One materialized request up to 5 GB | -| `put(stream)` | Multipart upload | Bounded concurrent parts and explicit completion | -| `delete()` | `DeleteObject` | Missing object is already removed | -| `list()` | `ListObjectsV2` | Prefix, delimiter, limit, continuation token | -| `copy()` small | `CopyObject` | Server-side copy through the 5 GB single-copy limit | -| `copy()` large | Multipart `UploadPartCopy` | Server-side ranged copy without JS body transfer | +| Object operation | S3 REST operation | Current behavior | +| ----------------- | -------------------------- | --------------------------------------------------- | +| `head()` | `HeadObject` | Exact-key metadata or `null` for `404` | +| `get()` | `GetObject` | Complete object or one `Range` | +| `put(Uint8Array)` | `PutObject` | One materialized request up to 5 GB | +| `put(stream)` | Multipart upload | Bounded concurrent parts and explicit completion | +| `delete()` | `DeleteObject` | Missing object is already removed | +| `list()` | `ListObjectsV2` | Prefix, delimiter, limit, continuation token | +| `copy()` small | `CopyObject` | Server-side copy through the 5 GB single-copy limit | +| `copy()` large | Multipart `UploadPartCopy` | Server-side ranged copy without JS body transfer | The S3-specific surface also exposes: @@ -77,14 +71,45 @@ completeUpload() abortUpload() ``` -Those operations are public because an S3 consumer can need storage class, -checksums, encryption, object lock, tagging, or another provider extension that -is not a filesystem concern. The low-level request method signs the request but -does not interpret every S3 feature on the caller's behalf. +Those operations are public because an S3 consumer can need storage class, checksums, encryption, object lock, tagging, +or another provider extension that is not a filesystem concern. The low-level request method signs the request but does +not interpret every S3 feature on the caller's behalf. + +The driver remains a separate public layer: + +```ts +import { createS3Client } from "@okikio/opfs/s3"; +import { createS3DriverFromClient } from "@okikio/opfs/driver/s3"; +import { createObjectAdapter } from "@okikio/opfs/adapter/object"; + +const client = createS3Client(options); +const driver = createS3DriverFromClient(client); +const adapter = createObjectAdapter(driver); +``` + +The driver adds provider requirements, hard limits, optimization state, deterministic planning, and physical request +metrics. The adapter then translates object prefixes and values into canonical filesystem primitives. + +## Optimizations are independently controllable + +The S3 client currently exposes two optimization switches. Both are visible through the S3 driver inspection. + +`delayedMultipart` : Defaults to true. For an unknown-length stream, the client buffers the first bounded multipart +chunk before creating an upload. If EOF arrives while the object can use one `PutObject`, the client avoids multipart +initiation. If more data arrives, or the first chunk is already above the single-PUT ceiling, it starts multipart and +replays the retained bytes into the multipart lane. Set this to false when exact multipart request lifecycle is required +for testing or application policy. + +`signingKeyCache` : Defaults to true. The client caches the derived SigV4 signing key for the current +credentials/date/region/service tuple. The cache is per client and invalidates when the tuple changes. Disabling it +forces key derivation on every signed request. The optimization changes CPU work but not the canonical request or +resulting signature. +The client reports both switches through `client.optimizations`, while `driver.inspect().optimizations` exposes the same +state in the generic driver model. Any future optimization that changes request count, provider resource lifetime, +failure timing, or storage layout must also be independently disableable. -Addressing and canonical request construction ---------------------------------------------- +## Addressing and canonical request construction The client supports path-style and virtual-hosted-style addressing. @@ -100,30 +125,28 @@ Virtual-hosted style: https://bucket.endpoint.example/path/to/object ``` -`addressing: "path"` is the compatibility-oriented default because many local -and third-party S3 implementations expose one HTTP endpoint without wildcard -DNS for buckets. +`addressing: "path"` is the compatibility-oriented default because many local and third-party S3 implementations expose +one HTTP endpoint without wildcard DNS for buckets. -The client preserves slash separators in object keys while percent-encoding each -path segment. Signature Version 4 is sensitive to exact path and query -serialization. The implementation therefore does not use locale-sensitive -sorting and does not normalize the canonical URI after object-key construction. +The client preserves slash separators in object keys while percent-encoding each path segment. Signature Version 4 is +sensitive to exact path and query serialization. The implementation therefore does not use locale-sensitive sorting and +does not normalize the canonical URI after object-key construction. For the canonical query string, the client: -1. expands repeated query values; -2. URI-encodes each name and value; -3. sorts the encoded names and then encoded values by code-unit order; -4. joins the pairs with `&`. +1. expands repeated query values; +2. URI-encodes each name and value; +3. sorts the encoded names and then encoded values by code-unit order; +4. joins the pairs with `&`. For canonical headers, the client: -1. lowercases header names; -2. normalizes internal whitespace; -3. sorts the signed header names; -4. includes the request authority as `host` even though browser Fetch controls - the actual Host or HTTP/2 `:authority` field; -5. includes required `x-amz-*` headers. +1. lowercases header names; +2. normalizes internal whitespace; +3. sorts the signed header names; +4. includes the request authority as `host` even though browser Fetch controls the actual Host or HTTP/2 `:authority` + field; +5. includes required `x-amz-*` headers. The signing flow is: @@ -148,64 +171,52 @@ kDate -> kRegion -> kService(s3) -> kSigning Authorization signature ``` -Credential sources can be static or refreshable. Refreshable credentials are -resolved immediately before each signed request. Session credentials add -`x-amz-security-token` before canonical signing. +Credential sources can be static or refreshable. Refreshable credentials are resolved immediately before each signed +request. Session credentials add `x-amz-security-token` before canonical signing. -The client uses Web Crypto directly because browser-compatible SHA-256 and HMAC -are already platform APIs. `@std/crypto` does not implement S3 Signature -Version 4, so adding it would not remove protocol code or improve ownership. +The client uses Web Crypto directly because browser-compatible SHA-256 and HMAC are already platform APIs. `@std/crypto` +does not implement S3 Signature Version 4, so adding it would not remove protocol code or improve ownership. ### Payload hashes -Replayable Web request bodies are SHA-256 hashed before signing. The client -hashes strings, `ArrayBuffer`, `ArrayBufferView`, `Blob`, and -`URLSearchParams` values. Requests without bodies use the standard SHA-256 of -an empty payload. +Replayable Web request bodies are SHA-256 hashed before signing. The client hashes strings, `ArrayBuffer`, +`ArrayBufferView`, `Blob`, and `URLSearchParams` values. Requests without bodies use the standard SHA-256 of an empty +payload. -A low-level caller can supply `payloadHash`. This exists for S3 modes such as -`UNSIGNED-PAYLOAD` where the selected provider accepts that contract. The -client does not silently choose an unsigned payload when the exact request -bytes can be determined without consuming or re-encoding caller state. +A low-level caller can supply `payloadHash`. This exists for S3 modes such as `UNSIGNED-PAYLOAD` where the selected +provider accepts that contract. The client does not silently choose an unsigned payload when the exact request bytes can +be determined without consuming or re-encoding caller state. -A streamed low-level body cannot be consumed once merely to calculate a hash and -then consumed again by Fetch. `FormData` has a related problem because Fetch -owns its multipart boundary and exact wire encoding. Those two low-level body -forms therefore use `UNSIGNED-PAYLOAD` unless the caller supplies an explicit -`payloadHash`. Callers that need AWS streaming-signature chunk framing must -implement that S3-specific streaming mode above `request()`. The normal -high-level streamed `put()` avoids this ambiguity by using multipart parts, -each of which is materialized before its individual signed request. +A streamed low-level body cannot be consumed once merely to calculate a hash and then consumed again by Fetch. +`FormData` has a related problem because Fetch owns its multipart boundary and exact wire encoding. Those two low-level +body forms therefore use `UNSIGNED-PAYLOAD` unless the caller supplies an explicit `payloadHash`. Callers that need AWS +streaming-signature chunk framing must implement that S3-specific streaming mode above `request()`. The normal +high-level streamed `put()` avoids this ambiguity by using multipart parts, each of which is materialized before its +individual signed request. - -Object and multipart limits are part of planning ------------------------------------------------- +## Object and multipart limits are part of planning `S3_LIMITS` records the protocol limits that affect this implementation: -| Limit | Value used by the client | Why it matters | -| ----- | ------------------------ | -------------- | -| Maximum object | 53,687,091,200,000 bytes | Exact 10,000 x 5 GiB multipart ceiling (48.8 TiB) | -| Single `PutObject` | 5,000,000,000 bytes | Larger replacement must use multipart upload | -| Single `CopyObject` | 5,000,000,000 bytes | Larger copy must use multipart copy | -| Minimum multipart part | 5 MiB | Every non-final upload part must meet the S3 minimum | -| Maximum multipart part | 5 GiB | Client part-size configuration cannot exceed it | -| Maximum multipart parts | 10,000 | Known-size streams must choose a large enough part size | - -The distinction between GB and GiB is intentional. AWS documents the -single-request PUT/copy threshold in decimal GB, while multipart part limits use -binary-sized MiB/GiB values. AWS product documentation often calls the maximum -object size 50 TB, while the current object guide and multipart arithmetic make -the exact ceiling 10,000 x 5 GiB = 53,687,091,200,000 bytes (48.8 TiB, about -53.7 TB). The client uses the exact multipart-derived value so it does not +| Limit | Value used by the client | Why it matters | +| ----------------------- | ------------------------ | ------------------------------------------------------- | +| Maximum object | 53,687,091,200,000 bytes | Exact 10,000 x 5 GiB multipart ceiling (48.8 TiB) | +| Single `PutObject` | 5,000,000,000 bytes | Larger replacement must use multipart upload | +| Single `CopyObject` | 5,000,000,000 bytes | Larger copy must use multipart copy | +| Minimum multipart part | 5 MiB | Every non-final upload part must meet the S3 minimum | +| Maximum multipart part | 5 GiB | Client part-size configuration cannot exceed it | +| Maximum multipart parts | 10,000 | Known-size streams must choose a large enough part size | + +The distinction between GB and GiB is intentional. AWS documents the single-request PUT/copy threshold in decimal GB, +while multipart part limits use binary-sized MiB/GiB values. AWS product documentation often calls the maximum object +size 50 TB, while the current object guide and multipart arithmetic make the exact ceiling 10,000 x 5 GiB = +53,687,091,200,000 bytes (48.8 TiB, about 53.7 TB). The client uses the exact multipart-derived value so it does not reject objects that S3 can legally assemble. -`partSize` defaults to 8 MiB. `copyPartSize` defaults to 1 GiB. Both are -validated when the client is created. +`partSize` defaults to 8 MiB. `copyPartSize` defaults to 1 GiB. Both are validated when the client is created. -If `ObjectPutOptionsType.size` is supplied, the upload planner calculates the -minimum part size required to stay at or below 10,000 parts and uses the larger -of that value and the configured `partSize`. +If `ObjectPutOptionsType.size` is supplied, the upload planner calculates the minimum part size required to stay at or +below 10,000 parts and uses the larger of that value and the configured `partSize`. ```text known body size @@ -220,26 +231,21 @@ ceil(size / 10,000) `--> reject if > 5 GiB ``` -When a stream size is unknown, the configured part size remains authoritative. -The chunk iterator fails before part 10,001 instead of creating an invalid -multipart request. A caller with a very large stream should supply `size` so -the planner can choose a safe part size before network work begins. - -The final upload byte count is compared with the declared `size`. A mismatch is -a caller/data-source error and the multipart upload is aborted. +When a stream size is unknown, the configured part size remains authoritative. The chunk iterator fails before part +10,001 instead of creating an invalid multipart request. A caller with a very large stream should supply `size` so the +planner can choose a safe part size before network work begins. -Failure and cancellation do not transfer cleanup authority to the caller. Once -`CreateMultipartUpload` succeeds, the high-level streamed `put()` owns that -upload until completion or abort. If a part fails or the caller cancels, the -client stops consuming the source, waits for admitted part requests to settle, -and then sends `AbortMultipartUpload` with a separate cleanup signal. -`abortTimeoutMs` bounds that best-effort cleanup and defaults to 30 seconds. -Using a separate signal matters because the caller's cancellation signal is -already aborted at exactly the time cleanup becomes necessary. +The final upload byte count is compared with the declared `size`. A mismatch is a caller/data-source error and the +multipart upload is aborted. +Failure and cancellation do not transfer cleanup authority to the caller. Once `CreateMultipartUpload` succeeds, the +high-level streamed `put()` owns that upload until completion or abort. If a part fails or the caller cancels, the +client stops consuming the source, waits for admitted part requests to settle, and then sends `AbortMultipartUpload` +with a separate cleanup signal. `abortTimeoutMs` bounds that best-effort cleanup and defaults to 30 seconds. Using a +separate signal matters because the caller's cancellation signal is already aborted at exactly the time cleanup becomes +necessary. -Multipart upload is a commit protocol -------------------------------------- +## Multipart upload is a commit protocol A streamed replacement follows four protocol stages: @@ -262,53 +268,42 @@ CompleteMultipartUpload `--> failure -> AbortMultipartUpload after active parts settle ``` -`@std/async/pool` owns the concurrency admission. The client does not start an -unbounded Promise for every part. +`@std/async/pool` owns the concurrency admission. The client does not start an unbounded Promise for every part. -Each successful `UploadPart` must return an ETag. The completion document -contains one ordered `PartNumber`/`ETag` record per uploaded part. Before the -client sends that XML, it verifies: +Each successful `UploadPart` must return an ETag. The completion document contains one ordered `PartNumber`/`ETag` +record per uploaded part. Before the client sends that XML, it verifies: - - at least one part exists; - - no more than 10,000 parts exist; - - every part number is an integer in the legal range; - - a part number is not duplicated; - - every ETag is non-empty; - - the final order is ascending by part number. +- at least one part exists; +- no more than 10,000 parts exist; +- every part number is an integer in the legal range; +- a part number is not duplicated; +- every ETag is non-empty; +- the final order is ascending by part number. -The XML document is built with `@std/xml/stringify`; protocol escaping is not a -hand-written string replacement. +The XML document is built with `@std/xml/stringify`; protocol escaping is not a hand-written string replacement. -The completion request can carry destination `If-Match`, `If-None-Match`, and -`x-amz-mp-object-size`. Preconditions belong to the commit stage rather than -multipart initiation because commit is the point where the destination object +The completion request can carry destination `If-Match`, `If-None-Match`, and `x-amz-mp-object-size`. Preconditions +belong to the commit stage rather than multipart initiation because commit is the point where the destination object becomes authoritative. -S3 has an unusual completion failure mode: the service can send HTTP `200 OK` -before final assembly finishes and then put an `` document in the -response body. `completeUpload()` therefore parses the success body and treats -an embedded `` as a failed commit. - -If any streamed part operation fails, the client waits for the pool's already -started requests to settle before it sends `AbortMultipartUpload`. This order -prevents a late `UploadPart` from racing after the abort request. +S3 has an unusual completion failure mode: the service can send HTTP `200 OK` before final assembly finishes and then +put an `` document in the response body. `completeUpload()` therefore parses the success body and treats an +embedded `` as a failed commit. -An abort failure does not replace the original upload failure. The original -operation remains the terminal error because it is what caused cleanup. -Unfinished multipart state can still remain at the provider when both the -operation and cleanup request fail, so production buckets should also use an S3 -lifecycle rule for stale multipart uploads. +If any streamed part operation fails, the client waits for the pool's already started requests to settle before it sends +`AbortMultipartUpload`. This order prevents a late `UploadPart` from racing after the abort request. +An abort failure does not replace the original upload failure. The original operation remains the terminal error because +it is what caused cleanup. Unfinished multipart state can still remain at the provider when both the operation and +cleanup request fail, so production buckets should also use an S3 lifecycle rule for stale multipart uploads. -Server-side copy has two distinct paths ---------------------------------------- +## Server-side copy has two distinct paths -`copy()` first performs `HeadObject` on the source. This confirms existence, -obtains size for planning, and supplies metadata needed by multipart initiation. +`copy()` first performs `HeadObject` on the source. This confirms existence, obtains size for planning, and supplies +metadata needed by multipart initiation. -For a source at or below the single-copy limit, the client sends `CopyObject`. -For a larger source, it creates a multipart upload at the destination and sends -one `UploadPartCopy` request per byte range. +For a source at or below the single-copy limit, the client sends `CopyObject`. For a larger source, it creates a +multipart upload at the destination and sends one `UploadPartCopy` request per byte range. ```text HEAD source @@ -325,73 +320,61 @@ HEAD source CompleteMultipartUpload ``` -Multipart copy chooses a range size large enough to keep the destination at or -below 10,000 parts and rejects a plan that would require a part above 5 GiB. -The range is inclusive because `x-amz-copy-source-range` uses inclusive byte +Multipart copy chooses a range size large enough to keep the destination at or below 10,000 parts and rejects a plan +that would require a part above 5 GiB. The range is inclusive because `x-amz-copy-source-range` uses inclusive byte positions. Source preconditions map to S3 copy-source headers: -| Package option | S3 header | -| -------------- | --------- | -| `sourceIfMatch` | `x-amz-copy-source-if-match` | -| `sourceIfNoneMatch` | `x-amz-copy-source-if-none-match` | -| `sourceIfModifiedSince` | `x-amz-copy-source-if-modified-since` | +| Package option | S3 header | +| ------------------------- | --------------------------------------- | +| `sourceIfMatch` | `x-amz-copy-source-if-match` | +| `sourceIfNoneMatch` | `x-amz-copy-source-if-none-match` | +| `sourceIfModifiedSince` | `x-amz-copy-source-if-modified-since` | | `sourceIfUnmodifiedSince` | `x-amz-copy-source-if-unmodified-since` | -Destination `ifMatch` and `ifNoneMatch` apply directly to `CopyObject` or to the -multipart completion request. They are deliberately removed from individual -`UploadPartCopy` requests because a destination object does not become the +Destination `ifMatch` and `ifNoneMatch` apply directly to `CopyObject` or to the multipart completion request. They are +deliberately removed from individual `UploadPartCopy` requests because a destination object does not become the completed value until commit. -`CopyObject` and `UploadPartCopy` can also encode a service error inside an HTTP -200 response. Both paths use the same success-XML inspection as multipart -completion. +`CopyObject` and `UploadPartCopy` can also encode a service error inside an HTTP 200 response. Both paths use the same +success-XML inspection as multipart completion. - -Listing preserves provider pagination -------------------------------------- +## Listing preserves provider pagination `list()` uses `ListObjectsV2` with these mappings: -| Package field | S3 query field | -| ------------- | -------------- | -| `prefix` | `prefix` | -| `delimiter` | `delimiter` | -| `limit` | `max-keys` | -| `cursor` | `continuation-token` | - -`Contents` records become `ObjectEntryType`. `CommonPrefixes` become child -prefixes. `NextContinuationToken` is returned as the next opaque cursor. +| Package field | S3 query field | +| ------------- | -------------------- | +| `prefix` | `prefix` | +| `delimiter` | `delimiter` | +| `limit` | `max-keys` | +| `cursor` | `continuation-token` | -The object-store filesystem adapter is responsible for consuming pages and -interpreting directory-marker metadata. The S3 client itself does not pretend -that prefixes are native directories. +`Contents` records become `ObjectEntryType`. `CommonPrefixes` become child prefixes. `NextContinuationToken` is returned +as the next opaque cursor. +The S3 object driver and object adapter are responsible for consuming pages and interpreting directory-marker metadata. +The S3 client itself does not pretend that prefixes are native directories. -Conditional writes are capability claims ------------------------------------------ +## Conditional writes are capability claims -The client advertises `conditionalWrite: true` by default because Amazon S3 -honors the preconditions used by this implementation. An S3-compatible service -that ignores or only partially implements these conditions must set +The client advertises `conditionalWrite: true` by default because Amazon S3 honors the preconditions used by this +implementation. An S3-compatible service that ignores or only partially implements these conditions must set `conditionalWrite: false` in `createS3Client()`. -That flag changes filesystem behavior. The object adapter will not claim that a -read-modify-write append/update is protected from concurrent replacement when -the backend cannot enforce the ETag precondition. - -`copy: false` similarly disables server-side copy for an endpoint whose S3 API -does not implement the required copy operations correctly. +That flag changes filesystem behavior. The S3 driver/object adapter will not claim that a read-modify-write +append/update is protected from concurrent replacement when the backend cannot enforce the ETag precondition. -These overrides are explicit because compatibility means "uses the S3 protocol" -not "implements every Amazon S3 behavior". +`copy: false` similarly disables server-side copy for an endpoint whose S3 API does not implement the required copy +operations correctly. +These overrides are explicit because compatibility means "uses the S3 protocol" not "implements every Amazon S3 +behavior". -Failures retain provider evidence ---------------------------------- +## Failures retain provider evidence -Non-success responses are parsed as S3 XML when possible. `S3Error` retains: +Non-success responses are parsed as S3 XML when possible. `S3Error` retains: ```text status @@ -401,114 +384,93 @@ hostId original Response ``` -The original `Response` remains available so a caller can inspect headers and -provider-specific diagnostics that the stable error fields do not model. +The original `Response` remains available so a caller can inspect headers and provider-specific diagnostics that the +stable error fields do not model. -A response body that is not valid S3 XML still produces a failure based on HTTP -status and available text. The client does not convert an unknown provider -response into a fake known S3 error code. +A response body that is not valid S3 XML still produces a failure based on HTTP status and available text. The client +does not convert an unknown provider response into a fake known S3 error code. -Cancellation uses `AbortSignal` on each Fetch request. Multipart cancellation -is not a distributed transaction: cancellation can stop local admission and -abort HTTP work, while an already accepted provider request may still have -created remote multipart state. Cleanup is therefore explicit. +Cancellation uses `AbortSignal` on each Fetch request. Multipart cancellation is not a distributed transaction: +cancellation can stop local admission and abort HTTP work, while an already accepted provider request may still have +created remote multipart state. Cleanup is therefore explicit. -Request retry is explicit and operation-aware. `request` in `S3ClientOptionsType` configures retries, exponential delay, jitter, -and an optional per-attempt timeout. The implementation uses `@std/async/retry` for 408, 429, 5xx, and transport failures. -Authorization is rebuilt for every attempt so refreshable credentials and SigV4 timestamps are current. Signed redirects are -manual and are returned to the caller instead of being followed to another authority. +Request retry is explicit and operation-aware. `request` in `S3ClientOptionsType` configures retries, exponential delay, +jitter, and an optional per-attempt timeout. The implementation uses `@std/async/retry` for 408, 429, 5xx, and transport +failures. Authorization is rebuilt for every attempt so refreshable credentials and SigV4 timestamps are current. Signed +redirects are manual and are returned to the caller instead of being followed to another authority. -Body replayability is only one admission condition. A one-shot `ReadableStream` receives one attempt. A mechanically replayable -body can still belong to a non-idempotent protocol operation, so low-level `request()` also accepts `retry: false`. The high-level -client disables automatic retry for `CreateMultipartUpload` and `CompleteMultipartUpload` because a lost response can make the -server-side outcome ambiguous. Stable part-number PUTs, reads, deletes, lists, and ordinary replacements use the configured -policy. `request: { retries: 0 }` disables automatic retry client-wide. +Body replayability is only one admission condition. A one-shot `ReadableStream` receives one attempt. A mechanically +replayable body can still belong to a non-idempotent protocol operation, so low-level `request()` also accepts +`retry: false`. The high-level client disables automatic retry for `CreateMultipartUpload` and `CompleteMultipartUpload` +because a lost response can make the server-side outcome ambiguous. Stable part-number PUTs, reads, deletes, lists, and +ordinary replacements use the configured policy. `request: { retries: 0 }` disables automatic retry client-wide. -`getMetrics()` returns direct HTTP request counts, retry counts, terminal failures, response counts, and optional Fetch duration. -Set `metrics: "none"` when measuring the raw protocol path, `basic` for counters, or `timing` for counters plus durations. +`getMetrics()` returns direct HTTP request counts, retry counts, terminal failures, response counts, and optional Fetch +duration. Set `metrics: "none"` when measuring the raw protocol path, `basic` for counters, or `timing` for counters +plus durations. +## Provider compatibility and known non-goals -Provider compatibility and known non-goals ------------------------------------------- - -The direct client supports custom endpoint, region, headers, path/virtual -addressing, copy capability, and conditional-write capability. This is enough -to use Amazon S3 and many S3-compatible products while keeping compatibility -choices visible. +The direct client supports custom endpoint, region, headers, path/virtual addressing, copy capability, and +conditional-write capability. This is enough to use Amazon S3 and many S3-compatible products while keeping +compatibility choices visible. The current client does **not** claim complete coverage of: - - SigV4 streaming chunk signatures; - - presigned URL creation; - - S3 Express directory-bucket session management; - - access points, Object Lambda, or Outposts host construction; - - Multi-Region Access Point SigV4A; - - SSE-C/SSE-KMS convenience APIs; - - checksum negotiation beyond the payload hash needed for signing; - - object tagging, ACLs, retention, legal hold, replication, or lifecycle APIs; - - version-ID aware filesystem paths; - - bucket creation or bucket policy management; - - adaptive throttling and provider-specific `Retry-After` scheduling beyond the shared exponential retry policy. - -A low-level signed request can still reach some provider features when the -caller knows the exact S3 REST contract. A feature should receive a typed -high-level API only after the library can document and test its semantics. +- SigV4 streaming chunk signatures; +- presigned URL creation; +- S3 Express directory-bucket session management; +- access points, Object Lambda, or Outposts host construction; +- Multi-Region Access Point SigV4A; +- SSE-C/SSE-KMS convenience APIs; +- checksum negotiation beyond the payload hash needed for signing; +- object tagging, ACLs, retention, legal hold, replication, or lifecycle APIs; +- version-ID aware filesystem paths; +- bucket creation or bucket policy management; +- adaptive throttling and provider-specific `Retry-After` scheduling beyond the shared exponential retry policy. +A low-level signed request can still reach some provider features when the caller knows the exact S3 REST contract. A +feature should receive a typed high-level API only after the library can document and test its semantics. -Validation strategy -------------------- +## Validation strategy S3 validation is intentionally split into protocol and provider tests. -`tests/s3.test.ts` is deterministic. It uses a controlled Fetch implementation -to inspect exact requests and covers: - - - Signature Version 4 canonicalization; - - deterministic timestamps and credentials; - - minimum/maximum multipart part configuration; - - multipart part ordering and duplicate rejection; - - `x-amz-mp-object-size` at completion; - - HTTP 200 embedded service errors; - - source and destination copy preconditions; - - large-copy multipart planning; - - part-count behavior. - -`tests/provider.test.ts` opens the pinned SeaweedFS S3-compatible server through -Testcontainers. It proves a real HTTP implementation can accept our signed -requests for PUT, HEAD, range GET, conditional create, multipart streaming -upload, copy, listing, delete, and the object-store filesystem adapter. The -fixture uses a random mapped host port so concurrent local runs do not share one -fixed endpoint. - -The provider container is SeaweedFS, not a statement that SeaweedFS defines the -S3 specification. The container proves interoperability with one independent -S3-compatible implementation. AWS-specific wire details remain covered by the +`tests/s3.test.ts` is deterministic. It uses a controlled Fetch implementation to inspect exact requests and covers: + +- Signature Version 4 canonicalization; +- deterministic timestamps and credentials; +- minimum/maximum multipart part configuration; +- multipart part ordering and duplicate rejection; +- `x-amz-mp-object-size` at completion; +- HTTP 200 embedded service errors; +- source and destination copy preconditions; +- large-copy multipart planning; +- part-count behavior. + +`tests/provider.test.ts` opens the pinned SeaweedFS S3-compatible server through Testcontainers. It proves a real HTTP +implementation can accept our signed requests for PUT, HEAD, range GET, conditional create, multipart streaming upload, +copy, listing, delete, the S3 driver, and the object filesystem adapter. The fixture uses a random mapped host port so +concurrent local runs do not share one fixed endpoint. + +The provider container is SeaweedFS, not a statement that SeaweedFS defines the S3 specification. The container proves +interoperability with one independent S3-compatible implementation. AWS-specific wire details remain covered by the deterministic protocol tests and the AWS documentation listed below. -Before release, the maintainer test matrix should also run an opt-in real Amazon -S3 suite with short-lived credentials when CI secret policy permits it. That -suite must use a dedicated disposable bucket/prefix and explicit cleanup. - - -Primary specification sources ------------------------------ - -The implementation and this guide should be checked against these primary AWS -sources when S3 behavior changes: - - - AWS Signature Version 4 canonical request: - https://docs.aws.amazon.com/AmazonS3/latest/API/sig-v4-header-based-auth.html - - Multipart upload limits: - https://docs.aws.amazon.com/AmazonS3/latest/userguide/qfacts.html - - CompleteMultipartUpload: - https://docs.aws.amazon.com/AmazonS3/latest/API/API_CompleteMultipartUpload.html - - CopyObject: - https://docs.aws.amazon.com/AmazonS3/latest/API/API_CopyObject.html - - UploadPartCopy: - https://docs.aws.amazon.com/AmazonS3/latest/API/API_UploadPartCopy.html - - ListObjectsV2: - https://docs.aws.amazon.com/AmazonS3/latest/API/API_ListObjectsV2.html - -Secondary S3-compatible provider documentation can explain provider-specific -configuration, but it must not override the Amazon S3 wire contract when the -client claims Amazon S3 behavior. +Before release, the maintainer test matrix should also run an opt-in real Amazon S3 suite with short-lived credentials +when CI secret policy permits it. That suite must use a dedicated disposable bucket/prefix and explicit cleanup. + +## Primary specification sources + +The implementation and this guide should be checked against these primary AWS sources when S3 behavior changes: + +- AWS Signature Version 4 canonical request: + https://docs.aws.amazon.com/AmazonS3/latest/API/sig-v4-header-based-auth.html +- Multipart upload limits: https://docs.aws.amazon.com/AmazonS3/latest/userguide/qfacts.html +- CompleteMultipartUpload: https://docs.aws.amazon.com/AmazonS3/latest/API/API_CompleteMultipartUpload.html +- CopyObject: https://docs.aws.amazon.com/AmazonS3/latest/API/API_CopyObject.html +- UploadPartCopy: https://docs.aws.amazon.com/AmazonS3/latest/API/API_UploadPartCopy.html +- ListObjectsV2: https://docs.aws.amazon.com/AmazonS3/latest/API/API_ListObjectsV2.html + +Secondary S3-compatible provider documentation can explain provider-specific configuration, but it must not override the +Amazon S3 wire contract when the client claims Amazon S3 behavior. diff --git a/docs/sources.md b/docs/sources.md index 528baa6..721e232 100644 --- a/docs/sources.md +++ b/docs/sources.md @@ -1,10 +1,10 @@ -Research and source register -============================ +# Research and source register Research date: 2026-08-15. -This register records the external contracts used to design and test the implementation. Source code and provider behavior can -change, so a release review should recheck current primary sources rather than assuming this date remains current. +This register records the external contracts used to design and test the implementation. Source code and provider +behavior can change, so a release review should recheck current primary sources rather than assuming this date remains +current. Use this authority order when sources disagree: @@ -13,11 +13,10 @@ Use this authority order when sources disagree: 3. current package implementation and tests; 4. older experiments and secondary performance reports. -The package intentionally distinguishes implemented behavior from provider claims and proposals. A compatibility note in this -file is not evidence that an adapter passed a live integration test against that provider. +The package intentionally distinguishes implemented behavior from provider claims and proposals. A compatibility note in +this file is not evidence that an adapter passed a live integration test against that provider. -Browser File System and OPFS ----------------------------- +## Browser File System and OPFS Primary sources: @@ -46,8 +45,7 @@ Secondary OPFS/performance context reviewed earlier in the project: These sources informed performance questions. They do not override the File System Standard or real browser tests. -Playwright ----------- +## Playwright Primary documentation: @@ -63,8 +61,7 @@ The canonical browser test architecture uses Playwright Test for Chromium, Firef ServiceWorker instrumentation because Playwright documents that inspection surface as Chromium-specific; observable ServiceWorker behavior remains a black-box test in the other browsers. -Testcontainers --------------- +## Testcontainers Primary documentation and source reviewed: @@ -76,18 +73,17 @@ Primary documentation and source reviewed: - Toxiproxy module: - repository: -The provider suite keeps `node:test` as the test runner and uses Testcontainers only for service lifecycle. The official Azurite -module is used instead of reproducing its container flags. SeaweedFS has no focused Testcontainers module, so the fixture uses -`GenericContainer` with the pinned image and a composite listening-port/HTTP wait. Random mapped host ports avoid collisions -between concurrent local runs. +The provider suite keeps `node:test` as the test runner and uses Testcontainers only for service lifecycle. The official +Azurite module is used instead of reproducing its container flags. SeaweedFS has no focused Testcontainers module, so +the fixture uses `GenericContainer` with the pinned image and a composite listening-port/HTTP wait. Random mapped host +ports avoid collisions between concurrent local runs. -Testcontainers currently documents Docker directly and Docker-compatible configuration for Podman, Colima, and Rancher Desktop. -Those runtimes have provider-specific caveats, including Ryuk behavior under Podman and delayed port forwarding under -Colima/Rancher. The package therefore treats Testcontainers as an interim test-resource implementation, not as the future -compute-provider API for OPFS. +Testcontainers currently documents Docker directly and Docker-compatible configuration for Podman, Colima, and Rancher +Desktop. Those runtimes have provider-specific caveats, including Ryuk behavior under Podman and delayed port forwarding +under Colima/Rancher. The package therefore treats Testcontainers as an interim test-resource implementation, not as the +future compute-provider API for OPFS. -Deno, Node, and Bun -------------------- +## Deno, Node, and Bun Primary runtime sources: @@ -99,111 +95,92 @@ Primary runtime sources: - Bun benchmarking guidance: - Node documentation: -Current Deno documentation treats `node:test` as a first-class test API and currently marks Deno KV unstable. The real Deno KV -suite therefore uses `--unstable-kv` without making that flag part of unrelated source imports. Current Deno KV documentation -states a 2 KiB serialized key limit, a 64 KiB serialized value limit, 1,000 mutations per atomic operation, and an 800 KiB total -atomic-operation limit. The Deno KV adapter exposes these constraints and uses a configurable manifest/part layout instead of -pretending one logical file must fit in one 64 KiB value. - -Current Bun compatibility documentation says its in-process `node:test` API works when files run under `bun test`, while some -advanced Node test-runner/reporting features remain incomplete. The repository uses the common `describe`/`it`/hooks subset and -keeps the test API itself as `node:test`. - -Bun's current File I/O documentation says `Bun.write(destination, Bun.file(source))` selects fast platform system calls for -file-to-file copies. The current Bun Rust source also keeps file-backed Blob state distinct so file-to-file paths can avoid a -naive user-space read/write loop. The benchmark therefore compares Bun's direct copy shape with Node-compatible `copyFile` -before changing the adapter implementation. - -Bun's S3 documentation exposes `S3Client`, `S3File`, `write`, `stat`, `stream`, and multipart `writer()` APIs. A Bun-only provider -benchmark now uses that native implementation as a second S3 baseline beside AWS SDK v3. The project does not treat Bun main -branch implementation work as proof about a released runtime; the mise pin remains the current released version selected by the -repository until a deliberate toolchain update. - -Deno standard libraries and Standard Schema -------------------------------------------- - -Primary sources reviewed from the current `denoland/std` repository and JSR -packages: - - - `@std/async`: - - `@std/bytes`: - - `@std/encoding`: - - `@std/expect`: - - `@std/fs`: - - `@std/http`: - - `@std/path`: - - `@std/streams`: - - `@std/xml`: - - `@std/crypto`: - - Standard Schema: - - Zod 4: - -The review was operation-led. A standard package replaces project code only -when its contract matches the filesystem or provider requirement without hiding -a stronger invariant. - -`@std/async/pool` owns bounded multipart and block concurrency. The stable -`pooledMap()` contract limits active requests and lets already-started requests -settle after one item fails. S3 cleanup waits for that settlement before it -sends `AbortMultipartUpload`, so a late part cannot arrive after the cleanup -request. - -`@std/async/retry` owns the direct clients' exponential backoff, jitter, AbortSignal, and retriable-error loop. The protocol -layer still classifies whether a request may enter that loop. One-shot streams are not replayed, and S3 multipart initiation and -completion disable automatic retry because a lost response can make the remote lifecycle outcome ambiguous. The low-level -request APIs also expose `retry: false` for provider-specific operations. - -`@std/bytes/concat` owns byte-array concatenation used by bounded chunk -assembly. The package does not maintain another concatenation implementation. - -`@std/streams` owns bounded materialization through -`LimitedBytesTransformStream` and final stream collection through `toBytes()`. -The current `FixedChunkStream` API is still marked unstable, so fixed-size -provider chunks remain in the package's small streaming adapter until that -standard contract is suitable for a public dependency. - -`@std/encoding` owns Base64 and hexadecimal encoding through their direct -subpaths. S3 uses hexadecimal SHA-256 output, Azure Shared Key and block IDs -use Base64, and record stores use Base64 for portable byte persistence. - -`@std/path` owns host path normalization and resolution for the Deno, Bun, and -Node adapters. The OPFS virtual path model remains project-owned because it -rejects and normalizes a different namespace than an operating-system path. - -`@std/fs` was reviewed for copy, move, walk, ensure, and host filesystem -operations. Those are intentionally not used inside the primitive Deno/Bun/Node -adapters. The public filesystem facade already owns recursive copy/move/walk, -overwrite, cancellation, and adapter-neutral semantics. Calling `@std/fs` from -one host adapter would duplicate that layer and introduce host-only symlink and -filesystem assumptions. `@std/path`, by contrast, directly replaces custom host -path manipulation without changing facade semantics. - -`@std/http/etag` was reviewed for conditional request support. The clients keep -provider ETags opaque instead of generating or evaluating them locally. S3 -multipart ETags and Azure ETags are provider tokens, not hashes that this -library should reinterpret. The package therefore forwards `If-Match` and -`If-None-Match` values to the provider rather than applying `@std/http/etag` in -the client. The unstable HTTP message-signature utilities also do not implement -AWS Signature Version 4 or Azure Shared Key. - -`@std/xml` owns provider control-document parsing and serialization. S3 list, -error, multipart, and copy responses and Azure list/error/block-list documents -use the standard XML tree instead of regular expressions or hand-written XML +Current Deno documentation treats `node:test` as a first-class test API and currently marks Deno KV unstable. The real +Deno KV suite therefore uses `--unstable-kv` without making that flag part of unrelated source imports. Current Deno KV +documentation states a 2 KiB serialized key limit, a 64 KiB serialized value limit, 1,000 mutations per atomic +operation, and an 800 KiB total atomic-operation limit. The Deno KV adapter exposes these constraints and uses a +configurable manifest/part layout instead of pretending one logical file must fit in one 64 KiB value. + +Current Bun compatibility documentation says its in-process `node:test` API works when files run under `bun test`, while +some advanced Node test-runner/reporting features remain incomplete. The repository uses the common +`describe`/`it`/hooks subset and keeps the test API itself as `node:test`. + +Bun's current File I/O documentation says `Bun.write(destination, Bun.file(source))` selects fast platform system calls +for file-to-file copies. The current Bun Rust source also keeps file-backed Blob state distinct so file-to-file paths +can avoid a naive user-space read/write loop. The benchmark therefore compares Bun's direct copy shape with +Node-compatible `copyFile` before changing the adapter implementation. + +Bun's S3 documentation exposes `S3Client`, `S3File`, `write`, `stat`, `stream`, and multipart `writer()` APIs. A +Bun-only provider benchmark now uses that native implementation as a second S3 baseline beside AWS SDK v3. The project +does not treat Bun main branch implementation work as proof about a released runtime; the mise pin remains the current +released version selected by the repository until a deliberate toolchain update. + +## Deno standard libraries and Standard Schema + +Primary sources reviewed from the current `denoland/std` repository and JSR packages: + +- `@std/async`: +- `@std/bytes`: +- `@std/encoding`: +- `@std/expect`: +- `@std/fs`: +- `@std/http`: +- `@std/path`: +- `@std/streams`: +- `@std/xml`: +- `@std/crypto`: +- Standard Schema: +- Zod 4: + +The review was operation-led. A standard package replaces project code only when its contract matches the filesystem or +provider requirement without hiding a stronger invariant. + +`@std/async/pool` owns bounded multipart and block concurrency. The stable `pooledMap()` contract limits active requests +and lets already-started requests settle after one item fails. S3 cleanup waits for that settlement before it sends +`AbortMultipartUpload`, so a late part cannot arrive after the cleanup request. + +`@std/async/retry` owns the direct clients' exponential backoff, jitter, AbortSignal, and retriable-error loop. The +protocol layer still classifies whether a request may enter that loop. One-shot streams are not replayed, and S3 +multipart initiation and completion disable automatic retry because a lost response can make the remote lifecycle +outcome ambiguous. The low-level request APIs also expose `retry: false` for provider-specific operations. + +`@std/bytes/concat` owns byte-array concatenation used by bounded chunk assembly. The package does not maintain another +concatenation implementation. + +`@std/streams` owns bounded materialization through `LimitedBytesTransformStream` and final stream collection through +`toBytes()`. The current `FixedChunkStream` API is still marked unstable, so fixed-size provider chunks remain in the +package's small streaming adapter until that standard contract is suitable for a public dependency. + +`@std/encoding` owns Base64 and hexadecimal encoding through their direct subpaths. S3 uses hexadecimal SHA-256 output, +Azure Shared Key and block IDs use Base64, and record stores use Base64 for portable byte persistence. + +`@std/path` owns host path normalization and resolution for the Deno, Bun, and Node adapters. The OPFS virtual path +model remains project-owned because it rejects and normalizes a different namespace than an operating-system path. + +`@std/fs` was reviewed for copy, move, walk, ensure, and host filesystem operations. Those are intentionally not used +inside the primitive Deno/Bun/Node adapters. The public filesystem facade already owns recursive copy/move/walk, +overwrite, cancellation, and adapter-neutral semantics. Calling `@std/fs` from one host adapter would duplicate that +layer and introduce host-only symlink and filesystem assumptions. `@std/path`, by contrast, directly replaces custom +host path manipulation without changing facade semantics. + +`@std/http/etag` was reviewed for conditional request support. The clients keep provider ETags opaque instead of +generating or evaluating them locally. S3 multipart ETags and Azure ETags are provider tokens, not hashes that this +library should reinterpret. The package therefore forwards `If-Match` and `If-None-Match` values to the provider rather +than applying `@std/http/etag` in the client. The unstable HTTP message-signature utilities also do not implement AWS +Signature Version 4 or Azure Shared Key. + +`@std/xml` owns provider control-document parsing and serialization. S3 list, error, multipart, and copy responses and +Azure list/error/block-list documents use the standard XML tree instead of regular expressions or hand-written XML escaping. Storage payloads themselves do not pass through XML parsing. -`@std/crypto` was reviewed but is not used for provider signing. Web Crypto -already exposes browser-compatible SHA-256 and HMAC-SHA256, while the standard -crypto package does not implement AWS Signature Version 4 or Azure Shared Key -canonicalization. Adding it would introduce a wrapper without removing the -protocol code that actually carries the risk. +`@std/crypto` was reviewed but is not used for provider signing. Web Crypto already exposes browser-compatible SHA-256 +and HMAC-SHA256, while the standard crypto package does not implement AWS Signature Version 4 or Azure Shared Key +canonicalization. Adding it would introduce a wrapper without removing the protocol code that actually carries the risk. -`@std/expect` remains the assertion API on top of `node:test`. Zod 4 implements -Standard Schema, so the repository exports the Zod schemas directly instead of -maintaining a second validation wrapper for Standard Schema consumers. +`@std/expect` remains the assertion API on top of `node:test`. Zod 4 implements Standard Schema, so the repository +exports the Zod schemas directly instead of maintaining a second validation wrapper for Standard Schema consumers. - -S3 and Signature Version 4 --------------------------- +## S3 and Signature Version 4 AWS primary references: @@ -222,7 +199,8 @@ AWS primary references: Implementation details derived from these contracts include: -- canonical signing includes `host` even though browser Fetch does not let application code set the Host header directly; +- canonical signing includes `host` even though browser Fetch does not let application code set the Host header + directly; - multipart upload parts are bounded and the destination publishes on CompleteMultipartUpload; - conditional `If-Match`/`If-None-Match` behavior belongs to multipart completion for the commit path used here; - CompleteMultipartUpload can return an HTTP 200 response whose XML body later reports an error; @@ -230,11 +208,10 @@ Implementation details derived from these contracts include: - CopyObject has a 5 GB source limit, so larger provider-side copies use UploadPartCopy; - multipart uploads permit at most 10,000 parts and have defined part-size limits. -The package uses Web Crypto and Web Fetch rather than the AWS SDK so the direct client remains small, runtime-neutral, and -explicit about the S3 protocol surface it actually implements. +The package uses Web Crypto and Web Fetch rather than the AWS SDK so the direct client remains small, runtime-neutral, +and explicit about the S3 protocol surface it actually implements. -S3-compatible providers ------------------------ +## S3-compatible providers Provider-specific primary sources reviewed for compatibility differences: @@ -258,11 +235,11 @@ Backblaze B2 S3-compatible API: - S3-compatible API: -These providers illustrate why capability overrides exist. Endpoint, region, addressing, copy support, multipart preconditions, -checksum behavior, and unsupported control-plane operations can differ even when basic object requests use the S3 protocol. +These providers illustrate why capability overrides exist. Endpoint, region, addressing, copy support, multipart +preconditions, checksum behavior, and unsupported control-plane operations can differ even when basic object requests +use the S3 protocol. -Azure Blob Storage ------------------- +## Azure Blob Storage Microsoft primary references: @@ -274,16 +251,16 @@ Microsoft primary references: - Copy Blob From URL: - Put Block From URL: - List Blobs: -- Versioning for Azure Storage services: +- Versioning for Azure Storage services: + -The implementation keeps the service version explicit because accepted block sizes and Shared Key canonicalization depend on -the service version. Shared Key support starts at the augmented Blob format introduced in `2009-09-19`; zero-length -`Content-Length` signing changes after `2014-02-14`, and empty `x-ms-*` header canonicalization changes at `2016-05-31`. -Current copy behavior uses synchronous Copy Blob From URL for the smaller path and Put Block From URL ranges for large -provider-side copies. +The implementation keeps the service version explicit because accepted block sizes and Shared Key canonicalization +depend on the service version. Shared Key support starts at the augmented Blob format introduced in `2009-09-19`; +zero-length `Content-Length` signing changes after `2014-02-14`, and empty `x-ms-*` header canonicalization changes at +`2016-05-31`. Current copy behavior uses synchronous Copy Blob From URL for the smaller path and Put Block From URL +ranges for large provider-side copies. -Unstorage ---------- +## Unstorage Primary sources: @@ -293,12 +270,11 @@ Primary sources: - custom drivers: - built-in driver catalog: -The forward bridge targets `Storage`, not individual unstorage drivers. The reverse driver implements the stable Driver subset -needed for values, raw bytes, metadata, keys, clear, and disposal. `maxDepth` is advertised because the reverse driver applies -the depth filter itself. +The forward driver targets unstorage `Storage`, not individual unstorage backend implementations. The reverse bridge +implements the stable unstorage `Driver` subset needed for values, raw bytes, metadata, keys, clear, and disposal. +`maxDepth` is advertised because the bridge applies the depth filter itself. -RxDB ----- +## RxDB Primary sources: @@ -306,11 +282,10 @@ Primary sources: - RxStorage interface: - RxCollection implementation: -The bridge accepts an RxCollection. RxDB retains responsibility for the selected RxStorage, replication, conflicts, +The driver accepts an RxCollection. RxDB retains responsibility for the selected RxStorage, replication, conflicts, multi-instance behavior, wrappers, and licensing. -db0 and Drizzle ---------------- +## db0 and Drizzle Primary sources: @@ -319,35 +294,38 @@ Primary sources: - Drizzle ORM: - Drizzle repository: -The db0 bridge targets the Database/dialect contract rather than connector names. Direct SQLite reuses that same record schema. -Drizzle keeps table/DDL ownership with the application because its schema builders and database behavior are dialect-specific. +The db0 record driver targets the Database/dialect contract rather than connector names. Direct SQLite reuses that same +record schema. The Drizzle record driver keeps table/DDL ownership with the application because its schema builders and +database behavior are dialect-specific. -Upstream issue and pull-request review --------------------------------------- +## Upstream issue and pull-request review -Current upstream issue/PR review was used to find failure modes that happy-path API docs do not reveal. The implementation does -not copy another library's behavior blindly; the issues are evidence for tests and invariants. +Current upstream issue/PR review was used to find failure modes that happy-path API docs do not reveal. The +implementation does not copy another library's behavior blindly; the issues are evidence for tests and invariants. -Bun S3/Rust work reviewed included fixes for retry coverage, exponential backoff, timeouts, manual redirect handling, option -propagation, multipart abort on writer error, long SigV4 inputs, in-place multipart part assembly, XML parsing, proxy handling, -and worker-termination lifetime safety. The repeated lessons are: signed redirects must not be followed automatically, remote -cleanup has its own lifecycle, part concurrency needs a memory budget, and retry policy must not be inferred from body type alone. +Bun S3/Rust work reviewed included fixes for retry coverage, exponential backoff, timeouts, manual redirect handling, +option propagation, multipart abort on writer error, long SigV4 inputs, in-place multipart part assembly, XML parsing, +proxy handling, and worker-termination lifetime safety. The repeated lessons are: signed redirects must not be followed +automatically, remote cleanup has its own lifecycle, part concurrency needs a memory budget, and retry policy must not +be inferred from body type alone. -AWS SDK v3 issues reviewed included very large upload memory growth, unknown-size multipart completion hangs, empty-stream lockups, -stream chunk-integrity regressions, conditional-header gaps in `lib-storage`, browser decompression/checksum mismatches, socket -exhaustion, and S3-compatible provider deserialization/endpoint regressions. The project benchmark keeps the AWS SDK as a -baseline while retaining a smaller direct protocol client with independently testable semantics. +AWS SDK v3 issues reviewed included very large upload memory growth, unknown-size multipart completion hangs, +empty-stream lockups, stream chunk-integrity regressions, conditional-header gaps in `lib-storage`, browser +decompression/checksum mismatches, socket exhaustion, and S3-compatible provider deserialization/endpoint regressions. +The project benchmark keeps the AWS SDK as a baseline while retaining a smaller direct protocol client with +independently testable semantics. -Azure SDK issues reviewed included paused-stream abort hangs, invalid upload buffer arguments producing zero-byte blobs, large -buffer/block-size constraints, historical stream/file data corruption, copy polling request noise, and concurrency/default-size -questions. These reinforce explicit size/concurrency limits, bounded block admission, real abort tests, and provider request-count -benchmarks. +Azure SDK issues reviewed included paused-stream abort hangs, invalid upload buffer arguments producing zero-byte blobs, +large buffer/block-size constraints, historical stream/file data corruption, copy polling request noise, and +concurrency/default-size questions. These reinforce explicit size/concurrency limits, bounded block admission, real +abort tests, and provider request-count benchmarks. -Unstorage issues reviewed included non-atomic filesystem writes, S3 pagination/prefix bugs, XML entity decoding, file/prefix -collisions, SQL disposal, binary Redis storage, and Cloudflare Cache method binding. RxDB issues reviewed included OPFS/Expo file -truncation after crashes or rapid writes, large-replication corruption, and concurrency/benchmark questions. db0 issues reviewed -included connector/dialect exposure, caller-owned connections, deprecated sqlite3, and Drizzle result-shape mismatches. These are -why the OPFS project keeps ownership, collision semantics, partial-result failure, and backend capability differences explicit. +Unstorage issues reviewed included non-atomic filesystem writes, S3 pagination/prefix bugs, XML entity decoding, +file/prefix collisions, SQL disposal, binary Redis storage, and Cloudflare Cache method binding. RxDB issues reviewed +included OPFS/Expo file truncation after crashes or rapid writes, large-replication corruption, and +concurrency/benchmark questions. db0 issues reviewed included connector/dialect exposure, caller-owned connections, +deprecated sqlite3, and Drizzle result-shape mismatches. These are why the OPFS project keeps ownership, collision +semantics, partial-result failure, and backend capability differences explicit. Deno KV issue review also covered historical reports about large prefix-list cost and selector/transaction limits: @@ -355,10 +333,11 @@ Deno KV issue review also covered historical reports about large prefix-list cos - - -The Deno KV physical key layout therefore indexes a logical entry by `(namespace, "entry", parentPath, name)`. Listing one -directory uses `(namespace, "entry", parentPath)` as the provider prefix, so descendants of a child directory are not part of -that prefix result. Physical body parts use the complete canonical path as one tuple component rather than expanding each path -segment into the provider prefix. This keeps exact lookup and direct-child enumeration aligned with the filesystem contract. +The Deno KV physical key layout therefore indexes a logical entry by `(namespace, "entry", parentPath, name)`. Listing +one directory uses `(namespace, "entry", parentPath)` as the provider prefix, so descendants of a child directory are +not part of that prefix result. Physical body parts use the complete canonical path as one tuple component rather than +expanding each path segment into the provider prefix. This keeps exact lookup and direct-child enumeration aligned with +the filesystem contract. Recent Drizzle issue review included SQLite/libSQL transaction-lifetime failures and migration/data-loss cases: @@ -368,12 +347,12 @@ Recent Drizzle issue review included SQLite/libSQL transaction-lifetime failures - - -These are not all adapter-runtime bugs, but they reinforce a deliberate contract here: the generic Drizzle bridge does not -claim universal cross-process atomic replacement or own application migrations. The caller keeps dialect/driver/table lifecycle -and can provide a stronger database-specific transaction strategy when that concrete driver proves the required semantics. +These are not all driver-runtime bugs, but they reinforce a deliberate contract here: the generic Drizzle record driver +does not claim universal cross-process atomic replacement or own application migrations. The caller keeps +dialect/driver/table lifecycle and can provide a stronger database-specific transaction strategy when that concrete +driver proves the required semantics. -Project architecture and writing sources ----------------------------------------- +## Project architecture and writing sources The implementation was reviewed against the attached/current project guides covering: @@ -386,13 +365,49 @@ The implementation was reviewed against the attached/current project guides cove - runtime-neutral TypeScript and explicit runtime subpaths; - verification against real runtimes and extracted release artifacts. -Older OPFS experiments were treated as intent/history only. The current repository, current project rules, and current upstream -contracts are the implementation authority for this pass. +Older OPFS experiments were treated as intent/history only. The current repository, current project rules, and current +upstream contracts are the implementation authority for this pass. + +## Client protocol handoffs + +The detailed implementation contracts live in [s3.md](./s3.md) and [azure.md](./azure.md). The Testcontainers-backed +interoperability matrix is documented in [providers.md](./providers.md). These files separate protocol requirements from +emulator evidence and record the unsupported surface explicitly. + +## Filesystem-client baselines + +AWS Mountpoint primary sources reviewed on 2026-08-15: + +- Amazon S3 Mountpoint overview: +- Mountpoint usage: +- Mountpoint configuration source: + +Mountpoint is a high-throughput S3 filesystem client, not a full POSIX filesystem. AWS documents that it can list/read +existing objects and create new files, while operations such as modifying existing files, symbolic links, and file +locking are not general Mountpoint capabilities. The benchmark therefore compares only operations whose semantics are +close enough to answer the same question. `--endpoint-url` and `--force-path-style` can support controlled endpoint +experiments, but an alternate endpoint is not an AWS compatibility guarantee. + +Azure BlobFuse primary sources reviewed on 2026-08-15: + +- BlobFuse repository: +- Microsoft limitations/known issues: + +BlobFuse translates Linux FUSE operations to Azure Blob requests. Its cache modes and unsupported/altered filesystem +operations are part of benchmark semantics. In particular, cache configuration can change freshness and request count, +while operations such as file locking and several extended/POSIX operations are not supported. Benchmark output must +therefore record cache/write mode rather than treating every mounted file operation as equivalent to direct REST access. + +## Mise and registry publishing +Primary tool/release sources reviewed on 2026-08-15: -Client protocol handoffs ------------------------- +- mise npm backend: +- mise settings: +- npm trusted publishing: +- npm provenance: -The detailed implementation contracts live in [s3.md](./s3.md) and [azure.md](./azure.md). The Testcontainers-backed interoperability -matrix is documented in [providers.md](./providers.md). These files separate protocol requirements from emulator evidence and -record the unsupported surface explicitly. +Mise can install npm-distributed CLI tools directly through the `npm:` backend. The release configuration therefore pins +the npm CLI through mise instead of depending on the npm version bundled with a selected Node release. npm's current +trusted-publishing requirements call for npm CLI 11.5.1 or later and Node 22.14.0 or later; the repository pins npm +11.18.0 for the publish job. diff --git a/docs/validation.md b/docs/validation.md index 52aa57c..564d061 100644 --- a/docs/validation.md +++ b/docs/validation.md @@ -1,330 +1,419 @@ -Validation strategy -=================== +# Validation strategy -The test architecture separates portable filesystem semantics from the runtimes and providers that supply concrete storage. -This is deliberate. A fast memory test should not be the evidence for browser OPFS interoperability, and a browser test should -not be the only evidence for a deterministic path or copy invariant. +## Purpose -The canonical layers are: +Validation follows the storage layers. A memory test is not evidence for browser OPFS interoperability. A fake HTTP test +is not evidence that an S3-compatible server accepts the request. A facade benchmark is not enough to identify whether +overhead came from the protocol client, driver, adapter, metrics, or provider. + +The canonical model is: ```text node:test + @std/expect - portable schemas, paths, facade behavior, record/object translations, - ecosystem bridges, S3/Azure protocol behavior with deterministic fakes - -real server runtimes - Deno host filesystem - Deno KV - Node host filesystem + node:sqlite - Bun host filesystem + schemas / paths / driver contracts / adapter translation / facade semantics + deterministic S3/Azure protocol tests + record/object/database contract tests -Testcontainers + node:test - SeaweedFS S3 compatibility / Azurite Blob interoperability - random host ports / readiness / owned cleanup +Deno / Node / Bun runtime tests + actual host filesystem behavior + Deno KV runtime behavior Playwright Test Chromium / Firefox / WebKit - Window / Worker / ServiceWorker / iframe / persistence / browser storage + Window / Worker / iframe / ServiceWorker / persistence -Mitata - raw backend baseline -> adapter primitive -> filesystem facade +Testcontainers + node:test + disposable SeaweedFS and Azurite provider interoperability + +Mitata + Playwright benchmarks + native/client -> driver -> adapter -> facade -> facade+metrics ``` -A test states the contract it protects. Avoid tests that only mirror the current implementation line by line. - -Portable tests protect filesystem semantics -------------------------------------------- - -The portable suite uses `node:test` with `describe` and `it`, plus `@std/expect` for expectations. Deno runs these same source -files directly. Node runs the same source. Bun runs the same `node:test` API through `bun test`. - -The suite covers: - -- schema acceptance/rejection and Standard Schema exposure; -- canonical path normalization and root-escape rejection; -- file and directory handle semantics; -- replace, append, update, truncate, and byte-range behavior; -- staged writable close versus abort; -- stream cancellation after the operation becomes terminal; -- bounded stream materialization for simple record adapters; -- Deno KV partitioned large-file stat/list/range/stream behavior, bounded append/update patching, and manifest-last visibility; -- optimization-disabled differential paths and matching preflight plans; -- filesystem route/peak-buffer metrics; -- copy/move overwrite and source/destination overlap protection; -- file mutation versus structural mutation coordination; -- queued cancellation recovery; -- sync-file lock lifetime and partial-write looping; -- adapter disposal ownership; -- record-store semantics; -- generic object-store directories, ranges, streaming replacement, optimistic read-modify-write, and native copy; -- foreign object layouts where an exact file key and a child prefix coexist; -- unstorage forward and reverse integration; -- the generic reverse key-value driver and collision-safe keys; -- RxDB, db0 dialect, Drizzle, and direct SQLite translation; -- S3 Signature Version 4, XML list/error parsing, multipart commit preconditions, HTTP-200 embedded failures, multipart - server-side copy, retry/backoff, timeout, manual redirects, one-shot body admission, and non-idempotent multipart lifecycle retry guards; -- Azure list/error parsing, large server-side range copy, bearer/SAS source authorization, provider request identities, - retry/backoff, and one-shot/explicit no-retry behavior. - -Focused commands: +## Repository command authority + +Mise owns tool versions and repository commands. ```sh -deno task test:portable -deno task test:node -deno task test:bun +mise install +mise run check +mise run test +mise run test-node +mise run test-deno +mise run test-bun +mise run test-browser +mise run test-providers +mise run bench +mise run bench-providers +mise run bench-browser +mise run bench-filesystem-clients +mise run quality ``` -The deterministic stress run shuffles and repeats the portable suite so hidden test order does not become a dependency: +GitHub Actions owns only GitHub-specific orchestration: triggers, permissions, matrices, secrets, outputs, immutable +release refs, and calls into those mise tasks. -```sh -deno task test:stress -``` +## Portable tests -Runtime suites prove runtime adapters against the real API ----------------------------------------------------------- +Portable tests use `node:test` with `describe`/`it` and `@std/expect`. Deno and Node consume the same TypeScript source. +Bun uses the same test contracts where its runner/runtime supports them. -`tests/deno.test.ts` exercises the real Deno host filesystem adapter. `tests/node.test.ts` exercises real Node host filesystem -operations and runs the SQL bridge against Node's built-in SQLite engine. `tests/bun.test.ts` exercises the Bun adapter against -Bun's actual runtime. +Important portable suites: -Deno KV has a separate real integration test because current Deno requires the unstable KV flag: +```text +tests/path.test.ts + canonical path parsing / root escape / names -```sh -deno task test:deno-kv -``` +tests/driver.test.ts + driver definition validation + requirement/limit provenance + behavior-changing optimization disableability + direct third-party driver planning -The adapter module is still type-checked with the server set. The unstable flag belongs to the real Deno KV execution, not to -unrelated package imports. +tests/memory.test.ts + deterministic record driver/adapter/facade behavior -The normal server-runtime matrix is: +tests/filesystem.test.ts + locks / staged writes / copy / move / cancellation / lifecycle -```sh -deno task test:deno -deno task test:deno-kv -deno task test:node -deno task test:bun -``` +tests/ecosystems.test.ts + unstorage / RxDB / db0 / Drizzle + integration direction metadata + real reverse unstorage bridge -The pinned mise task runs these after installing the frozen dependency graph: +tests/object.test.ts + generic object driver -> object adapter -> facade contract -```sh -mise run test +tests/s3.test.ts + deterministic S3 REST/SigV4/multipart/copy/retry behavior + +tests/azure.test.ts + deterministic Azure REST/auth/block/copy/retry behavior + +tests/deno-kv-partition.test.ts + partition layout using an in-memory Deno KV contract double + +tests/sqlite.test.ts + direct SQLite row-driver behavior ``` -GitHub Actions does not recreate this runtime setup with separate Node, Deno, and Bun setup actions. `jdx/mise-action` installs -the pinned mise release and only the tools required by the current job. The job then calls the same focused mise task a -maintainer can run locally, such as `mise run test-deno`, `mise run test-node`, or `mise run test-bun`. This keeps tool versions -and test commands in the repository instead of duplicating them in workflow YAML. +A test should identify the contract it protects. Avoid tests that merely restate private implementation steps. -Testcontainers owns provider-service lifecycle ---------------------------------------------- +## Driver tests -`tests/provider.test.ts` remains a `node:test` suite. `tests/provider/fixture.ts` uses Testcontainers only to supply real local -services. SeaweedFS runs through `GenericContainer`; Azurite uses the official `@testcontainers/azurite` module. Testcontainers -selects free host ports, applies readiness checks, and owns container cleanup. No Docker Compose subprocess or project polling -loop is required. +A driver is a public extension seam and therefore receives direct tests before an adapter exists. -```sh -mise run test-providers -``` +The generic suite proves: -The provider job in GitHub Actions installs Deno and Node through mise and calls that same task. Docker-compatible runtime -selection remains Testcontainers configuration, not package runtime logic. This is an interim fixture layer and does not constrain -a future provider abstraction to the Docker API. +- structured requirements are retained; +- provider, implementation, user, and probe limits keep their provenance; +- an optimization with `changesBehavior: true` cannot declare `disableable: false`; +- driver planning can reject a known impossible input without storage I/O; +- structured problems/actions are stable machine data. -Playwright owns browser installation and browser lifecycle ---------------------------------------------------------- +Backend-specific driver tests then protect physical rules. Deno KV is the most important reference because byte size, +path/key size, partition policy, and provider ceilings all affect admission. -The browser suite lives under `tests/browser/`. There is no custom browser-launch loop or custom test-result protocol. -Playwright owns browser installation, contexts, server lifecycle, traces, retries, and test attribution. +## Deno KV validation -Install the compatible browser builds, then run the matrix: +The portable Deno KV partition suite uses a deterministic contract double. It proves: -```sh -deno task test:browser:install -deno task test:browser -``` +- no stored test value crosses the documented serialized value ceiling enforced by the double; +- conservative `partBytes`/`inlineBytes` safety budgets reject unsafe configuration; +- a long physical key is rejected during driver preflight before provider I/O; +- `partition: "never"` returns structured `change-policy`/`select-driver` actions for oversized input; +- large logical files reconstruct exactly; +- directory listing does not load file body parts; +- range reads touch only overlapping parts; +- streamed replacement uses bounded partition writes without facade buffering; +- disabling the facade stream-write optimization forces the bounded facade fallback; +- append/update preserve untouched bytes; +- manifest-last replacement never publishes a partial new generation. -or: +The Deno-native suite uses the real Deno KV API when the runtime is available. It remains the release evidence for +actual serialization/provider behavior; the contract double does not replace it. -```sh -mise run test-browser -``` +## Host filesystem runtime tests -The same semantic tests run in Chromium, Firefox, and WebKit. Tests probe runtime capability and then assert the actual result. -They do not encode statements such as "Firefox has no sync OPFS" or "WebKit always rejects this iframe" into the test logic. -Those are exactly the assumptions an interoperability suite is supposed to detect when browser behavior changes. +Node, Deno, and Bun each run real host-file tests because a shared structural contract cannot prove runtime I/O +behavior. -The browser cases include: +The runtime suites cover the applicable routes: ```text -Window - OPFS probe - async write/read - abort before commit - -DedicatedWorker - async OPFS - synchronous handle probe and open attempt +replace / append / update +ranges +native streams +native copy +native move +asynchronous positional files +sync random access +flush +close/disposal +cancellation +``` -SharedWorker - async OPFS through the actual SharedWorker realm +Bun tests also verify the Bun-specific read/replace route rather than only the Node-compatible fallback. -ServiceWorker - black-box registration + postMessage in all browsers - deeper serviceWorkers() instrumentation in Chromium only +## Browser tests use Playwright -iframes - same-origin - cross-origin - opaque sandbox +Playwright owns browser lifecycle and cross-browser orchestration. The matrix covers Chromium, Firefox, and WebKit +instead of encoding a Chromium-only browser assumption. -storage lifecycle - fresh BrowserContext isolation - persistent profile close/reopen +Browser cases include: +```text +Window async OPFS +DedicatedWorker +SharedWorker +ServiceWorker observable behavior +same-origin iframe +cross-origin iframe +opaque sandbox iframe +persistent/reopen behavior browser record adapters - localStorage - IndexedDB - Cache Storage +locking/cancellation where the browser exposes the capability ``` -The iframe and ServiceWorker tests report unsupported runtime APIs as capability skips. A supported realm whose OPFS root is -rejected is not silently skipped; the test asserts that a normalized root failure is present. +Synchronous OPFS access is probed from the actual realm/handle. A test does not infer support from browser name or +worker type. + +Playwright's deeper ServiceWorker instrumentation is browser-specific, so cross-browser service-worker tests use a +page-owned registration/message path when direct runner instrumentation is unavailable. -Benchmarks measure overhead against the direct backend ------------------------------------------------------- +## Provider tests use Testcontainers -A benchmark without a raw baseline cannot tell whether the adapter is fast or merely whether one code path is faster than -another OPFS code path. The benchmark layout therefore keeps three layers visible: +`tests/provider/fixture.ts` owns disposable provider services through Testcontainers. ```text -raw backend - | - v -adapter primitive - | - v +ProviderFixture + | + +-- SeaweedFS S3-compatible endpoint + `-- Azurite Blob endpoint +``` + +Testcontainers selects mapped host ports and owns readiness/disposal. The repository does not keep a parallel Docker +Compose, fixed-port, curl-polling lifecycle. + +Provider tests exercise: + +```text +protocol client + -> provider + +driver + -> client/provider + +adapter + -> driver + FileSystemType - | - +-- coordination: none - `-- coordination: local + -> adapter ``` -`bench/memory.bench.ts` measures raw `Map`, direct `RecordStoreType`, direct memory adapter, and facade overhead. This exposes the -cost of record serialization separately from the higher-level filesystem contract. +The provider suite proves interoperability with SeaweedFS/Azurite. It does not redefine Amazon S3 or Azure Blob +specifications. Deterministic protocol tests continue to protect exact signing, conditions, limits, and error parsing. -`bench/node.bench.ts`, `bench/deno.bench.ts`, and `bench/bun.bench.ts` compare raw host filesystem reads/writes and copy with -the direct adapter and facade. Bun measures both Node-compatible `copyFile` and `Bun.write(destination, Bun.file(source))` so a -future adapter change has a runtime baseline instead of an assumption. `bench/deno-kv.bench.ts` does the same for real local Deno -KV, and `bench/sqlite.bench.ts` compares a raw SQLite BLOB row with the direct record adapter and facade. Metrics are disabled for -facade baseline measurements. +## Stress and lifecycle tests -```sh -deno task bench:memory -deno task bench:deno -deno task bench:deno-kv -deno task bench:node -deno task bench:sqlite -deno task bench:bun +`test:stress` runs the portable suite repeatedly with a fixed shuffle seed. Its purpose is to expose ordering, lock, +cleanup, and state-sharing defects that one deterministic order can hide. + +Lifecycle-sensitive code must test: + +- successful cleanup; +- cleanup after failure; +- caller cancellation; +- post-open stream cancellation; +- close exactly once; +- abort exactly once; +- use after close/abort; +- ownership transfer versus borrowed resources; +- cleanup with a separate signal when the caller signal is already aborted. + +## Coverage + +Coverage is useful evidence, not architectural proof. The coverage task exists to find unexecuted branches in portable +code. A high line percentage does not prove that provider limits, cancellation, or resource ownership are correct. + +## Benchmarks measure each layer + +Every benchmark should identify the cost added by one layer. + +For an object protocol: + +```text +official/native SDK baseline + | +project protocol client + | +project object driver + | +project object adapter + | +FileSystemType metrics:none + | +FileSystemType metrics:basic ``` -`mise run bench` runs the server/memory set with the pinned runtimes. +For a host/native filesystem: -The browser benchmark keeps the raw browser API, direct adapter, and facade visible in each real browser. Native OPFS uses 25 -replace/read iterations with a 64 KiB payload. localStorage, IndexedDB, and Cache Storage use 20 iterations with a 16 KiB -payload. Each sample records the raw, adapter, and facade durations plus the adapter/raw, facade/raw, and facade/adapter ratios -as a Playwright attachment. +```text +raw runtime filesystem API + | +project file driver + | +project file adapter + | +FileSystemType +``` -```sh -deno task bench:browser -# or -mise run bench-browser +For memory/record storage: + +```text +raw Map/value structure + | +record driver + | +record adapter + | +FileSystemType ``` -Microbenchmarks are evidence about overhead in the measured operation. They are not universal provider throughput numbers. -Object-store latency, geographical distance, TLS, provider multipart behavior, and connection reuse can dominate the small -client/facade cost. +The benchmark result should include throughput/latency plus semantic context. A faster route is not a valid substitute +if it has different supported operations, consistency, atomicity, or caching semantics. -`mise run bench-providers` starts the pinned SeaweedFS/Azurite services through Testcontainers and compares official SDK/native -runtime baselines with the direct protocol clients, object adapters, and filesystem facade. `bench/providers.ts` owns provider -startup before it launches benchmark programs, so image pull/readiness time is outside Mitata samples. S3 includes AWS SDK v3 and a Bun-native `S3Client` run; Azure -uses `@azure/storage-blob`. The small write baseline includes the same follow-up stat/properties request as the direct project -client, and multipart/block cases are separate. `metrics: "none"` versus `metrics: "basic"` makes instrumentation overhead -visible instead of hiding it. Real-cloud performance still requires an opt-in controlled provider benchmark. +## Provider benchmarks -Type, lint, format, and documentation gates stay separate ---------------------------------------------------------- +`bench/provider.bench.ts` uses the Testcontainers provider fixture and compares: -`deno task check` type-checks the code in environment-focused groups so unrelated ambient globals do not accidentally make an -invalid target look valid: +S3: ```text -check:core - root/core + provider-neutral adapters/clients + reverse drivers +AWS SDK +project S3 client +project S3 driver +project object adapter +facade metrics:none +facade metrics:basic +``` -check:browser - Window/browser storage adapters + Playwright specs/config +Azure: -check:workers - DedicatedWorker / SharedWorker / ServiceWorker fixtures with WebWorker libs +```text +Azure SDK +project Azure client +project Azure driver +project object adapter +facade metrics:none +facade metrics:basic +``` + +`bench/bun-provider.bench.ts` also compares Bun's native S3 client against the same project layers when Bun is +available. -check:server - Deno / Deno KV / Node / Bun / SQLite adapters and server benchmarks +Provider container startup/readiness happens before measured samples. Container pull/start time is not benchmark data. -check:tests - portable + runtime test source +## Filesystem-client baselines -check:deno-kv - Deno KV test source with the unstable KV flag +`bench/filesystem-provider.bench.ts` compares already-mounted provider filesystem clients through the same local-file +staircase: -check:providers - Testcontainers fixture + provider tests + provider benchmark orchestration +```text +raw mounted path + -> Node file driver + -> file adapter + -> FileSystemType ``` -The normal quality gates are: +Environment variables select mounted roots: ```sh -deno ci -deno task check -deno task lint -deno task doc -deno task fmt:check +OPFS_MOUNTPOINT_S3_ROOT=/mnt/s3 \ +OPFS_BLOBFUSE_ROOT=/mnt/azure \ +mise run bench-filesystem-clients ``` -`deno ci` is important because the committed manifests and lockfile must describe one dependency graph. A changed dependency is -not ready for release until the real lockfile has been regenerated and the frozen install succeeds. +The external mounts are intentionally not started inside the normal Testcontainers fixture. AWS Mountpoint and Azure +BlobFuse are FUSE/system clients with host privileges, installation, mount, and unmount lifecycle beyond a normal +application container. A dedicated benchmark runner can provision them and then call the same mise task. -Release validation checks the artifact, not only the source tree ---------------------------------------------------------------- +Only comparable operations should be measured. Unsupported filesystem operations are capability differences, not +benchmark failures. -Before publication, the repository runs the complete source-level checks plus registry dry-runs. npm packaging uses Deno's -package output and then adjusts Drizzle from a normal generated dependency to the optional peer relationship authored by this -project. +## Browser benchmarks -The artifact gate should verify: +The Playwright benchmark compares raw native OPFS calls against the package's OPFS driver, adapter, and facade where +practical. Each browser result is separate. A result from one browser is not generalized to another engine. -1. the public export map contains every intended subpath and no internal-only file; -2. the generated npm package has JavaScript/declarations that import in Node, Deno, and Bun; -3. browser-safe imports bundle without pulling server-only adapters into the root graph; -4. optional Drizzle remains optional until its subpath is imported; -5. package files exclude tests, benchmarks, coverage, temporary output, and repository-only tooling; -6. the extracted artifact passes the same checks that are meaningful after packaging. +## Metrics cost is measurable -The release command is: +Facade metrics support: -```sh -deno task release:check +```text +none +basic +timing ``` -Browser tests and browser benchmarks remain explicit matrix jobs because downloading three browser engines is a large operation -and should not be hidden inside every local unit-test invocation. +`none` is the instrumentation baseline. `basic` records counters without per-operation timing. `timing` adds monotonic +clock work. Benchmarks keep those modes separate so metrics overhead cannot hide inside the main facade result. + +Driver physical metrics are also distinct from facade metrics. S3/Azure request/retry/part work should not be inferred +from one logical filesystem write. + +## Quality gate + +`mise run quality` owns the Deno-centric release quality gate: + +```text +frozen dependency install +strict check graph +lint +public documentation lint +format check +stress tests +coverage tests +JSR dry-run +npm/deno package dry-run +``` + +`mise run test`, browser tests, provider tests, and runtime matrix jobs add the environment-specific evidence. + +## Agent validation + +A ChatGPT/agent host can lack Deno, Bun, Docker, mise, package registry access, or Playwright browsers. Temporary +validation support belongs under `.agents/` and never becomes production code. + +Allowed fallback rules: + +1. Keep production source Deno/browser/server-native. +2. Use the installed Node.js/TypeScript toolchain for supplemental strict checks. +3. Add narrow `.agents/` declarations/stubs only for dependencies unavailable in the host. +4. Do not change production imports merely to satisfy the agent host. +5. Report missing canonical runtime gates explicitly. + +A validation-only type stub can prove project TypeScript structure. It cannot prove the external dependency's real +runtime or full type contract. Release CI must run against the actual dependency graph. + +## Artifact verification + +Before delivering a modified ZIP: + +1. run every available strict/type/behavior/configuration check on the working tree; +2. inspect stale exports/imports and documentation terminology; +3. inspect package exports and publish payload; +4. remove generated validation/build dependency state; +5. create the ZIP; +6. extract that exact ZIP to a clean directory; +7. recreate only validation-side host declarations if needed; +8. rerun the available checks against the extracted artifact; +9. compare source/extracted file lists; +10. record SHA-256. + +The extracted artifact is the final thing that must pass the claimed checks. A green mutable working tree is not enough. -Testcontainers provider tests ------------------------------ +## Release evidence -`mise run test-providers` runs `tests/provider.test.ts` under Node. The suite opens pinned SeaweedFS and Azurite services through -Testcontainers, uses random mapped host ports, waits for provider readiness, and releases every owned container after the suite. -It proves real HTTP/signing/interoperability for the direct clients without a repository-owned Docker Compose lifecycle. It does -not replace deterministic request-shape tests or real-cloud conformance. See [providers.md](./providers.md) for the exact coverage -and limitations. `mise run bench-providers` starts the same fixture outside the timed benchmark programs. +A release-ready claim requires all applicable canonical gates, including Deno, Node, Bun, Playwright, provider +containers, package dry-runs, and lockfile validation. If the current host cannot run one of those environments, the +result is recorded as unverified rather than passed. -- 2.51.2