From 04c2a3ed7ea485800c580d00c9327432d50a48ea Mon Sep 17 00:00:00 2001 From: Okiki Ojo Date: Mon, 17 Aug 2026 03:52:08 -0400 Subject: [PATCH] docs: add comprehensive RDF and SPARQL repository guide Signed-off-by: Okiki Ojo --- AGENTS.md | 82 +++++ README.md | 920 +++++++++++------------------------------------------- 2 files changed, 261 insertions(+), 741 deletions(-) create mode 100644 AGENTS.md diff --git a/AGENTS.md b/AGENTS.md new file mode 100644 index 0000000..ec5e017 --- /dev/null +++ b/AGENTS.md @@ -0,0 +1,82 @@ +# RDF and SPARQL repository guide + +Read [`docs/architecture.md`](./docs/architecture.md) before architecture or public API changes. + +## Runtime and dependency rules + +- Production code targets Deno v2, strict TypeScript, ESM, explicit `.ts` source imports, and JavaScript-native TypeScript. +- The same source should remain usable from Node.js, Bun, browsers, and Workers where the capability applies. +- Importing a package or definition must not start network work, initialize Wasm, open files, configure logging, or acquire unrelated resources. +- Keep `@okikio/rdf` root lightweight. Optional processors and host-specific implementations belong behind explicit subpaths. +- Do not add a dependency only to save a small amount of code. A dependency must materially improve correctness, standards conformance, performance, maintenance, or interoperability. + +## Naming + +- Prefer one-word names when the package or module supplies enough context. +- Use concrete verbs such as `get`, `create`, `open`, `read`, `write`, `parse`, `inspect`, `select`, `close`, and `cancel`. +- Avoid vague `generate`, `execute`, `handle`, `process`, `manager`, `helper`, `common`, `shared`, and `misc` names unless an external protocol requires them. +- Project-owned runtime schemas end in `Schema`. +- Project-owned data types normally end in `Type`. +- Established RDF/SPARQL protocol nouns such as `Quad`, `Literal`, `Dataset`, and `Queryable` keep their established names. +- Generated vocabulary terms, types, and schemas use direct PascalCase imports when the vocabulary term is PascalCase: `Product`, `ProductType`, `ProductSchema`. +- Namespace imports are preferred for coherent operation families: `import * as rdf` and `import * as sparql`. + +## Parsers + +- Prefer semantic event streams over compulsory AST construction. +- Use data-oriented scanner state on measured hot paths. +- Do not allocate token/AST objects when the normal consumer only needs quads or semantic records. +- Preserve source ranges and diagnostics when tolerant or tooling-oriented parsing needs them. +- Cancellation must stop pending source reads and upstream parser work. +- Buffer enough semantic work to avoid leaking partial statements before a recoverable syntax error. +- Add a limit before accepting any parser path whose memory or work can grow with untrusted input. + +## RDF semantics + +- W3C RDF semantics are authoritative. RDF/JS is an interoperability target. +- Do not reinterpret RDFS domain/range as JSON requiredness. +- Preserve unsupported/unknown ontology or SHACL assertions instead of discarding them. +- Keep draft-sensitive RDF/SPARQL/SHACL behavior versioned and documented. +- Do not claim a standards feature until the relevant conformance corpus has passed. + +## Lifecycles + +- The caller owns injected resources unless an API explicitly transfers ownership. +- Use `AbortSignal` for cooperative cancellation. +- Use `Disposable` / `AsyncDisposable` for owned cleanup where appropriate. +- Returning early from an async iterable must release/cancel upstream work. +- A promise represents terminal authority; event streams are observational. + +## Tests and benchmarks + +- Co-locate tests as `*_test.ts` and benchmarks as `*_bench.ts`. +- Use `node:test` with `@std/expect` for expectations. +- Test success, failure, cancellation, early return, cleanup, limits, malformed inputs, hostile chunk splits, and recovery. +- Benchmark semantically equivalent work. State the oracle in the benchmark source. +- Construct deterministic fixtures outside timed callbacks. +- Separate cold start from warm throughput. +- Measure memory and compiler cost when those can change an architectural decision. +- Keep durable benchmark definitions in the owning package or `bench/`. Assistant-only raw benchmark captures must stay outside the repository; do not promote one run into a universal claim. +- A surprising benchmark result is a profiling lead. Identify the mechanism, change one thing, and rerun the same oracle. + +## Documentation + +- Document public APIs and important internal invariants with TSDoc. +- Comments explain why, ownership, limits, standards rules, and non-obvious invariants. Do not narrate obvious syntax. +- Examples must compile against the current public API. +- Distinguish implemented behavior, blocked validation, and proposals. +- Do not preserve obsolete APIs merely for compatibility before the first stable release. + +## Validation + +Canonical Deno gates: + +```sh +deno task fmt --check +deno task lint +deno task check +deno task test +deno task bench +``` + +When Deno or JSR is unavailable, keep production code Deno-native. Temporary Node validation is assistant-only scratch state and must stay outside the committed repository. `.agents/` is ignored as an additional safeguard. diff --git a/README.md b/README.md index c823170..36b0b52 100644 --- a/README.md +++ b/README.md @@ -1,842 +1,280 @@ -# SPARQL Query Builder +# RDF and SPARQL for TypeScript -Writing SPARQL queries by hand means string concatenation, manual escaping, and hunting through parentheses when something breaks. This library gives you a type-safe query builder with a fluent API. Write `v('age').gte(18)` instead of `FILTER(?age >= 18)`, chain operations like `v('price').mul(1.2).round()`, and get autocomplete in your editor. TypeScript catches errors at compile time, values are properly escaped automatically, and you can compose queries from reusable pieces. +This repository is the home of a small RDF and SPARQL package family for Deno, Node.js, Bun, browsers, and Workers. -## Installation +The packages are library-first. Importing a package describes or creates values. It does not open files, start query engines, initialize WebAssembly, perform network requests, or configure global state. -```bash -deno add @okikio/sparql +```text +@okikio/rdf + │ + ├── @okikio/sparql + │ ├── @okikio/oxigraph + │ └── @okikio/comunica + │ + ├── @okikio/vocab + └── @okikio/triplestore ``` -```ts -import { select, triple, node, v, filter } from '@okikio/sparql' -``` - -## Quick Start - -Start with a simple query using triple patterns: - -```ts -const adults = select(['?name', '?age']) - .where(triple('?person', 'foaf:name', '?name')) - .where(triple('?person', 'foaf:age', '?age')) - .filter(v('age').gte(18)) - -const sparql = adults.build() -const results = await adults.execute({ - endpoint: 'http://localhost:3030/dataset/sparql' -}) -``` - -That `v('age').gte(18)` is the fluent API - variables become values with chainable methods. Compare values, do math, transform strings, all with natural dot notation. +The repository targets current RDF 1.2 semantics while keeping draft-dependent behavior explicit. SPARQL and SHACL 1.2 features are versioned because those specifications are still evolving. -## SPARQL Mapping +## Packages -Every library feature maps directly to standard SPARQL 1.1. The library provides 100% spec coverage with enhanced developer experience through type safety, fluent chaining, and multiple pattern styles. +| Package | Responsibility | Runtime dependency posture | +| --- | --- | --- | +| `@okikio/rdf` | RDF terms, datasets, parsers, ontology and shape models | root module imports no third-party processor; processor subpaths are isolated | +| `@okikio/sparql` | SPARQL construction, lexical inspection, result contracts, HTTP protocol client | core depends only on `@okikio/rdf` | +| `@okikio/vocab` | Ontology-to-TypeScript compiler and generated vocabulary runtime | depends on `@okikio/rdf` | +| `@okikio/triplestore` | Crash-recoverable persistent RDF dataset | depends on `@okikio/rdf`; borrows a structural filesystem | +| `@okikio/oxigraph` | Adapter from a caller-owned Oxigraph `Store` to the SPARQL query contract | Oxigraph remains caller-owned | +| `@okikio/comunica` | Adapter from a caller-owned Comunica `QueryEngine` to the SPARQL query contract | Comunica remains caller-owned | -**Key mappings:** -- `v('age').gte(18)` → `?age >= 18` -- `select([v('price').mul(1.2).as('total')])` → `SELECT (?price * 1.2 AS ?total)` -- `triple('?s', 'rdf:type', 'ex:Person')` → `?s rdf:type ex:Person .` -- `md5(v('email'))` → `MD5(?email)` -- `now()` → `NOW()` - -**See [sparql-mapping.md](./docs/sparql-mapping.md) for:** -- Complete function reference (85+ functions) -- Library → SPARQL examples for all features -- SPARQL → Library migration guide -- DX enhancements beyond the spec - -Call `.build().value` on any query to see the generated SPARQL string. - -## Pattern Styles - -The library supports multiple ways to describe graph patterns. Use triples for simple cases, nested objects for complex structures, or ASCII art when you want visual clarity. Every pattern compiles to standard SPARQL, so choose based on readability. - -Basic triples work for straightforward queries: - -```ts -select(['?name', '?age']) - .where(triple('?person', 'foaf:name', '?name')) - .where(triple('?person', 'foaf:age', '?age')) -``` +## RDF -Nested object notation handles complex graphs without repetition. Properties can contain other nodes: +Use a namespace import for RDF operations. The namespace gives short operations enough context without forcing long exported names. ```ts -select(['?title', '?publisherName', '?city']) - .where( - node('product', 'schema:Product', { - 'schema:name': v('title'), - 'schema:price': v('price'), - 'schema:publisher': node('publisher', 'schema:Organization', { - 'schema:name': v('publisherName'), - 'schema:location': node('location', 'schema:Place', { - 'schema:city': v('city'), - 'schema:country': v('country') - }) - }) - }) - ) -``` - -That nesting generates all the triples automatically. Product has a publisher, publisher has a location, location has city and country. You write the structure as you think about it. - -ASCII art syntax emphasizes visual clarity. The cypher template tag lets you draw connections: - -```ts -const product = node('product', 'schema:Product', { - 'schema:name': v('title') -}) - -const publisher = node('publisher', 'schema:Organization', { - 'schema:name': v('pubName') -}) - -const query = select(['?title', '?pubName']) - .where(cypher`${product}-[schema:publisher]->${publisher}`) -``` - -The arrow `->` shows the relationship direction visually. This generates the same triples as the object notation, but reads like a diagram. - -Combine patterns with `match()` when you want to build complex structures from separate pieces: - -```ts -const pattern = match( - node('person', 'foaf:Person', { 'foaf:name': v('personName') }), - rel('person', 'foaf:knows', 'friend'), - node('friend', 'foaf:Person', { 'foaf:name': v('friendName') }) -) - -select(['?personName', '?friendName']).where(pattern) -``` +import * as rdf from '@okikio/rdf' -Mix patterns in the same query. Use triples for simple bindings, objects for nested structures, ASCII art for visual relationships: +const product = rdf.namedNode('https://example.com/products/1') +const name = rdf.namedNode('https://schema.org/name') -```ts -select(['?person', '?skill', '?friendName']) - .where(triple('?person', 'ex:hasSkill', '?skill')) - .where( - node('person') - .prop('foaf:knows', node('friend', { - 'foaf:name': v('friendName') - })) - ) -``` - -## Fluent Operations - -Chain operations to build expressions. Arithmetic works left to right: - -```ts -const pricing = select(['?product', '?total']) - .where(triple('?product', 'schema:price', '?basePrice')) - .bind( - v('basePrice') - .mul(1.2) // Apply markup - .add(5) // Add shipping - .round() // Clean up decimals - .as('total') - ) - .filter(v('total').gte(20)) -``` - -String operations chain naturally. Build display names with fallbacks: - -```ts -select(['?displayName']) - .where(triple('?person', 'foaf:firstName', '?first')) - .where(triple('?person', 'foaf:lastName', '?last')) - .optional(triple('?person', 'foaf:nickname', '?nick')) - .bind( - substr(coalesce(v('nick'), v('first')).ucase(), 1, 10) - .concat(' ') - .concat(v('last').ucase().substr(1, 1)) - .concat('.') - .as('displayName') - ) -``` - -Read it step by step: "Use nickname if available, otherwise first name. Uppercase it. Take first 10 characters. Add space. Add uppercased first letter of last name. Add period." - -Note: `substr()` is a standalone function, not a fluent method. Most string operations like `ucase()`, `lcase()`, `concat()`, `strlen()` are available as fluent methods for chaining. - -Conditional logic stays readable with nested fluent operations: - -```ts -select(['?item', '?price', '?status']) - .where(triple('?item', 'schema:basePrice', '?base')) - .where(triple('?item', 'schema:inStock', '?stock')) - .bind( - ifElse( - v('stock').gt(10), - v('base').mul(0.85), - ifElse( - v('stock').gt(0), - v('base').mul(0.95), - v('base').add(20) - ) - ).round().as('price') - ) - .bind( - ifElse( - v('stock').gt(10), - 'In Stock', - ifElse( - v('stock').gt(0), - 'Low Stock', - 'Out of Stock' - ) - ).as('status') - ) -``` - -The nested structure mirrors the decision tree. Good stock gets 15% off, low stock gets 5% off, out of stock adds a premium. - -## Aggregations and Grouping - -Count, average, sum - aggregations work with the fluent API: - -```ts -const analytics = select([ - v('country'), - count().as('users'), - avg(v('age')).as('avgAge'), - countDistinct(v('city')).as('cities'), - sum(v('purchases')).as('revenue') +const data = rdf.dataset([ + rdf.quad(product, name, rdf.literal('Widget')), ]) - .where(triple('?user', 'schema:country', '?country')) - .where(triple('?user', 'foaf:age', '?age')) - .where(triple('?user', 'schema:city', '?city')) - .where(triple('?user', 'ex:totalPurchases', '?purchases')) - .groupBy('?country') - .having(count().gte(10)) - .orderBy('?revenue', 'DESC') -``` - -Notice `count().gte(10)` in the having clause - aggregations return fluent values. Everything chains consistently. -## Subqueries and Composition - -Break complex logic into pieces: - -```ts -// Find top 10 best-selling products -const topSellers = select([v('product'), count().as('sales')]) - .where(triple('?order', 'schema:product', '?product')) - .where(triple('?order', 'schema:date', '?date')) - .filter(v('date').gte('2024-01-01')) - .groupBy('?product') - .orderBy('?sales', 'DESC') - .limit(10) - -// Enrich with product details -const enriched = select(['?product', '?name', '?price', '?sales']) - .where(subquery(topSellers)) - .where( - node('product', { - 'schema:name': v('name'), - 'schema:price': v('price') - }) - ) - .orderBy('?sales', 'DESC') +for (const quad of data.match(product)) { + console.log(quad.object.value) +} ``` -Each piece is simple. Together they solve the complex problem. - -## Property Paths - -Navigate graph structures without manual recursion. Find all contacts through any number of "knows" relationships: +The root exports the semantic model only. Syntax-specific code is opt-in: ```ts -const network = select(['?person', '?contact']) - .where(triple('?person', zeroOrMore('foaf:knows'), '?contact')) - .filter(v('person').neq(v('contact'))) +import * as nquads from '@okikio/rdf/nquads' +import * as turtle from '@okikio/rdf/turtle' +import * as jsonld from '@okikio/rdf/jsonld' +import * as canon from '@okikio/rdf/canon' ``` -That `zeroOrMore` expands to any number of hops. The SPARQL engine handles it efficiently. +Implemented format and semantic subpaths currently include: -Navigate nested properties with sequences: - -```ts -const cities = select(['?person', '?city']) - .where(triple('?person', sequence('schema:address', 'schema:city'), '?city')) +```text +@okikio/rdf/ntriples N-Triples 1.2 parsing and serialization +@okikio/rdf/nquads N-Quads 1.2 parsing and serialization +@okikio/rdf/turtle Turtle 1.2 streaming parser and conservative serializer +@okikio/rdf/trig TriG 1.2 streaming parser and conservative serializer +@okikio/rdf/jsonld JSON-LD processor adapter with bounded document loading +@okikio/rdf/xml RDF/XML streaming adapter +@okikio/rdf/rdfa RDFa 1.1 streaming adapter +@okikio/rdf/microdata Microdata-to-RDF streaming adapter +@okikio/rdf/canon RDFC-1.0 canonicalization +@okikio/rdf/ontology RDFS/OWL ontology interpretation model +@okikio/rdf/shape loss-preserving SHACL shape model ``` -"Follow address property, then city property" - one pattern instead of two triples. +`@okikio/rdf` uses `Iterable`, `AsyncIterable`, `ReadableStream`, `AbortSignal`, and explicit disposal where those shapes match the workload. RDF/JS is an interoperability target, not a required core dependency. -Use alternatives when property names vary: +## SPARQL -```ts -const names = select(['?person', '?name']) - .where(triple('?person', alternative('foaf:name', 'schema:name'), '?name')) -``` - -Combine path operators for complex navigation: +Build queries independently from the engine that executes them. ```ts -// Manager or manager's manager -const bosses = select(['?employee', '?boss']) - .where( - triple( - '?employee', - sequence( - oneOrMore('org:reportsTo'), - alternative('org:manages', 'org:supervises') - ), - '?boss' - ) - ) -``` +import * as sparql from '@okikio/sparql' -## Updates and Modifications +const query = sparql.select(['?product', '?name']) + .prefix('schema', 'https://schema.org/') + .where(sparql.triple('?product', 'schema:name', '?name')) + .orderBy('?name') + .limit(25) -Updates use the same fluent patterns. Increment everyone's age: - -```ts -const birthday = modify() - .delete(triple('?person', 'foaf:age', '?oldAge')) - .insert(triple('?person', 'foaf:age', v('oldAge').add(1))) - .where(triple('?person', 'foaf:age', '?oldAge')) - .where(filter(v('oldAge').gte(0))) - .done() - -await birthday.execute({ endpoint: 'http://localhost:3030/dataset/update' }) +console.log(query.build().value) ``` -Notice `v('oldAge').add(1)` works in the insert template - fluent operations work everywhere. - -Conditional inserts: +Execution is explicit: ```ts -const markSeniors = modify() - .insert( - node('person', { - 'ex:seniorCitizen': true, - 'ex:discount': 0.15 - }) - ) - .where(triple('?person', 'foaf:age', '?age')) - .where(filter(v('age').gte(65))) - .done() -``` +import { createClient } from '@okikio/sparql/http' -Delete patterns with conditions: +const client = createClient({ endpoint: 'https://example.com/sparql' }) -```ts -const cleanup = modify() - .delete( - node('account', { - 'ex:status': v('status'), - 'ex:lastLogin': v('lastLogin') - }) - ) - .where(triple('?account', 'ex:status', 'inactive')) - .where(triple('?account', 'ex:lastLogin', '?lastLogin')) - .where(filter(v('lastLogin').lt('2023-01-01'))) - .done() +for await (const row of await client.queryBindings(query)) { + console.log(row.get('name')?.value) +} ``` -## Complete Example - -Real-world e-commerce query combining multiple patterns: +The result modes stay separate: ```ts -const productSearch = select([ - v('title'), - v('displayPrice'), - v('stockStatus'), - v('categoryName'), - v('averageRating') -]) - .where( - node('product', 'schema:Product', { - 'schema:name': v('title'), - 'schema:price': v('basePrice'), - 'schema:inventory': v('stock'), - 'schema:category': node('category', 'schema:Category', { - 'schema:name': v('categoryName') - }) - }) - ) - .optional( - node('product') - .prop('schema:review', node('review', 'schema:Review', { - 'schema:ratingValue': v('rating') - })) - ) - .bind( - ifElse( - v('stock').gt(5), - v('basePrice').mul(0.9), - ifElse( - v('stock').gt(0), - v('basePrice').mul(0.95), - v('basePrice').mul(1.1) - ) - ).round().as('displayPrice') - ) - .bind( - ifElse( - v('stock').gt(5), - 'In Stock', - ifElse( - v('stock').gt(0), - v('stock').concat(' left'), - 'Out of Stock' - ) - ).as('stockStatus') - ) - .filter(v('displayPrice').gte(10)) - .groupBy('?product', '?title', '?displayPrice', '?stockStatus', '?categoryName') - .bind(avg(v('rating')).as('averageRating')) - .having(countDistinct(v('review')).gte(5)) - .orderBy('?averageRating', 'DESC') - .limit(50) - -const results = await productSearch.execute({ - endpoint: 'http://localhost:3030/catalog/sparql' -}) +client.queryBindings(query) // SELECT +client.queryQuads(query) // CONSTRUCT / DESCRIBE +client.queryBoolean(query) // ASK +client.update(update) // UPDATE ``` -That query handles nested relationships (product → category), optional patterns (reviews), computed values (discounted price), conditional logic (stock messages), aggregations (average rating), and filtering. - -## Named Graphs and Federation +This prevents graph results from being coerced into binding rows and keeps RDF values as RDF terms instead of silently converting datatypes to JavaScript primitives. -Query specific graphs: +`@okikio/sparql/syntax` exposes a source-ranged token and feature event stream for inspection, diagnostics, editors, formatters, and future grammar/algebra parsing. It is intentionally not advertised as a complete SPARQL AST today. -```ts -const graphData = select(['?s', '?p', '?o']) - .fromNamed('http://example.org/metadata') - .where(graph('?g', triple('?s', '?p', '?o'))) -``` +## Generated vocabularies -Combine data from multiple endpoints: +Generated vocabulary data uses direct, tree-shakeable imports. Runtime values, schemas, and types use PascalCase when the vocabulary term is PascalCase. ```ts -const federated = select(['?person', '?name', '?birthPlace', '?abstract']) - .where(triple('?person', 'foaf:name', '?name')) - .where( - service( - 'http://dbpedia.org/sparql', - triple('?person', 'dbo:birthPlace', '?birthPlace'), - triple('?person', 'dbo:abstract', '?abstract') - ) - ) - .filter(v('abstract').regex('scientist')) +import { + Product, + ProductSchema, + type ProductType, + name, + offers, +} from '@okikio/vocab/schema' ``` -Federation lets you query distributed data sources in a single query. The SERVICE clause transparently handles remote execution. +This avoids awkward APIs such as `schema.ProductSchema` while still allowing operation-oriented namespaces elsewhere. -## Type Safety +A generated class normally provides: -Everything is fully typed. TypeScript catches errors at compile time: - -```ts -const age = v('age') - -age.gte(18) // ✓ Returns SparqlValue for filters -age.add(5) // ✓ Returns FluentValue, can chain -age.add(5).mul(2) // ✓ Chains continue naturally -age.gte('not a number') // ✗ TypeScript error +```text +Product RDF named node +ProductType TypeScript JSON-LD/value shape +ProductSchema Standard Schema + Standard JSON Schema validator +ProductPropertiesType ``` -You get autocomplete in your editor. The library guides you toward correct code. Generated SPARQL is safe from injection attacks because values are properly escaped automatically. - -## Query Composition and Reuse +The compiler consumes RDF ontology quads through `@okikio/rdf/ontology`. It does not parse one specific source syntax itself. Schema.org-specific `domainIncludes` and `rangeIncludes` aliases are compiler policy, not RDF core semantics. -Building complex queries from reusable pieces makes your code cleaner and more maintainable. These patterns show how to compose queries, extract common logic, and build flexible query templates. +The checked-in Schema.org module is currently a bootstrap slice used to validate the public API and compiler. The complete upstream vocabulary must be regenerated from an authoritative Schema.org release before publication. -### Reusable Pattern Fragments +## Persistent RDF -Extract common triple patterns into functions. This reduces duplication and makes queries easier to understand: +`@okikio/triplestore` is a persistent RDF dataset, not an OPFS-specific package. It borrows any filesystem satisfying its small structural contract. `@okikio/opfs` is the intended browser/runtime filesystem provider. ```ts -import { RDF, FOAF, SCHEMA } from '@okikio/sparql' - -// Define reusable pattern fragments -function personPattern(personVar = 'person') { - return node(personVar, FOAF.Person, { - [FOAF.name]: v('name'), - [FOAF.age]: v('age') - }) -} - -function addressPattern(personVar = 'person') { - return node(personVar) - .prop(SCHEMA.address, node('address', SCHEMA.PostalAddress, { - [SCHEMA.addressLocality]: v('city'), - [SCHEMA.addressCountry]: v('country') - })) -} +import * as rdf from '@okikio/rdf' +import { open } from '@okikio/triplestore' -// Use them in queries -const adults = select(['?name', '?age']) - .where(personPattern()) - .filter(v('age').gte(18)) +await using store = await open(fileSystem, { path: '/knowledge' }) -const peopleWithAddress = select(['?name', '?city', '?country']) - .where(personPattern()) - .where(addressPattern()) +await store.add(rdf.quad( + rdf.namedNode('https://example.com/1'), + rdf.namedNode('https://schema.org/name'), + rdf.literal('Widget'), +)) ``` -Patterns compose naturally. Mix and match to build exactly the query you need without repeating yourself. - -### Filter Builder Functions +The baseline store uses immutable segments and immutable generation records. Recovery scans for the newest fully valid committed generation and ignores incomplete publications. It deliberately does not depend on filesystem rename being atomic across every adapter. -Extract filter logic into reusable functions. This makes query intentions clearer: +The current implementation is one-writer-per-path. Cross-process writer coordination and physical garbage collection are not claimed yet. -```ts -// Define filter builders -function olderThan(age: number) { - return v('age').gte(age) -} +## Oxigraph and Comunica -function nameMatches(pattern: string, caseInsensitive = true) { - return caseInsensitive - ? v('name').regex(pattern, 'i') - : v('name').regex(pattern) -} +The engine integration packages wrap resources that the caller creates and owns. -function inCountry(country: string) { - return v('country').eq(country) -} +```ts +import { Store } from 'oxigraph' +import { createClient } from '@okikio/oxigraph' -// Use them in queries -const query = select(['?name', '?age']) - .where(personPattern()) - .where(addressPattern()) - .filter(olderThan(21)) - .filter(nameMatches('^John')) - .filter(inCountry('USA')) +const store = new Store() +const client = createClient(store) ``` -Read the query like English: "Select name and age where person is older than 21, name matches '^John', and in country USA." - -### Query Templates with Parameters - -Build flexible query templates that adapt based on parameters. This pattern works great for search interfaces: - ```ts -interface SearchFilters { - minAge?: number - maxAge?: number - namePattern?: string - city?: string - country?: string -} - -function searchPeople(filters: SearchFilters) { - let query = select(['?name', '?age', '?city', '?country']) - .prefix('rdf', getNamespaceIRI(RDF)) - .prefix('foaf', getNamespaceIRI(FOAF)) - .prefix('schema', getNamespaceIRI(SCHEMA)) - .where(personPattern()) - .where(addressPattern()) - - // Add filters conditionally - if (filters.minAge !== undefined) { - query = query.filter(v('age').gte(filters.minAge)) - } - - if (filters.maxAge !== undefined) { - query = query.filter(v('age').lte(filters.maxAge)) - } - - if (filters.namePattern) { - query = query.filter(v('name').regex(filters.namePattern, 'i')) - } - - if (filters.city) { - query = query.filter(v('city').eq(filters.city)) - } - - if (filters.country) { - query = query.filter(v('country').eq(filters.country)) - } - - return query -} +import { QueryEngine } from '@comunica/query-sparql' +import { createClient } from '@okikio/comunica' -// Use it with different filter combinations -const youngAdults = searchPeople({ minAge: 18, maxAge: 25 }) -const londonResidents = searchPeople({ city: 'London' }) -const ukSeniors = searchPeople({ minAge: 65, country: 'UK' }) +const engine = new QueryEngine() +const client = createClient(engine, { + context: () => ({ sources: [/* caller-selected sources */] }), +}) ``` -The query adapts automatically. Only the filters you provide get added. +Importing either adapter does not create an engine or initialize unrelated runtime state. -### Base Query Extension +## Parser model -Start with a base query and extend it for different use cases: +Hot parsers use a data-oriented streaming model: -```ts -// Base query everyone shares -const baseProductQuery = select(['?product', '?name', '?price']) - .prefix('schema', getNamespaceIRI(SCHEMA)) - .where( - node('product', SCHEMA.Product, { - [SCHEMA.name]: v('name'), - [SCHEMA.price]: v('price') - }) - ) - -// Extend for specific needs -const affordableProducts = baseProductQuery - .filter(v('price').lte(50)) - .orderBy('?price') - .limit(20) - -const expensiveProducts = baseProductQuery - .filter(v('price').gte(100)) - .orderBy('?price', 'DESC') - .limit(10) - -const searchResults = baseProductQuery - .filter(v('name').regex(userInput, 'i')) - .limit(50) +```text +source windows + ↓ +mutable scanner state + ↓ +semantic events + ↓ +quads / directives / diagnostics ``` -The base query stays immutable. Each extension creates a new query without affecting the original. +The normal RDF path does not create an AST. Source-ranged event streams exist where callers need tolerant parsing, diagnostics, inspection, or incremental tooling. This follows the same useful principle as `@okikio/wikitext`: the event stream is the factual interchange layer, while materialized structures are optional consumers. -### Combining Subqueries +## Validation and benchmarks -Break complex logic into subqueries then combine them: +Package tests live beside the package code. Runtime benchmarks use Mitata. Compiler/type benchmarks use isolated TypeScript subprocesses because compiler memory and type-instantiation cost are different measurements from hot runtime throughput. -```ts -// Find top sellers -const topSellers = select([v('product'), count().as('sales')]) - .where(triple('?order', SCHEMA.product, '?product')) - .where(triple('?order', SCHEMA.orderDate, '?date')) - .filter(v('date').gte('2024-01-01')) - .groupBy('?product') - .orderBy('?sales', 'DESC') - .limit(10) - -// Enrich with product details -const enrichedProducts = select([ - '?product', - '?name', - '?price', - '?sales', - '?category' -]) - .where(subquery(topSellers)) - .where( - node('product', { - [SCHEMA.name]: v('name'), - [SCHEMA.price]: v('price'), - [SCHEMA.category]: v('category') - }) - ) - .orderBy('?sales', 'DESC') -``` +The permanent benchmark questions include: -Each subquery solves one piece of the problem. Combine them to solve the whole thing. +- RDF term creation and equality +- Dataset scan versus indexed matching +- N-Triples/N-Quads parsing and chunk sizes +- Turtle/TriG parsing +- SPARQL syntax streaming versus materialization +- persistent store read/write/recovery behavior +- vocabulary generation time and emitted bytes +- TypeScript compiler time, memory, symbols, types, and instantiations +- generated schema validation +- optional engine integration where the upstream package is available -### Pattern Collections +Benchmarks must state the semantic oracle. A faster result is not accepted when it performs less work or changes the result. -Group related patterns together for complex domains: +## Development -```ts -// E-commerce patterns -const ecommerce = { - product(productVar = 'product') { - return node(productVar, SCHEMA.Product, { - [SCHEMA.name]: v('productName'), - [SCHEMA.price]: v('price'), - [SCHEMA.sku]: v('sku') - }) - }, - - order(orderVar = 'order') { - return node(orderVar, SCHEMA.Order, { - [SCHEMA.orderDate]: v('orderDate'), - [SCHEMA.orderNumber]: v('orderNumber'), - [SCHEMA.customer]: v('customer') - }) - }, - - customer(customerVar = 'customer') { - return node(customerVar, SCHEMA.Person, { - [SCHEMA.name]: v('customerName'), - [SCHEMA.email]: v('email') - }) - }, - - // Relationship connectors - orderContains(orderVar = 'order', productVar = 'product') { - return triple(`?${orderVar}`, SCHEMA.orderedItem, `?${productVar}`) - }, - - orderBy(orderVar = 'order', customerVar = 'customer') { - return triple(`?${orderVar}`, SCHEMA.customer, `?${customerVar}`) - } -} +The production project is Deno-native: -// Use pattern collection -const orderAnalysis = select([ - '?orderNumber', - '?customerName', - '?productName', - '?price' -]) - .where(ecommerce.order()) - .where(ecommerce.orderBy()) - .where(ecommerce.customer()) - .where(ecommerce.orderContains()) - .where(ecommerce.product()) - .filter(v('orderDate').gte('2024-01-01')) +```sh +deno task fmt +deno task lint +deno task check +deno task test +deno task bench +deno task verify ``` -Pattern collections organize domain knowledge. Your queries become high-level descriptions of what you want to find. - -### Dynamic Query Construction - -Build queries programmatically from user input or configuration: - -```ts -interface FieldSelection { - fields: string[] - filters: Array<{ field: string, operator: string, value: any }> - sort?: { field: string, direction: 'ASC' | 'DESC' } - limit?: number -} - -function buildDynamicQuery(config: FieldSelection) { - // Map user field names to actual predicates - const fieldMap: Record = { - name: FOAF.name, - age: FOAF.age, - email: SCHEMA.email, - city: SCHEMA.addressLocality - } - - // Start with base pattern - let query = select(config.fields.map(f => `?${f}`)) - .where(triple('?person', RDF.type, uri(FOAF.Person))) - - // Add triples for each requested field - for (const field of config.fields) { - const predicate = fieldMap[field] - if (predicate) { - query = query.where(triple('?person', predicate, `?${field}`)) - } - } - - // Apply filters - for (const filter of config.filters) { - const varRef = v(filter.field) - switch (filter.operator) { - case 'eq': - query = query.filter(varRef.eq(filter.value)) - break - case 'gt': - query = query.filter(varRef.gt(filter.value)) - break - case 'lt': - query = query.filter(varRef.lt(filter.value)) - break - case 'contains': - query = query.filter(varRef.contains(filter.value)) - break - } - } - - // Apply sorting - if (config.sort) { - query = query.orderBy(`?${config.sort.field}`, config.sort.direction) - } - - // Apply limit - if (config.limit) { - query = query.limit(config.limit) - } - - return query -} +Vocabulary compilation is a library capability. File I/O for the repository task lives under `.mise/tasks/`: -// Use with different configurations -const query1 = buildDynamicQuery({ - fields: ['name', 'email'], - filters: [{ field: 'name', operator: 'contains', value: 'John' }], - limit: 10 -}) +```sh +deno task vocab --input ontology.nq --out packages/vocab/example --name example --namespace https://example.com/ --prefix ex -const query2 = buildDynamicQuery({ - fields: ['name', 'age', 'city'], - filters: [ - { field: 'age', operator: 'gt', value: 18 }, - { field: 'city', operator: 'eq', value: 'London' } - ], - sort: { field: 'age', direction: 'DESC' } -}) +# Reproduce the pinned complete Schema.org release when network access is available. +deno task vocab:schema ``` -Dynamic construction turns configuration into queries. Good for building query builders, APIs, or UI-driven search. - -### Cached Query Builders - -Pre-configure query builders for common operations: +`.agents/` is reserved for disposable assistant validation only and is ignored by the repository. Permanent tests, benchmarks, fixtures, and release evidence belong in project-owned package/workspace locations. See [`docs/testing.md`](./docs/testing.md). -```ts -// Create specialized builders -class PersonQueries { - private readonly baseQuery: QueryBuilder - - constructor() { - this.baseQuery = select(['?person', '?name', '?age']) - .prefix('rdf', getNamespaceIRI(RDF)) - .prefix('foaf', getNamespaceIRI(FOAF)) - .where(triple('?person', RDF.type, uri(FOAF.Person))) - .where(triple('?person', FOAF.name, '?name')) - .where(triple('?person', FOAF.age, '?age')) - } - - all() { - return this.baseQuery - } - - adults() { - return this.baseQuery.filter(v('age').gte(18)) - } - - children() { - return this.baseQuery.filter(v('age').lt(18)) - } - - byName(name: string) { - return this.baseQuery.filter(v('name').eq(name)) - } - - olderThan(age: number) { - return this.baseQuery.filter(v('age').gt(age)) - } - - byAgeRange(min: number, max: number) { - return this.baseQuery - .filter(v('age').gte(min)) - .filter(v('age').lte(max)) - } -} +## Current implementation status -// Use it -const people = new PersonQueries() - -const adults = await people.adults().execute(config) -const seniors = await people.olderThan(65).execute(config) -const alice = await people.byName('Alice').execute(config) -const millennials = await people.byAgeRange(25, 40).execute(config) -``` +Implemented and locally validated without optional upstream packages: -Query builders encapsulate domain logic. Each method returns a ready-to-execute query. +- RDF 1.2-native core terms, triple terms, directional literals, Dataset indexes +- streaming N-Triples/N-Quads +- streaming Turtle/TriG semantic parser +- RDF ontology and SHACL shape IRs +- SPARQL builder migration and explicit result-mode contracts +- SPARQL HTTP protocol client +- source-ranged SPARQL syntax inspection +- vocabulary compiler/runtime/bootstrap Schema.org surface +- crash-recoverable triplestore +- caller-owned Oxigraph and Comunica adapters +- lifecycle adapters for JSON-LD, RDF/XML, RDFa, Microdata, and RDFC-1.0 -## Further Reading +Still requiring canonical upstream validation before release claims: -Check the module documentation for architecture details and progressive examples. The pattern files (triples.ts, objects.ts, cypher.ts) show different syntax styles with comprehensive examples. Implementation docs explain technical decisions and trade-offs. +- Deno formatting, linting, type checks, and package-local tests on a Deno/JSR-capable host +- W3C RDF/Turtle/TriG/SPARQL/SHACL conformance corpora +- real `jsonld`, `rdf-canonize`, RDF/XML, RDFa, and Microdata dependency integration +- real Oxigraph Wasm and Comunica engine integration +- complete generated Schema.org vocabulary +- package builds and JSR/npm publication artifacts -Start simple with basic queries. Add complexity as you need it. The patterns stay consistent - learning one part teaches you the rest. +See [`docs/standard-schema.md`](./docs/standard-schema.md) for generated validation, [`docs/vocabulary-generation.md`](./docs/vocabulary-generation.md) for the `ttl-to-ts` replacement, [`docs/testing.md`](./docs/testing.md) for permanent test ownership, [`docs/migration.md`](./docs/migration.md) for intentional 0.0.2 API changes, [`docs/benchmarks.md`](./docs/benchmarks.md) for benchmark-derived decisions, and [`VALIDATION.md`](./VALIDATION.md) for the exact passed and blocked release gates. ## License -MIT \ No newline at end of file +MIT -- 2.51.2