From 835b461475cfb36ef11f9da04e62b5f240b8c254 Mon Sep 17 00:00:00 2001 From: Okiki Ojo Date: Wed, 22 Apr 2026 02:48:56 -0400 Subject: [PATCH] docs: add comprehensive examples for parser usage and diagnostics Variety of examples demonstrating how to effectively use the parser. The examples cover different scenarios, including building a tree, inspecting headings, and handling diagnostics. This addition aims to enhance user understanding and facilitate better integration of the parser into various workflows. Additionally, it outlines the future direction of the parser, emphasizing the importance of maintaining reusable primitives and avoiding product-specific behavior. The document also discusses the evaluation of block-to-inline handoff and the core parser's handling of bare URLs, ensuring clarity on the parser's capabilities and limitations. Signed-off-by: Okiki Ojo --- docs/api-reference.md | 92 ++++ docs/architecture/api-direction.md | 232 +++++++++ docs/architecture/diagnostic-anchors.md | 78 +++ docs/architecture/diagnostics-first.md | 166 ++++++ docs/architecture/parser-contracts.md | 156 ++++++ docs/architecture/performance.md | 98 ++++ docs/architecture/sessions-and-streaming.md | 117 +++++ docs/architecture/utility-first.md | 101 ++++ docs/diagnostics-first-redesign.md | 550 ++++++++++++++++++++ docs/examples.md | 163 ++++++ docs/future-direction.md | 168 ++++++ docs/handoff-and-bare-url-notes.md | 165 ++++++ docs/parser-architecture-comparison.md | 150 ++++++ 13 files changed, 2236 insertions(+) create mode 100644 docs/api-reference.md create mode 100644 docs/architecture/api-direction.md create mode 100644 docs/architecture/diagnostic-anchors.md create mode 100644 docs/architecture/diagnostics-first.md create mode 100644 docs/architecture/parser-contracts.md create mode 100644 docs/architecture/performance.md create mode 100644 docs/architecture/sessions-and-streaming.md create mode 100644 docs/architecture/utility-first.md create mode 100644 docs/diagnostics-first-redesign.md create mode 100644 docs/examples.md create mode 100644 docs/future-direction.md create mode 100644 docs/handoff-and-bare-url-notes.md create mode 100644 docs/parser-architecture-comparison.md diff --git a/docs/api-reference.md b/docs/api-reference.md new file mode 100644 index 0000000..6062706 --- /dev/null +++ b/docs/api-reference.md @@ -0,0 +1,92 @@ +# API Reference + +This page is the compact reference index for the currently shipped public API. + +If you are new to the package, start with [readme.md](../readme.md) first. +That page explains what the parser is for, how to install it, and which output +shape to choose. + +The reference below assumes the same core model as the architecture docs: +parser outputs are range-first, source-backed where practical, and grounded in +UTF-16 offsets. + +This page is for lookup, not for the full design explanation. + +- for install and first use, go to [readme.md](../readme.md) +- for task-focused snippets, go to [examples.md](./examples.md) +- for parser trade-offs, malformed-input behavior, and tree-policy reasoning, + go to [architecture/README.md](./architecture/README.md) + +## Pipeline modules + +| Module | Purpose | Status | +|--------|---------|--------| +| `text_source.ts` | Abstracts the backing text store (string, rope, CRDT) while preserving UTF-16 source access | Published | +| `token.ts` | Token type constants and `Token` interface | Published | +| `events.ts` | Event stream types, constructors, type guards, and diagnostics | Published | +| `ast.ts` | Wikist AST node types, type guards, and builders | Published | +| `tokenizer.ts` | `charCodeAt` generator scanner over `TextSource`, producing offset-based tokens | Published | +| `block_parser.ts` | Block-level event emitter over source-backed token and text ranges | Published | +| `inline_parser.ts` | Inline event enrichment over source-backed text ranges | Published | +| `parse.ts` | Orchestration (tokenizer, block, inline, tree) | Published | +| `tree_builder.ts` | `buildTree(events, { source })` to `WikistRoot`, preserving source-backed positions | Published | +| `stringify.ts` | AST to wikitext (round-trip) | Not yet implemented | +| `filter.ts` | Filter and visit utilities for trees and event streams | Published | +| `session.ts` | Stateful `Session` wrapper for repeated sync access | Published | + +## Parsing and serialization + +| Function | Description | +|----------|-------------| +| `parse(input)` | Parse wikitext into a `WikistRoot` AST using the cheapest default tree lane. | +| `parseWithDiagnostics(input)` | Parse to `{ tree, diagnostics }` with the default tolerant materialization and preserved diagnostics. | +| `parseStrictWithDiagnostics(input)` | Parse to `{ tree, diagnostics }` with the conservative source-strict materialization. | +| `stringify(tree)` | Serialize a wikist tree back to wikitext. Not yet implemented. | +| `events(input)` | Full event stream (block + inline), with source-backed text ranges and optional diagnostics. | +| `outlineEvents(input)` | Block-only event stream over the same source-backed structure. | +| `parseChunked(chunks)` | Progressive completed block nodes. Not yet implemented. | + +## Low-level streams and tree building + +| Function | Description | +|----------|-------------| +| `tokenize(input)` | Raw token generator stream with offset-based token spans. | +| `tokens(input)` | Raw token generator alias for the sync public API. | +| `blockEvents(source, tokens)` | Block-level event stream from tokens. | +| `inlineEvents(source, blockEvents)` | Inline event enrichment over block events, preserving source-backed text ranges where possible. | +| `buildTree(events, { source })` | Build AST from an event iterable plus source. | +| `buildTreeWithDiagnostics(events, { source })` | Build `{ tree, diagnostics }` with the default HTML-like materialization. | +| `buildTreeStrict(events, { source })` | Build `{ tree, diagnostics }` with the conservative source-strict materialization. | +| `buildTreeWithRecovery(events, { source })` | Build `{ tree, diagnostics, recovered }` with the default tree plus an explicit recovery summary. | + +## Tree utilities + +| Function | Description | +|----------|-------------| +| `filter(tree, type)` | Get all nodes of a type recursively. | +| `visit(tree, visitor)` | Pre-order tree walker. | +| `resolveTreePath(tree, path)` | Resolve a root-relative child-index path back to a node. | +| `resolveDiagnosticAnchor(tree, anchor)` | Resolve a diagnostic anchor back to the nearest node in one materialized tree. | +| `locateDiagnostic(tree, diagnostic)` | Resolve a diagnostic's anchor back to the nearest node while the diagnostic's source span stays authoritative. | +| `createSession(source)` | Cached sync wrapper for repeated access to one source input. | + +## Foundation exports + +| Export | Module | Description | +|--------|--------|-------------| +| `TextSource` | `text_source.ts` | Interface for backing text stores with UTF-16 source access. | +| `slice()` | `text_source.ts` | Resolve an offset range to a string when a caller needs materialized text. | +| `TokenType` | `token.ts` | Constant map of all token types. | +| `Token` | `token.ts` | Token interface with type plus UTF-16 start and end offsets. | +| `isToken()` | `token.ts` | Type guard for token validation. | +| `tokenize()` | `tokenizer.ts` | Generator-based `charCodeAt` scanner. | +| `blockEvents()` | `block_parser.ts` | Block-level event generator from tokens. | +| `EnterEvent`, `ExitEvent`, `TextEvent`, `TokenEvent`, `ErrorEvent` | `events.ts` | Concrete event interfaces, including source-backed text and diagnostic events. | +| `WikitextEvent` | `events.ts` | Discriminated union of event kinds. | +| `ErrorEventOptions`, `DiagnosticSeverity` | `events.ts` | Diagnostic support types. | +| `enterEvent()`, `exitEvent()`, ... | `events.ts` | Event constructors. | +| `isEnterEvent()`, `isExitEvent()`, ... | `events.ts` | Event type guards. | +| `WikistNode`, `WikistRoot`, `WikistNodeType` | `ast.ts` | Core AST unions and aliases. | +| `WikistParent`, `WikistLiteral`, `WikistVoid` | `ast.ts` | Structural AST category unions. | +| `root()`, `heading()`, `text()`, ... | `ast.ts` | AST builder functions. | +| `isRoot()`, `isHeading()`, `isParent()`, ... | `ast.ts` | AST type guards. | \ No newline at end of file diff --git a/docs/architecture/api-direction.md b/docs/architecture/api-direction.md new file mode 100644 index 0000000..046497f --- /dev/null +++ b/docs/architecture/api-direction.md @@ -0,0 +1,232 @@ +# API Direction + +This note explains the current public surface and the likely direction of the +next cleanup. + +It is not a promise that the public API will change immediately. It is a way to +write down the target shape clearly enough that the docs, code, and tests can +move toward the same model. + +## Why this note exists + +The parser docs have been trying to explain three things at once: + +- what callers can use today +- what those current wrappers really mean +- where the public surface likely wants to go next + +Those are related, but they are not the same question. + +This note keeps the future-facing public-surface discussion separate from the +user-facing "which result do I want?" explanation in +[choosing-a-parser-result.md](./choosing-a-parser-result.md). + +It also assumes the same range-first base as the rest of the architecture docs: +diagnostics, events, and text-like tree content should stay anchored to source +spans and UTF-16 offsets for as long as possible. + +## Current public shape + +Today the main wrappers are: + +```text +parse() + -> default tree + +parseWithDiagnostics() + -> default tree + diagnostics + +parseWithRecovery() + -> default tree + diagnostics + recovery summary + +parseStrictWithDiagnostics() + -> conservative tree + diagnostics +``` + +That surface is already more honest than the older docs, but it still mixes two +different questions: + +1. do you want diagnostics? +2. how should malformed regions be materialized? + +## The direction the surface is moving toward + +The cleaner long-term public story is: + +```text +diagnostics are one choice +materialization policy is another choice +findings are the replayable bridge between them +``` + +That makes the API easier to reason about. + +Instead of teaching several top-level parser truths, the surface can teach a +smaller set of orthogonal choices: + +- tree now, or findings first +- diagnostics off or on +- package-owned or caller-owned materialization policy + +## One concrete target shape + +One reasonable target would look like this. + +```ts +interface ParseOptions { + readonly diagnostics?: boolean; + readonly materialization?: 'default-html-like' | 'source-strict'; +} + +interface ParseOutput { + readonly tree: WikistRoot; + readonly diagnostics?: readonly ParseDiagnostic[]; +} +``` + +That does not have to be the final naming. The point is the shape: + +- one option controls whether diagnostics are collected +- another option controls the final tree policy + +`diagnostics` should be the public boolean name everywhere. It reads like a +caller-facing lane choice instead of an internal implementation switch. + +Under that surface, the parser should still keep its range-first behavior. +Changing the API shape should not imply a move toward eagerly copied strings or +away from source-backed text interpretations. + +Convenience wrappers could still exist on top of that base surface. + +## Lane 3: concrete `analyze()` proposal + +There is one stronger diagnostics-first lane that the public API does not fully +offer yet. + +```text +diagnostics and recovery data are exposed +but final materialization is delayed or caller-owned +``` + +That would go beyond `parseStrictWithDiagnostics()`. `parseStrictWithDiagnostics()` still chooses a final tree +policy for the caller. A true `analyze()` lane would expose parser findings +first and let later tools decide which repairs to apply. + +That kind of lane is especially compatible with a range-first design because it +lets the parser preserve diagnostics and source spans now, then defer heavier +materialization choices until a caller actually needs them. + +One practical shape could look more like this: + +```ts +interface AnalyzeOptions { + readonly recovery?: boolean; +} + +interface ParseRecovery { + readonly kind: + | 'missing-close' + | 'unterminated-opener' + | 'eof-autoclose' + | 'mismatched-exit'; + readonly anchor: ParseDiagnosticAnchor; + readonly node_type?: WikistNodeType; + readonly policies: readonly TreeMaterializationPolicy[]; +} + +interface ParseFindings { + readonly source: TextSource; + readonly events: readonly WikitextEvent[]; + readonly diagnostics: readonly ParseDiagnostic[]; + readonly recovery?: readonly ParseRecovery[]; +} + +function analyze(source: TextSource, options?: AnalyzeOptions): ParseFindings; + +function materialize( + findings: ParseFindings, + options?: { readonly policy?: TreeMaterializationPolicy }, +): ParseOutput; +``` + +The important part is the ownership model: + +- the parser exposes replayable findings +- the caller chooses whether to materialize them at all +- if the caller does materialize them, the tree policy is explicit + +That is why a plain one-shot generator is probably not enough as the whole +public lane. The event stream should stay the core primitive, but many callers +will need to inspect diagnostics and materialize more than once without +rerunning the parse. + +The main design bet is that findings should stay narrow and factual. A useful +first public version should probably include: + +- replayable events +- diagnostics +- a small, parser-owned recovery vocabulary + +It should not start by exposing a broad plugin or callback system. + +## Lane 4: exploratory policy proposal + +There is one possible lane after that, but it should stay explicitly more +tentative. + +```text +the caller starts from parse findings +the caller chooses some recovery outcomes itself +the final tree still comes from a later materialization step +``` + +One concrete exploration could look like this: + +```ts +type RecoveryDecision = 'keep-structural' | 'collapse-to-text'; + +interface CustomMaterializeOptions { + readonly policy: 'custom'; + readonly resolve_recovery: ( + recovery: ParseRecovery, + ) => RecoveryDecision; +} + +function materialize( + findings: ParseFindings, + options: CustomMaterializeOptions, +): ParseOutput; +``` + +That would be powerful for editors, migration tools, and profile-specific +compatibility layers, but it also hardens the recovery taxonomy and +callback timing much earlier than the package may want. + +That is the core trade-off. + +- Lane 3 makes parser facts explicit. +- Lane 4 starts making parser policy negotiable. + +Lane 3 is easier to keep stable because it mostly exposes what the parser +already knows. Lane 4 is harder because it starts exposing when and how those +findings are turned into structure. + +The likely order is: + +1. make the `analyze()` lane real +2. make the package-owned materializers explicit +3. only then decide whether a caller-owned policy lane belongs in public API + +That is the kind of thing I meant earlier by an "API proposal doc": a short +design note that describes a possible future public surface before code changes +lock it in. + +## What this note is not + +This note is not: + +- a release plan +- a promise that names or wrappers will change now +- a replacement for the user-facing parser-result docs + +It is just a way to make the future public-surface direction legible. \ No newline at end of file diff --git a/docs/architecture/diagnostic-anchors.md b/docs/architecture/diagnostic-anchors.md new file mode 100644 index 0000000..064bc1a --- /dev/null +++ b/docs/architecture/diagnostic-anchors.md @@ -0,0 +1,78 @@ +# Diagnostic Anchors + +Diagnostics live beside the tree, not inside it. + +That means the parser needs a small way to point back from a diagnostic to the +relevant region of the final materialized tree. + +That design works best because diagnostics are also range-first. They start as +parser findings about positions in the original source, then carry a narrow +tree-local anchor so tooling can reconnect those findings to one materialized +view. + +## What an anchor is + +Today each diagnostic carries an `anchor`. + +That anchor is intentionally narrow. It is a tree-path snapshot to the nearest +materialized node around the diagnostic point. + +```text +root +├─ paragraph path [0] +│ └─ bold path [0, 0] +└─ table path [1] +``` + +That shape is enough for tooling to answer practical questions such as: + +- which node is nearest to this diagnostic? +- where in the final tree should I highlight or inspect? + +The important part is what the anchor does not replace. It does not replace the +source range. The diagnostic still fundamentally refers to a location in the +original text. The anchor just helps map that finding into one tree. + +## Why diagnostics are not tree nodes + +The AST is meant to represent document structure and meaning. + +Diagnostics are parser findings about malformed input or continuation events. +They are important, but they are not normal content nodes. + +Keeping them beside the tree instead of inside it makes the AST easier to use +for normal document transforms while still letting tooling recover the relevant +location when needed. + +That separation also keeps the tree from pretending diagnostics are ordinary +content. A malformed span can still be source-backed text or a recovered +structure in the tree while the diagnostic remains a parallel parser finding. + +## Current helper path + +Callers do not have to resolve anchors by hand. + +Use: + +- `resolveDiagnosticAnchor(tree, diagnostic.anchor)` +- `locateDiagnostic(tree, diagnostic)` + +These helpers turn the stored path back into a concrete node, parent, and child +index. + +## What anchors do not promise yet + +Current anchors are tree-local, not edit-stable. + +They resolve against one final materialized tree. They do not yet promise: + +- stable identity across later edits +- slot semantics across reparses +- session-backed cross-edit anchor durability + +That stronger anchor model belongs to later session and edit-tracking work. + +Until then, the safest thing a caller can trust is: + +- the diagnostic's source range is authoritative +- the anchor is a helper for one final materialized tree \ No newline at end of file diff --git a/docs/architecture/diagnostics-first.md b/docs/architecture/diagnostics-first.md new file mode 100644 index 0000000..b2c290a --- /dev/null +++ b/docs/architecture/diagnostics-first.md @@ -0,0 +1,166 @@ +# Diagnostics-First Model + +This note explains the diagnostics-first idea without the full redesign +discussion. + +Use it when you want the architecture model, not the whole design-history note. + +If you want the caller-facing tree choice first, read +[choosing-a-parser-result.md](./choosing-a-parser-result.md). + +If you want the future public-surface direction, read +[api-direction.md](./api-direction.md). + +## The short version + +Diagnostics-first means the parser should be clearest about what it found +before it is clever about how to repair it. + +That does not remove tolerant parsing. It separates three jobs that were too +easy to blur together. + +## The core split + +The parser does three different jobs that used to blur together in the docs. + +```text +diagnostic = factual parser finding +continuation = minimal internal behavior that lets parsing proceed +materialization = caller-visible choice about the final tree shape +``` + +That split matters because `never throw` does not mean "the parser must pick +one canonical repaired tree for everyone." It means the parser keeps going and +preserves what it found. + +In plain English: + +- diagnostics answer "what went wrong here?" +- continuation answers "how did the parser keep going?" +- materialization answers "what final tree should the caller get?" + +When a caller wants to delay that third choice, it needs an `analyze()` lane: +a replayable package of events, diagnostics, and later possible recovery data +that can feed one or more explicit materializers. + +There is also one more exploratory step beyond that: a policy lane where the +caller decides some recoveries itself. That should be treated as a layer on top +of `analyze()`, not as a new parser truth. + +## What callers are really choosing today + +Most callers are choosing between these lanes: + +```text +parse() + -> cheap default tree + +parseWithDiagnostics() + -> default tolerant tree + diagnostics + +parseStrictWithDiagnostics() + -> conservative tree + diagnostics + +parseWithRecovery() + -> default tolerant tree + diagnostics + recovery summary +``` + +The key point is that `parseWithDiagnostics()` and `parseStrictWithDiagnostics()` are not two +different parser truths. They are two materializations of the same parser +findings. + +That is the heart of the model. Malformed input should not force the docs into +teaching several competing parser realities. + +## What diagnostics-first does not mean + +It does not mean the parser stops doing continuation work internally. + +The parser still has to: + +- keep the event stream well-formed +- survive malformed input +- auto-close or normalize enough internal state to stay usable + +The real question is whether those continuation steps become one mandatory +caller-facing tree shape or remain facts plus materialization choices. + +That is why diagnostics-first is not a "strict parser" slogan. The parser can +stay forgiving and still avoid turning one repair policy into the only public +story. + +## Why this matters for malformed input + +Malformed input is where the distinction becomes visible. + +The parser should stay explicit about structural evidence and commitment points. + +```text +before commitment + -> keep text-backed interpretation + +after commitment + -> keep the structural finding + -> let the chosen tree policy decide how much survives in the final tree +``` + +That is why the same malformed input can yield: + +- a tolerant structural node in the default lane +- plain text in the conservative lane +- and the same diagnostics in both + +The parser finding stays the same. What changes is the caller-facing tree +policy. + +## Why this matters for docs and APIs + +If the docs lead with wrapper names alone, they make it sound like the parser +owns several different malformed-input truths. + +The cleaner explanation is: + +1. the parser finds something malformed +2. the parser keeps going +3. the caller chooses whether to ask for a tree now, or keep findings first +4. if the caller does want a tree, the caller chooses how much repair should + show up in the final tree + +That same split should shape the public API over time. + +## The still-open next lane + +There is one stronger diagnostics-first idea that is not fully a public lane +yet: + +```text +parser findings are preserved +recovery data is exposed +final materialization is delayed or caller-owned +``` + +That would go beyond `parseStrictWithDiagnostics()`. `parseStrictWithDiagnostics()` still chooses a final +tree policy for the caller. A true findings lane would expose findings first +and let the caller decide later which recoveries to keep, discard, or replace. + +The event stream should remain the primitive underneath that lane, but a raw +generator is probably not enough as the public shape. A caller may need to +read diagnostics, compare recoveries, and materialize multiple +different tree policies from the same parse. + +The practical target is closer to this: + +```text +source + -> analyze() + events + diagnostics + recovery data + -> materialize(default-html-like) when wanted + -> materialize(source-strict) when wanted + -> materialize(custom policy) only if that later lane proves worth exposing +``` + +That is a real future direction, but the docs should name it as an open design +question rather than implying that the current API already provides it. + +The longer reasoning, rollout questions, and open design questions stay in +[docs/diagnostics-first-redesign.md](../diagnostics-first-redesign.md). \ No newline at end of file diff --git a/docs/architecture/parser-contracts.md b/docs/architecture/parser-contracts.md new file mode 100644 index 0000000..2c41e0e --- /dev/null +++ b/docs/architecture/parser-contracts.md @@ -0,0 +1,156 @@ +# Parser Contracts + +These are the rules every parser API path needs to preserve. + +If a future optimization is faster but breaks one of these contracts, that is a +behavior change, not a harmless refactor. + +## Why this matters + +This repo exposes several views of the same source: + +- tokens +- block-only events +- full events +- tree materializations +- later session and streaming views + +Those outputs are only useful together if they keep telling the same story. + +That shared story is range-first. The parser should preserve the original +source spans and UTF-16 offsets consistently enough that callers can move +between tokens, events, diagnostics, and trees without losing their bearings. + +## Contract 1: event well-formedness + +Every `enter(X)` must have a matching `exit(X)`. + +In plain English, the event stream has to nest like balanced parentheses. + +That matters because tree builders, stream consumers, and direct event walkers +all rely on the same stack discipline. + +## Contract 2: UTF-16 offsets are authoritative + +`position.offset` uses UTF-16 code unit indexing. + +That matches: + +- `string.charCodeAt(i)` +- `string.slice(start, end)` +- the Language Server Protocol's required UTF-16 compatibility path + +When a caller needs exact source fidelity, offsets and source slices win over +derived convenience fields. + +That also means text-like data should stay source-backed for as long as +possible. A later tree node, event, or diagnostic may expose a convenient text +view, but the authoritative grounding is still the original source range. + +## Contract 3: never throw + +The parser must produce a usable result for any input. + +That does not mean every malformed region becomes a clean structure. +It means malformed input should become: + +- text when the source never committed to structure +- tolerant structure plus diagnostics when the source did commit +- or a more conservative tree when the caller asks for it + +But the parser itself does not get to crash. + +The never-throw contract does not authorize silent rewriting of the source. It +authorizes continuation while keeping malformed regions attached to either a +source-backed text interpretation or a clearly signaled structural finding. + +## Contract 4: determinism + +The same input and the same config must produce the same output. + +That includes: + +- token order +- event nesting +- diagnostics +- tree shape for a given materialization policy + +Determinism is what makes corpus regression testing and later incremental work +trustworthy. + +## Contract 5: cross-mode block consistency + +The default structural story must stay aligned across the main parser views. + +```text +outlineEvents() + == block structure from events() + == block structure from parse() +``` + +This matters because `outlineEvents()` is meant to be the cheap structural +overlay. A caller should be able to use it first, then pay for more detail +later, without discovering that the parser changed its mind about which blocks +exist. + +`parseStrictWithDiagnostics()` is the one intentional exception. It may collapse a malformed +committed region back to text, so it is allowed to diverge from that default +overlay. + +Even there, "back to text" should still mean back to a source-backed text span, +not an invented replacement string. + +## Contract 6: commitment points are real boundaries + +A suspicious opener is not enough to create structure. + +The parser should commit structure only when the source crosses the relevant +recognition boundary. + +The clearest current example is HTML-like and extension-like tags: + +```text +before `>` -> keep text-backed interpretation +after `>` -> opener is structurally real +``` + +That is why these two cases behave differently: + +```text + source-backed text in the most conservative interpretation + +body + -> committed structure in the default tolerant lane +``` + +Boundary cases may still be policy questions for more HTML-like recovery. The +contract here is narrower: the parser should use real recognition boundaries +instead of inventing structure from weak evidence. + +## Contract 7: diagnostics are factual parser findings + +Diagnostics should describe what the parser found and what went wrong. + +They should not silently pretend that one recovery materialization is the only +possible meaning of malformed input. + +That is why the repo increasingly treats diagnostics and final tree shape as +separate concerns. + +## What callers can trust after malformed input + +These are the practical trust rules. + +1. Offsets and source slices stay authoritative. +2. The default tolerant lane keeps the main structural overlay stable. +3. `parseWithDiagnostics()` keeps the same default tolerant tree as `parse()`. +4. `parseStrictWithDiagnostics()` is intentionally more conservative and may collapse a + committed malformed region back to text. +5. If a construct never reaches its commitment point, it stays text-backed in + every current tree lane. +6. Diagnostic anchors are tree-local today. They resolve against one final tree + and do not yet promise edit-stable identity across later session changes. + +In all of those cases, "text-backed" means rooted in the original source span, +not eagerly copied into fresh replacement strings. \ No newline at end of file diff --git a/docs/architecture/performance.md b/docs/architecture/performance.md new file mode 100644 index 0000000..3c235a2 --- /dev/null +++ b/docs/architecture/performance.md @@ -0,0 +1,98 @@ +# Performance Model + +The parser's performance work follows one rule: + +remove repeated work without changing the source ranges the parser reports. + +That wording is deliberate. In this repo, source ranges are not an incidental +detail. They are the main way the parser preserves exact source fidelity while +still avoiding unnecessary string allocation. + +## What kinds of changes are valid + +These are the kinds of optimizations that fit the architecture. + +- coalescing adjacent block-parser text events into larger contiguous ranges +- skipping inline rescans when a merged text group contains no possible inline + opener +- avoiding temporary arrays or closures in hot parser paths +- using compact lookup tables for fixed parser vocabularies where repeated + membership checks matter + +These optimizations all fit the same pattern: do less repeated work while still +pointing at the same original text. + +## What kinds of changes are not valid + +An optimization is not acceptable if it: + +- trims real user content that still belongs to a block +- rewrites spacing inside ordinary text ranges +- removes structural boundaries such as table-cell separators +- makes block parsing depend on inline meaning + +The parser can merge ranges and skip redundant work. It cannot change what part +of the source those ranges actually mean. + +It also should not eagerly replace range-first data with copied strings unless a +consumer-facing need clearly justifies the cost. + +## The main current handoff optimization + +The most important recent example is the handoff from `blockEvents()` to +`inlineEvents()`. + +The useful change is: + +```text +old shape + many neighboring text fragments + -> inline parser merges them again + +new shape + one larger contiguous prose range + -> inline parser scans once +``` + +That saves work because the inline parser only needs accurate source coverage. +It does not need old tokenizer-sized boundaries if they no longer carry useful +meaning. + +That is a good example of the repo's performance philosophy. The parser does +not get faster by knowing less about the source. It gets faster by carrying the +same source meaning with fewer duplicated objects and fewer duplicate scans. + +## Why eager positions still matter + +The parser still pays a real cost for eagerly materialized positions. + +Each event currently carries nested start and end points with: + +- line +- column +- offset + +That is valuable for tooling, but it is also a meaningful allocation and +computation cost. That is why benchmark work should keep isolating how much of +the total time comes from: + +- event creation +- position calculation +- nested position-object allocation + +The same kind of scrutiny should apply to eager string materialization. If a +hot path starts producing copied text eagerly where ranges used to be enough, +that is a real performance regression candidate. + +## What this means for future optimization + +If a proposed optimization does not preserve source fidelity and structural +contracts, it is the wrong optimization. + +If it does preserve them, then the next question is simple: + +```text +does this remove repeated work on real inputs? +``` + +That is the bar the performance work should keep using. \ No newline at end of file diff --git a/docs/architecture/sessions-and-streaming.md b/docs/architecture/sessions-and-streaming.md new file mode 100644 index 0000000..26838f0 --- /dev/null +++ b/docs/architecture/sessions-and-streaming.md @@ -0,0 +1,117 @@ +# Sessions and Streaming + +The stateless parser functions are the core API. + +The session API exists for the cases where a caller wants to ask more than one +question about the same source, or where the source changes over time. + +## What a session is + +A session is a cached wrapper around the same parser pipeline. + +```text +createSession(source) + -> cached outline events + -> cached full events + -> cached tree materializations + -> later streaming and incremental edit support +``` + +It is not a different parser. + +It is still the same range-first parser. The session mainly caches views over +the same source and, later, over changes to that source. + +## What sessions buy you + +Without a session, each call starts again from the original source. + +With a session, one caller can ask: + +- `outline()` +- `events()` +- `parse()` +- `parseWithDiagnostics()` +- `parseStrictWithDiagnostics()` + +and reuse earlier work where that reuse is valid. + +That matters most for editors, live previews, and repeated structural queries. + +That reuse works because the underlying outputs stay anchored to the same +source ranges and UTF-16 offsets. If the parser eagerly rewrote text into new +strings at every step, much less of that work would compose cleanly. + +## Cache lanes follow the same cost model + +The session should not force every caller onto the expensive diagnostics path. + +That is why its caches follow the same split as the stateless APIs: + +```text +diagnostics off + -> cheapest event and tree access + +diagnostics on + -> diagnostics-preserving event and tree access +``` + +If a caller only wants the cheap lane, the session should not allocate or keep +diagnostics-enabled state unless some other call already paid for it. + +The same rule applies to text materialization. Sessions should prefer caching +range-first parser products over eagerly caching many duplicate string views of +the same source. + +## Streaming and the stability frontier + +Streaming input creates one extra problem: the parser needs to distinguish what +is already stable from what may still change. + +```text +stable prefix | provisional tail +``` + +The boundary between them is the stability frontier. + +Before the frontier, events are safe to treat as committed. +After the frontier, later source may still close an open construct and change +the parse. + +That is the core idea behind future streaming support such as: + +- `session.write(chunk)` +- stable-event draining +- provisional-tail replacement + +## Incremental editing + +Incremental parsing is the later session feature for edited text, not just +append-only streaming. + +The parser direction here is: + +- find the smallest safe region to reparse +- reuse the rest of the prior work +- return a `PositionMap` so cursors, selections, and comment-like anchors can + be translated across edits + +That only works if positions and source spans stay authoritative enough to map +old work onto new source confidently. + +That is how the parser supports editor-like workflows without turning the core +pipeline into an editor framework. + +## Why this belongs outside the main architecture overview + +Sessions, streaming, and incremental reparsing are important, but they are not +the first thing most readers need. + +Most people first need to know: + +- what the parser outputs are +- which tree lane to choose +- how the pipeline roughly works + +This note exists so those live-use-case details stop crowding the main +architecture entry point. \ No newline at end of file diff --git a/docs/architecture/utility-first.md b/docs/architecture/utility-first.md new file mode 100644 index 0000000..333bc2b --- /dev/null +++ b/docs/architecture/utility-first.md @@ -0,0 +1,101 @@ +# Utility-First Design + +This parser is utility-first. + +That means it tries to expose stable primitives that other tools can build on, +instead of turning the core parser into a giant plugin host. + +## What utility-first means here + +The package exposes building blocks such as: + +- `TextSource` +- tokens +- events +- tree builders +- tree utilities +- session-oriented wrappers + +Those are the surfaces other tools should compose. + +Just as importantly, those surfaces preserve the parser's native data model: +source-backed ranges first, richer interpretations second. + +## Why the parser does not lead with hooks + +The repo does not currently treat deep parser hooks as the main extension +story. + +That is intentional. + +Deep hooks freeze a lot of internal design too early: + +- ambiguity handling +- malformed-input boundaries +- hot-path control flow +- recovery and materialization policy + +If those internals are still evolving, a broad hook system creates long-term +API cost very quickly. + +## What to do instead + +The preferred extension model is to build on the parser's public primitives. + +In practice, that usually means one of these: + +1. consume tokens and emit your own higher-level events +2. consume the event stream and emit your own derived events +3. build a focused mini-parser for one domain-specific construct and merge its + output with the main parser's results +4. walk the tree and add your own interpretation layer later + +This keeps the core parser small and predictable while still giving consumers a +real extension surface. + +It also lets downstream tools choose when they actually need strings. Many +extensions can do their work directly from tokens, source ranges, event slices, +or tree nodes that still point back into the original source. + +## The recommended mental model + +Think of the package less like this: + +```text +one parser you patch from the inside +``` + +and more like this: + +```text +shared parser primitives + -> your domain-specific logic + -> your own events, trees, or transforms +``` + +If you need special behavior for one feature, it is often better to write a +small focused parser that consumes the source ranges or event slices you care +about than to ask the core parser to grow a new generic hook point. + +That recommendation is partly architectural and partly practical: range-first +primitives are cheaper to compose than hook systems that force the core parser +to materialize and expose every intermediate detail eagerly. + +## What stays public versus internal + +Public and intended for downstream use: + +- source abstractions such as `TextSource` +- token, event, and AST interfaces and unions +- builder functions and type guards +- parser stage entry points such as `tokenize()`, `blockEvents()`, and + `inlineEvents()` + +Still internal: + +- scanner-local context records +- one-stage matcher result shapes +- low-level continuation helpers whose contracts are still evolving + +That boundary is what keeps the package utility-first without making every +local implementation detail part of the public API. \ No newline at end of file diff --git a/docs/diagnostics-first-redesign.md b/docs/diagnostics-first-redesign.md new file mode 100644 index 0000000..add43fc --- /dev/null +++ b/docs/diagnostics-first-redesign.md @@ -0,0 +1,550 @@ +# Diagnostics-First Parser Redesign + +Status: the first public realignment has landed. The parser now exposes +`parseStrictWithDiagnostics()` and `buildTreeStrict()`, and event-stream diagnostics are +opt-in by default. The sections below still capture the reasoning and the +remaining architectural direction behind that shift. + +If you want the smaller architecture-facing explanation first, start with: + +- [docs/architecture/diagnostics-first.md](./architecture/diagnostics-first.md) +- [docs/architecture/choosing-a-parser-result.md](./architecture/choosing-a-parser-result.md) +- [docs/architecture/malformed-input.md](./architecture/malformed-input.md) +- [docs/architecture/api-direction.md](./architecture/api-direction.md) +- [docs/architecture/utility-first.md](./architecture/utility-first.md) +- [docs/architecture/performance.md](./architecture/performance.md) +- [docs/architecture/parser-contracts.md](./architecture/parser-contracts.md) + +This note audits the current parser surface against the intended +diagnostics-first model and proposes a concrete redesign. + +Use this note for the deeper design argument, the migration pressure, the +remaining rollout questions, and the open questions. For the shorter +architecture explanation, use +[docs/architecture/diagnostics-first.md](./architecture/diagnostics-first.md). + +The goal is not to remove tolerant defaults. The goal is to separate three +different ideas that have drifted together in the current API: + +- detecting malformed input +- continuing parsing without throwing +- materializing one particular recovered tree shape + +That separation matters because the parser should be explicit about its +defaults without compelling every consumer to accept the parser's own recovery +materialization. + +## Problem + +The current public story treats recovery as a parser-owned output lane. + +```text +current public framing + +source + ├─► parse() -> default tree + ├─► parseWithDiagnostics() -> default tree + diagnostics + └─► parseWithRecovery() -> default tree + diagnostics + recovery summary +``` + +The intended model is: + +- diagnostics are the primitive +- the parser may continue internally so it does not throw +- recovery materialization is an optional consumer response to diagnostics +- downstream consumers may expand parser diagnostics into their own + domain-specific diagnostics + +In other words, `never throw` means the parser keeps going and preserves the +problem. It does not mean the parser must publish one canonical recovered tree +shape as the meaning of malformed input. + +The current code still has useful pieces of that original design. In +[events.ts](events.ts), error events are optional, structured, and already +described as consumer-facing signals that can be logged, surfaced, ignored, or +expanded. The drift happens higher in the stack, where the parser and tree +builder elevate recovery materialization into first-class parser lanes. + +## Goals + +- Keep the parser never-throw contract. +- Stay close to HTML-like parsing with explicit commitment points and tolerant + defaults. +- Make diagnostics the primary public primitive for malformed input. +- Keep diagnostic emission optional so consumers that do not want diagnostics + do not pay for them. +- Surface recovery materialization options explicitly. +- Avoid compelling consumers to accept a recovered tree if they only want + diagnostics, source-backed text, or their own downstream recovery logic. +- Preserve the ability for downstream tools such as renderers and editors to + emit their own diagnostics on top of parser diagnostics. + +## Non-goals + +- Remove all internal continuation heuristics from the parser. +- Make malformed input fatal. +- Remove tolerant default behavior for consumers that do want it. +- Solve edit-stable diagnostic anchors in this redesign. + +## Current Drift + +### The event model is mostly aligned + +The event layer in [events.ts](events.ts) is already close to the intended +shape: + +- diagnostics are optional `error` events +- diagnostic codes are stable parser-owned facts +- diagnostic metadata is open enough for downstream consumers +- docs already describe consumer choice as the point of those diagnostics + +This is the strongest part of the current design and should remain the base +primitive. + +### The parser API still makes materialization feel more parser-owned than it should + +The orchestration API in [parse.ts](parse.ts) currently frames the parser as +several top-level tree lanes: + +- `parse()` returns the default HTML-like tree +- `parseWithDiagnostics()` returns that same default tree plus diagnostics +- `parseStrictWithDiagnostics()` returns the source-strict tree plus diagnostics +- `parseWithRecovery()` returns the default tree plus diagnostics and `recovered` + +That is better than the older framing, because diagnostics are now opt-in by +default and `parseStrictWithDiagnostics()` names the conservative lane more honestly. But it +still makes materialization feel like a parser-owned truth instead of a +consumer policy layered over the same syntax findings. + +### The tree builder currently owns recovery-shape policy + +[tree_builder.ts](tree_builder.ts) currently exposes `TreeBuildMode` as +`'strict' | 'loose'` and documents those modes as recovery-shape policies. + +That creates two problems: + +- tree materialization policy is presented as if it were part of parsing truth +- diagnostics and recovered tree shape are coupled in one API family + +The addition of `buildTreeWithLooseDiagnostics()` helps operationally because +it separates `recovered` from `diagnostics`, but it still keeps the public +story centered on parser-owned recovery shape. + +### The current docs teach the drift as architecture + +[docs/architecture.md](docs/architecture.md) and [readme.md](readme.md) +currently describe `strict` and `loose` as top-level parser choices and teach +the tree lanes as the main way to think about malformed input. + +That is the opposite of the intended mental model. It teaches readers that the +parser owns recovery semantics, when the intended design is that the parser +owns diagnostics and continuation, and consumers own recovery responses. + +### Tests currently lock in parser-owned recovery semantics + +[parse_test.ts](parse_test.ts) and [session_test.ts](session_test.ts) assert +that: + +- `parseWithDiagnostics()` and `parseWithRecovery()` are separate wrappers even + though they now share the same default tree shape +- strict versus loose tree shape is a parser-level distinction +- recovery-specific wrapper nodes are part of the parser's promised result + +Those tests are useful because they show the exact behavioral drift. They will +need to move as the API is realigned. + +## Design Distinction To Restore + +The core distinction should be: + +```text +diagnostic = factual parser finding +continuation = minimal internal behavior that lets parsing proceed +materialization = consumer-visible choice about how to represent malformed input +``` + +That gives a cleaner mental model: + +```text +source + ├─► parser detects malformed input + ├─► parser records a diagnostic when requested + ├─► parser continues internally so it never throws + └─► consumer chooses whether to: + - ignore the diagnostic + - surface the diagnostic + - materialize a tolerant tree + - materialize a conservative tree + - emit richer downstream diagnostics +``` + +The parser still needs continuation heuristics. A block parser cannot keep +streaming without making some local choice when a table never closes or an +event stream ends with open frames. But those choices should be treated as +survivability mechanics, not as the only public recovery meaning. + +## Proposal + +### 1. Make diagnostics the primary malformed-input contract + +Keep parser diagnostics as structured event-layer facts. + +The parser should continue to emit stable `DiagnosticCode` values and optional +details, but the docs should stop describing those diagnostics as evidence of a +parser-owned recovery lane. They are parser findings that consumers can act on +or translate. + +### 2. Rename the internal concept from recovery to continuation where possible + +Inside parser implementation docs and comments, prefer `continuation` for the +minimum logic required to keep scanning or keep the event stream well-formed. + +Use `recovery` for consumer-visible responses to diagnostics or for explicit +materialization helpers that are intentionally opt-in. + +This is especially important in: + +- [block_parser.ts](block_parser.ts) +- [inline_parser.ts](inline_parser.ts) +- [tree_builder.ts](tree_builder.ts) + +### 3. Move tree shape choices out of the parser-success narrative + +Tree shape should be described as a materialization policy, not as the parser's +acceptance or recovery mode. + +A clearer conceptual split is: + +- parser options decide whether diagnostics are emitted +- materializer options decide how malformed regions are represented + +That means the parser surface should stop teaching `strict` and `loose` as the +main mental model for malformed input. + +### 4. Keep diagnostic emission optional and explicit + +This part is already directionally correct in the implementation and should be +preserved. + +If a caller says they do not want diagnostics, block and inline parsers should +not emit diagnostic events at all. + +That cost model should be made explicit in the public API: + +```text +diagnostics off + -> no diagnostic events emitted + -> no diagnostic collection + -> cheapest parse lane + +diagnostics on + -> diagnostic events emitted + -> downstream recovery-aware materializers may consume them + -> caller accepts the added allocation and processing cost +``` + +### 5. Reframe public tree APIs around materialization policy + +Replace the current recovery-centered API story with a diagnostics-first, +materialization-second story. + +One concrete direction: + +```ts +interface ParseOptions { + readonly diagnostics?: boolean; + readonly materialization?: 'default-html-like' | 'source-strict'; +} + +interface ParseOutput { + readonly tree: WikistRoot; + readonly diagnostics?: readonly ParseDiagnostic[]; +} +``` + +With convenience wrappers if desired: + +- `parse(source)` + default HTML-like materialization, no diagnostics +- `parseWithDiagnostics(source)` + default HTML-like materialization plus diagnostics +- `parseStrictWithDiagnostics(source)` + source-strict materialization plus diagnostics by default + +The important change is not the exact names. The important change is that the +API no longer teaches "diagnostics lane versus recovery lane" as if those were +peer parser truths. Instead it teaches: + +- do you want diagnostics? +- if yes, how do you want malformed regions materialized? + +For that public-facing option shape, `diagnostics` should be the public name. +It is shorter, and it describes the caller's choice instead of the +implementation detail of whether diagnostic events are being emitted +internally. + +### 6. Treat explicit recovery materialization as an opt-in helper layer + +If the package wants to preserve strong support for tolerant consumers, keep +that support as explicit helpers rather than as the parser's main malformed- +input identity. + +For example: + +- `buildTree(events, { source })` + default HTML-like materialization +- `buildTreeWithDiagnostics(events, { source })` + same materialization plus anchored diagnostics +- `materializeDiagnostics(events, { source, policy: 'source-strict' })` + conservative materialization for linting or editor diagnostics +- `materializeDiagnostics(events, { source, policy: 'default-html-like' })` + tolerant materialization for rendering + +This keeps recovery materialization surfaced, explicit, and optional. + +### 7. Add materialization hints to diagnostics instead of baking one response into the tree API + +If parser diagnostics need to help consumers choose a response, add stable +metadata that describes the kind of malformed region without forcing one public +repair. + +For example, the diagnostic payload could eventually carry fields such as: + +- `code` +- `source` +- `details` +- `continuation_kind` +- `suggested_materializations` + +That would let a renderer, editor, or linter make an informed choice without +requiring the parser to publish one canonical recovered tree for every case. + +This should stay conservative. The parser should expose facts and hints, not a +large strategy engine. + +### 8. Leave room for a true `analyze()` lane + +There is one stronger diagnostics-first option that the current public API does +not fully provide yet. + +Today the parser offers: + +- a cheap tree with no preserved diagnostics +- the default tolerant tree with diagnostics +- a conservative tree with diagnostics + +What it does not fully offer yet is this: + +```text +diagnostics and possible recoveries are exposed +but the caller chooses the final materialization later +``` + +That would be more diagnostics-first than `parseStrictWithDiagnostics()`. `parseStrictWithDiagnostics()` +still chooses a final conservative tree policy on the caller's behalf. A true +`analyze()` lane would expose findings first and make the final repair policy +more explicitly caller-owned. + +The practical shape matters here. The best public answer is probably not just a +bare one-shot event generator. The parser is already event-stream-first, but a +control-heavy caller often needs more than one pass over the same information. + +That caller may need to: + +- inspect diagnostics before choosing a tree policy +- compare more than one candidate repair path +- materialize both tolerant and conservative trees from the same parse +- cache findings inside a session or incremental workflow + +That points toward a replayable findings utility built on the event stream. +Conceptually: + +```text +source + -> analyze() + -> findings + - replayable events + - diagnostics + - recovery data + -> materialize(findings, tolerant) + -> materialize(findings, conservative) +``` + +This should stay an open design question until the package has a clearer model +for what a "possible recovery" surface actually looks like and how much of that +surface belongs to the parser instead of later consumers. + +## Trust rules and practical contracts + +The practical trust rules now live in the smaller architecture notes so this +design note does not have to reteach the same ground every time. + +- [docs/architecture/parser-contracts.md](./architecture/parser-contracts.md) + covers the invariants the parser should preserve. +- [docs/architecture/malformed-input.md](./architecture/malformed-input.md) + covers commitment points and malformed-input behavior. +- [docs/architecture/diagnostic-anchors.md](./architecture/diagnostic-anchors.md) + covers how diagnostics point back into the final tree. + +## How To Tell Whether The Block/Inline Split Is A Problem + +The existence of two parser stages is not evidence of a design problem. The +useful question is whether the handoff causes correctness drift or measurable +extra work. + +Check it in this order: + +1. Correctness contract: + `outlineEvents()`, `events()`, and `parse()` should agree on block + structure for the same input in the default tolerant lane. +2. Handoff behavior: + merged block text groups should not trim content, invent boundaries, or + force the inline parser to reconstruct information the block parser already + knew. +3. Cost profile: + measure text-event counts, merged text-group counts, inline candidate + scans, allocations, and total throughput on prose-heavy, syntax-heavy, and + malformed corpora. +4. Change threshold: + only collapse the stages or redesign the boundary if the split causes a + real correctness bug or a benchmarked throughput or allocation regression. + +The useful comparison is not "two stages versus one stage" in the abstract. +It is "does this handoff preserve the contracts while reducing repeated work?" + +```text +bad split + block parser fragments text + -> inline parser has to reconstruct the same logical group + -> correctness or cost regresses + +good split + block parser emits one accurate contiguous prose range + -> inline parser scans once for real inline openers + -> contracts stay intact and work drops +``` + +That is why the next verification step should be contract tests plus focused +benchmarks, not a premature rewrite into one monolithic parser. + +## Primitive-first extension model + +The longer architecture-facing explanation now lives in +[docs/architecture/utility-first.md](./architecture/utility-first.md). + +The short version here is still important: the redesign works better if the +parser leads with stable primitives and downstream composition instead of a deep +hook surface that freezes hot-path behavior too early. + +## Ambiguity hotspots worth hardening next + +The narrow handoff and bare-URL follow-up now lives in +[docs/handoff-and-bare-url-notes.md](./handoff-and-bare-url-notes.md). + +That keeps this redesign note focused on the tree and diagnostics model instead +of turning it into a second all-purpose architecture file. + +## HTML-Like Parsing Guidance + +Staying close to HTML-like parsing still fits this redesign. + +The parser should continue to use explicit commitment points. The existing +HTML-like tag rule in [inline_parser.ts](inline_parser.ts) is a good example: + +- before the opener reaches `>`, do not commit a real tag node +- after the opener reaches `>`, the opener is structurally real +- if the close tag never arrives, emit a diagnostic and let materializers + choose how much tolerant structure to preserve + +That is a good default because it is explicit and predictable. What should +change is not the commitment rule. What should change is the public framing of +who owns the final recovery materialization. + +## Concrete API direction + +The future public-surface cleanup now lives in +[docs/architecture/api-direction.md](./architecture/api-direction.md). + +That keeps this redesign note focused on the larger architectural problem while +still preserving one place for the future wrapper and option-shape discussion. + +## Performance Model + +This redesign keeps the fast path clear. + +```text +cheap lane + diagnostics off + -> tokenizer + -> block parser without diagnostic events + -> inline parser without diagnostic events + -> default materialization if requested + +diagnostics lane + diagnostics on + -> tokenizer + -> block parser with diagnostic events + -> inline parser with diagnostic events + -> diagnostics-preserving materialization if requested +``` + +This matches the intended trade-off: + +- consumers that do not want diagnostics do not pay for them +- consumers that want diagnostics and recovery-aware handling accept the cost + +That cost model should be explicit in docs, code comments, and session cache +structure. + +## Rollout Plan + +1. Rewrite docs so diagnostics are the primary malformed-input concept and + tree shape is described as materialization policy. +2. Deprecate recovery-centered names where they teach the wrong model, + especially `parseWithRecovery()` as the main public malformed-input lane. +3. Introduce a findings utility and tree APIs that separate diagnostic emission + from materialization policy. +4. Move tests from "parser owns recovery tree semantics" to + "materializer policy changes representation of the same parser findings". +5. Keep compatibility wrappers temporarily if migration cost matters. +6. Add contract tests for cross-mode block consistency, commitment-point + behavior, and the default-versus-source-strict trust boundaries. +7. Add focused benchmarks or counters around the block-to-inline handoff before + reconsidering the two-stage design. + +This section stays here for now because it is still part of the redesign +discussion, not yet a settled implementation plan. + +## Open Questions + +- Should the default `events()` lane include diagnostics, or should the + cheapest no-diagnostics lane be the default and diagnostics require explicit + opt-in? +- Should the package keep a tolerant convenience wrapper equivalent to the + current `parseWithRecovery()`, or should that become a more clearly named + materialization helper? +- Should the `analyze()` lane expose replayable arrays, a session-backed view, or + another reusable iterable shape? +- How much structured response metadata should diagnostics carry before the + parser starts looking like a strategy engine? +- Should tree-stage findings such as mismatched exits remain public parser + diagnostics, or be treated as internal materializer diagnostics unless the + caller explicitly requests low-level stream integrity facts? +- How narrow should the core bare-URL recognizer stay before profiles or + downstream consumers take over richer URL or IRI handling? + +## Summary + +The current implementation already has the right low-level foundation: +optional structured diagnostics and explicit cost gates for emitting them. + +The drift is mainly in the public framing and the tree APIs. They currently +teach recovery materialization as a parser-owned lane. + +The redesign should restore this rule: + +- the parser owns diagnostics and continuation +- consumers own whether and how malformed input is materially recovered + +That keeps the parser explicit, tolerant, HTML-like, and performant without +forcing downstream consumers into one recovery worldview. \ No newline at end of file diff --git a/docs/examples.md b/docs/examples.md new file mode 100644 index 0000000..25e0d06 --- /dev/null +++ b/docs/examples.md @@ -0,0 +1,163 @@ +# Examples + +These examples show the most common ways to use the parser without reading the +deeper architecture notes first. + +They start with the simplest paths and then move toward more inspection-heavy +flows. + +## Build a tree for normal document work + +Use `parse()` when you want a tree and you do not need diagnostics. + +```ts +import { parse } from '@okikio/wikitext'; + +const tree = parse('== Heading ==\n\nA paragraph with [[Main Page|a link]].'); +console.log(tree.type); // 'root' +``` + +This is the cheapest tree path. It is a good fit when the source is mostly +ordinary wikitext and you do not need the parser's explanation of malformed +regions. + +## Inspect headings without building a tree + +Use `outlineEvents()` when you only care about block structure such as headings +or lists. + +```ts +import { outlineEvents } from '@okikio/wikitext'; + +const headings = []; + +for (const event of outlineEvents('== One ==\n\nText\n\n=== Two ===')) { + if (event.kind === 'enter' && event.node_type === 'heading') { + headings.push(event); + } +} +``` + +## Keep diagnostics while preserving the default tree + +Use `parseWithDiagnostics()` when you want the parser's default tolerant tree +and also want to know where malformed input was detected. + +```ts +import { parseWithDiagnostics } from '@okikio/wikitext'; + +const result = parseWithDiagnostics('Paragraph with note'); + +console.log(result.tree.children[0]?.type); +console.log(result.diagnostics.map((diagnostic) => diagnostic.code)); +``` + +This is the best current fit for the "recover on my behalf, but tell me what +you did" style of usage. + +## Ask for the conservative tree instead + +Use `parseStrictWithDiagnostics()` when diagnostics matter and you want the final tree to be +more conservative about malformed committed structure. + +```ts +import { parseStrictWithDiagnostics } from '@okikio/wikitext'; + +const result = parseStrictWithDiagnostics('Paragraph with note'); + +console.log(result.tree.children[0]?.type); +console.log(result.diagnostics.length > 0); +``` + +This is useful for inspection or linting flows that do not want the final tree +to keep as much repaired structure. + +## Compare the current tree lanes on one malformed input + +The easiest way to understand the tree differences is to run the same input +through more than one lane. + +```ts +import { parse, parseStrictWithDiagnostics, parseWithDiagnostics } from '@okikio/wikitext'; + +const input = 'Paragraph with note'; + +const fast_tree = parse(input); +const recovery_tree = parseWithDiagnostics(input); +const conservative_tree = parseStrictWithDiagnostics(input); + +console.log(fast_tree.children[0]?.type); +console.log(recovery_tree.tree.children[0]?.type); +console.log(conservative_tree.tree.children[0]?.type); +console.log(recovery_tree.diagnostics.map((diagnostic) => diagnostic.code)); +``` + +This kind of comparison is useful when you are deciding whether your tool wants +speed first, tolerant structure first, or conservative diagnostics first. + +## Resolve a diagnostic back to the tree + +Diagnostics live beside the tree, but you can still resolve them back to the +nearest materialized node. + +```ts +import { locateDiagnostic, parseWithDiagnostics } from '@okikio/wikitext'; + +const result = parseWithDiagnostics('{|\n| Cell'); +const location = locateDiagnostic(result.tree, result.diagnostics[0]); + +console.log(location?.node.type); +console.log(location?.parent?.type); +console.log(location?.index); +``` + +This is useful for editor tooling, lint output, or any workflow that wants to +connect a parser finding back to a concrete part of the final tree. + +## Reuse one source through a session + +Use a session when you want to ask more than one question about the same input. + +```ts +import { createSession } from '@okikio/wikitext'; + +const session = createSession('== Heading ==\n\nParagraph with [[Main Page]].'); + +const outline = Array.from(session.outline()); +const full_events = Array.from(session.events()); +const tree = session.parse(); + +console.log(outline.length > 0); +console.log(full_events.length > 0); +console.log(tree.type); +``` + +This avoids redoing the same work from scratch each time you ask a different +question about one source string. + +## Build a focused parser on top of events + +This package is utility-first. If you need domain-specific behavior, prefer +building on the public primitives instead of expecting deep hooks in the core +parser. + +```ts +import { events } from '@okikio/wikitext'; + +function collectLinks(source: string): string[] { + const links: string[] = []; + + for (const event of events(source)) { + if (event.kind === 'enter' && event.node_type === 'wikilink') { + links.push(String(event.props.target)); + } + } + + return links; +} + +console.log(collectLinks('Visit [[Main Page]] and [[Help:Contents]].')); +``` + +If you want the reasoning behind that design, see +[docs/architecture/utility-first.md](./architecture/utility-first.md). \ No newline at end of file diff --git a/docs/future-direction.md b/docs/future-direction.md new file mode 100644 index 0000000..0a338d1 --- /dev/null +++ b/docs/future-direction.md @@ -0,0 +1,168 @@ +# Future Direction + +This note captures the broader direction behind the current parser work. + +The parser is the current product focus, but it is not the full long-term +product story. Writing that down matters because parser architecture choices +look different when the parser is understood as a base layer rather than the +final user-facing system. + +## What this repo is building now + +Right now this repository is building a standards-aligned wikitext source +parser. + +Its job is to turn source text into reusable parser primitives: + +- tokens +- events +- trees +- diagnostics +- session-oriented access patterns later on + +Those primitives need to be good enough for extraction, rendering, +transformation, inspection, and editor-facing tooling. + +## What comes next + +Two higher-level libraries should sit above the parser once the core surface is +ready enough. + +```text +@okikio/wikitext + ├─► @okikio/wikidoc + └─► @okikio/wiki-extract +``` + +### `@okikio/wikidoc` + +This layer should focus on document-oriented work built on top of parser +primitives. + +That likely includes: + +- higher-level document transforms +- structure-aware editing helpers +- richer composition around trees, diagnostics, and sessions +- bridges toward editor and CMS-style workflows + +### `@okikio/wiki-extract` + +This layer should focus on pulling useful structured information out of parsed +content. + +That likely includes: + +- template and infobox extraction +- category and link extraction +- reference extraction +- selective structure queries and summaries + +The key point is that extraction and document composition should not distort the +parser core. They should consume parser primitives instead. + +## Longer-term direction + +Longer term, the same primitives may expand into a broader profile-driven +document engine. + +That direction is bigger than "a better wikitext parser." It points toward a +system that can support: + +- wikitext as the proving ground +- additional markup or rich-text families later +- CMS-style structured blocks +- editor-facing session and incremental workflows +- local-first or collaboration-aware tooling +- LLM-oriented document transforms built on explicit structure rather than raw + text guessing + +The right way to read this is as a direction of travel, not as a promise that +all of those layers belong in this package soon. + +## Why this matters to the parser now + +This longer horizon changes what "good parser architecture" means. + +The parser should optimize for reusable primitives, not just one rendering +pipeline. + +That means: + +- events need to stay first-class +- diagnostics need to stay factual and reusable +- offsets need to stay authoritative +- extension boundaries need to stay narrow and composable +- session and incremental APIs need to build on the same core primitives + +It also means the parser should avoid baking in product-specific behavior that +really belongs one layer up. + +## Non-goals for the core parser + +The parser core should not silently become: + +- a full MediaWiki runtime +- a CMS block model +- an editor framework +- a collaboration engine +- an extraction product + +Those may all become consumers or adjacent layers. They should not become the +hot loop. + +## Design pressure this adds today + +This broader direction puts useful pressure on current decisions. + +### Diagnostics versus materialization + +Different higher-level consumers need different responses to malformed input. +An extractor, a renderer, and an editor may all want the same parser finding +but a different final representation. + +That is why diagnostics should stay parser facts and final tree shape should +stay a materialization choice. + +### Primitive-first extension model + +If the long-term system grows into multiple products, deep parser hooks become +too expensive too early. + +The core should prefer: + +- stable tokens, events, and trees +- narrow feature gates +- explicit tag or profile handlers +- downstream enrichment passes + +### Profile-driven behavior + +The broader engine direction does not require one giant parser that tries to do +everything at once. + +It points instead toward profiles layered over shared primitives: + +```text +shared parser primitives + ├─► syntax profile + ├─► mediawiki-compatible profile + └─► later document-oriented profiles +``` + +That keeps the core maintainable while leaving room for stricter or more +product-specific behavior later. + +## How to use this note + +Use this note when judging parser design trade-offs. + +If two parser designs look equally good for today's tests, prefer the one that: + +- keeps primitives reusable +- keeps product policy out of the hot path +- keeps future extraction and document tooling possible +- preserves optionality for profile-driven evolution + +If a design only makes sense for one consumer and hardens that assumption into +the parser core, it is probably the wrong long-term move. \ No newline at end of file diff --git a/docs/handoff-and-bare-url-notes.md b/docs/handoff-and-bare-url-notes.md new file mode 100644 index 0000000..92abaed --- /dev/null +++ b/docs/handoff-and-bare-url-notes.md @@ -0,0 +1,165 @@ +# Block/Inline Handoff and Bare-URL Notes + +This note records two narrow follow-ups from the diagnostics-first redesign: + +1. how to tell whether the block-to-inline split is helping or hurting +2. which bare-URL rules belong in the core parser instead of downstream tools + +The goal is to keep the work measurable and local. Neither topic needs a broad +parser rewrite to become clearer. + +## What this note is for + +The parser already has a strong default shape: + +```text +TextSource -> tokenizer -> blockEvents() -> inlineEvents() -> tree consumers +``` + +That shape is worth keeping as long as the handoff preserves structure and +reduces repeated work. The right next step is not "merge the stages because two +stages feel suspicious." The right step is to measure the handoff directly and +to tighten the ambiguity rules that still look too ad hoc. + +Bare URLs are the clearest current example of that second category. + +## How to evaluate the block-to-inline handoff + +There are two separate questions. + +First, does the handoff preserve the parser contracts? + +- `outlineEvents()` and `events()` must agree on block structure. +- `parse()` must preserve that same block structure in the default lane. +- contiguous prose ranges must not lose spaces, punctuation, or boundaries + just because they were merged before inline parsing. + +Second, does the handoff actually save work? + +The useful benchmark comparison is: + +```text +fragmented handoff + many neighboring text events + -> inline parser merges them before scanning + +merged handoff + one larger contiguous text event + -> inline parser scans once +``` + +If merged handoff is not faster, or if it causes structural drift, then the +boundary needs work. If it is faster and keeps the contracts intact, then the +split is doing its job. + +That is why the benchmark work here isolates merged versus fragmented text +groups for the same input instead of comparing unrelated documents. + +## What belongs in the core bare-URL matcher + +The core parser does need bare-URL recognition, but it should stay narrow and +cheap. + +The matcher should own only the rules needed to answer this parser question: + +```text +does this source range clearly contain a link-shaped URL token here? +``` + +That means the core matcher should handle: + +- accepted schemes in the core lane +- start boundaries, so URLs do not begin in the middle of ordinary words +- end boundaries, so trailing prose punctuation does not become part of the URL +- lightweight balancing rules for common prose wrappers such as a trailing `)` + +That does not mean the core parser should do full URL normalization, +base-URL-aware resolution, or browser-style canonicalization. Those jobs belong +to downstream consumers. + +## Current acceptance matrix + +The current matcher keeps two different promises. + +- Explicit bracketed external-link syntax is broad. If the author wrote + `[scheme:payload Label]`, the parser should usually trust that they meant a + link and leave scheme-policy decisions to downstream consumers. +- Bare autolinks in prose are narrower. They should catch the common cases and + a strong set of opaque URI shapes, but they should not overmatch ordinary + colon prose. + +That gives this working matrix: + +| Case | Bare prose | Bracketed explicit syntax | Reason | +| --- | --- | --- | --- | +| `https://example.com` | accept | accept | authority form is high-confidence | +| `file:///Users/example/report.txt` | accept | accept | authority form is high-confidence | +| `foo+bar://example.service/path` | accept | accept | authority form is high-confidence | +| `mailto:editor@example.org` | accept | accept | `@` is strong opaque-URI evidence | +| `urn:isbn:0451450523` | accept | accept | repeated `:` plus digits is strong opaque-URI evidence | +| `data:text/plain,hello` | accept | accept | `/` and `,` make the payload structurally URI-like | +| `tel:+12025550123` | accept | accept | leading `+` plus digits is strong opaque-URI evidence | +| `magnet:?xt=urn:btih:...` | accept | accept | `?` and `=` make the payload structurally URI-like | +| `longcustomscheme:alpha` | reject | accept | too little structure for safe bare autolinking | +| `note:abc` | reject | accept | reads like ordinary colon prose in bare text | +| `chapter:one` | reject | accept | reads like ordinary colon prose in bare text | + +The general rule is simple: if a URI is not distinct enough to be recognized +confidently in bare prose, it should stay text and let an end consumer extend +that policy if they care about a niche scheme. + +The stable low-level boundary rules are still: + +- a bare URL must start at a word boundary, not in the middle of an ASCII word +- the scan stops before whitespace, quotes, `]`, `<`, or `>` +- trailing prose punctuation such as `.`, `,`, `;`, `:`, `!`, and `?` is + trimmed from the URL +- a trailing `)` is trimmed only when it is unmatched within the scanned URL + +That gives a practical parser rule: + +```text +Visit https://example.com. + ^^^^^^^^^^^^^^^^^^^ external-link + . plain text +``` + +and: + +```text +(https://example.com/path(test)) + ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ external-link + ) plain text +``` + +The parser keeps the balanced inner parentheses because they can be part of a +real URL path, but it leaves the extra outer prose wrapper behind. + +## What still stays out of scope + +The core matcher still does not try to solve every URI or IRI question. + +That is intentional. + +- RFC 3986 and RFC 3987 are useful references for boundary decisions. +- WHATWG URL behavior is useful for web-facing expectations. +- MediaWiki behavior still matters where it intentionally differs. + +But the hot path should not instantiate `URL` objects or require a base URL +just to recognize a bare link-shaped span in plain text. That would increase +allocation cost and couple parser recognition to downstream normalization. + +## What to watch next + +If future work expands bare-URL support, the next useful checks are: + +1. whether any rejected bare opaque cases deserve promotion into the core + acceptance matrix or should stay consumer-owned extensions +2. whether any IRI behavior is important enough to justify extra scan cost +3. whether MediaWiki-compatible punctuation trimming needs a more exact rule +4. whether URL-heavy prose changes the value of the current block-to-inline + merge strategy + +Those are good follow-ups only if the current tests and benchmarks show a real +gap. Until then, the better strategy is to keep the core matcher small, +predictable, and benchmarked. \ No newline at end of file diff --git a/docs/parser-architecture-comparison.md b/docs/parser-architecture-comparison.md new file mode 100644 index 0000000..a4a3151 --- /dev/null +++ b/docs/parser-architecture-comparison.md @@ -0,0 +1,150 @@ +# Parser Architecture Comparison + +This note compares the parser families that matter most to this repo and names +the practical lessons worth carrying forward. + +The goal is not to pick a winner. The goal is to avoid borrowing the wrong +mental model from the wrong ecosystem. + +## The short answer + +This repo should stay closest to an event-stream-first parser with explicit +commitment points and tolerant defaults. + +That means: + +- not AST-first by default +- not pure Markdown-style fallback-to-text +- not a full HTML-spec tree-construction clone +- not a MediaWiki rendering engine hidden inside a source parser + +The right shape is a hybrid tuned for wikitext. + +## Comparison table + +| Family | Typical strengths | Where it breaks for wikitext | Useful lesson | +| --- | --- | --- | --- | +| Compiler parsers such as SWC, Oxc, Babel | Strong contracts, explicit recovery, precise spans, clear grammar ownership | Wikitext has more legacy ambiguity and more malformed-but-still-intended input | Keep commitment points, diagnostics, and offsets explicit | +| Markdown parsers such as cmark, micromark, goldmark, Comrak | Clean block-inline layering, event pipelines, extension discipline, spec or corpus-driven test culture | Markdown often tolerates unknown syntax by flattening it to text more aggressively than wikitext should | Keep layered parsing and corpus discipline, but do not import Markdown's fallback instinct wholesale | +| HTML parsers and html5lib-style conformance suites | Forgiving parsing after commitment, explicit tokenizer states, rich error categories | HTML's insertion modes and DOM rules are too specific to copy directly into wikitext | Use commitment-driven tolerance and error taxonomy ideas | +| MediaWiki core and Parsoid | Real wikitext correctness pressure, extension boundaries, round-trip expectations | Operationally heavy and too tied to full MediaWiki semantics to use as the core architecture model | Treat them as correctness and corpus references | + +## Why compiler architecture is only partially transferable + +Compiler parsers earn their simplicity from stronger grammars. + +JavaScript, Rust, or TypeScript parsers still have ambiguity, but they usually +know the space of valid forms ahead of time. Wikitext has more legacy syntax, +profile-specific behavior, extension-driven constructs, and malformed content +that users still expect to "work well enough." + +So the lesson from compiler parsers is not "use recursive descent everywhere" +or "turn the grammar into a formal parser generator input." The lesson is to +be explicit about these boundaries: + +- when structure becomes committed +- what counts as a factual parser finding +- what offsets and ranges mean +- which behavior is parser truth versus consumer policy + +That matches the repo's diagnostics-first direction. + +## Why Markdown architecture is useful but incomplete + +Markdown parsers are closer to this repo operationally because they often split +block and inline work and they often support event or token pipelines. + +micromark is especially relevant because it treats events as the primary +interchange layer and lets AST builders sit on top. That is already a strong +fit for this repo. + +But Markdown has one dangerous influence here: many implementations can safely +say "if we cannot prove this syntax, just keep the text." Wikitext cannot +always do that, because users often write malformed markup that is still +clearly trying to be a table, tag, link, or template. + +That is why this repo should keep the pipeline lesson from Markdown, but not +the full malformed-input philosophy. + +## Why HTML-style commitment matters more for malformed input + +HTML parsing gives a better model for the repo's tolerant default lane. + +The useful pattern is: + +```text +before commitment -> keep text-backed interpretation +after commitment -> preserve structural intent +malformed continue -> emit diagnostics and keep going +strict consumer -> optionally collapse back to text later +``` + +This repo should keep that shape for tags and, where it fits, for other +constructs with meaningful commitment points. + +The important part is discipline. A forgiving parser is not a guessing parser. +It still needs explicit rules for when an opener became real. + +## Why MediaWiki and Parsoid are the correctness references + +MediaWiki core and Parsoid are where the ecosystem has already paid for years +of parser mistakes. + +They matter less because their internals are pretty and more because their test +corpora capture: + +- real editorial habits +- malformed inputs that still need usable output +- extension interactions +- round-trip expectations +- regression history + +This repo should not clone their architecture. It should borrow pressure from +their corpora and use that pressure to validate its own simpler architecture. + +## Recommended architecture stance + +This is the working recommendation. + +### Core parser + +- Keep the tokenizer, block parser, and inline parser hand-written. +- Keep events as the primary interchange layer. +- Keep offsets and positions range-first. + +### Malformed-input model + +- Treat diagnostics as parser facts. +- Treat continuation as internal survivability logic. +- Treat final tree shape as materialization policy. + +### Public API philosophy + +- Default lane keeps tolerant structure. +- Strict lane stays conservative. +- Recovery summary is optional metadata, not a separate parser truth. + +### Extension philosophy + +- Prefer primitive-first composition over deep parser plugins. +- Keep extension boundaries explicit, especially for tag-like and parser- + function-like constructs. + +## What to test because of this comparison + +Architecture choices only matter if tests make them real. + +The most important tests implied by this comparison are: + +1. Block structure stays consistent across `outlineEvents()`, `events()`, and + `parse()` in the default tolerant lane. +2. A construct that never reaches its commitment point stays text-backed in all + lanes. +3. A construct that did commit may stay structural in the default lane and + collapse in the strict lane without changing diagnostics. +4. Extension-boundary constructs such as `` and parser functions stay + covered by their own reduced corpora instead of leaking into generic tests. +5. Later round-trip work should use Parsoid-style corpus pressure instead of + synthetic only-happy-path fixtures. + +The corresponding corpus plan lives in [docs/corpus-matrix.md](./corpus-matrix.md). \ No newline at end of file -- 2.51.2