diff --git a/docs/api-reference.md b/docs/api-reference.md new file mode 100644 index 0000000..6062706 --- /dev/null +++ b/docs/api-reference.md @@ -0,0 +1,92 @@ +# API Reference + +This page is the compact reference index for the currently shipped public API. + +If you are new to the package, start with [readme.md](../readme.md) first. +That page explains what the parser is for, how to install it, and which output +shape to choose. + +The reference below assumes the same core model as the architecture docs: +parser outputs are range-first, source-backed where practical, and grounded in +UTF-16 offsets. + +This page is for lookup, not for the full design explanation. + +- for install and first use, go to [readme.md](../readme.md) +- for task-focused snippets, go to [examples.md](./examples.md) +- for parser trade-offs, malformed-input behavior, and tree-policy reasoning, + go to [architecture/README.md](./architecture/README.md) + +## Pipeline modules + +| Module | Purpose | Status | +|--------|---------|--------| +| `text_source.ts` | Abstracts the backing text store (string, rope, CRDT) while preserving UTF-16 source access | Published | +| `token.ts` | Token type constants and `Token` interface | Published | +| `events.ts` | Event stream types, constructors, type guards, and diagnostics | Published | +| `ast.ts` | Wikist AST node types, type guards, and builders | Published | +| `tokenizer.ts` | `charCodeAt` generator scanner over `TextSource`, producing offset-based tokens | Published | +| `block_parser.ts` | Block-level event emitter over source-backed token and text ranges | Published | +| `inline_parser.ts` | Inline event enrichment over source-backed text ranges | Published | +| `parse.ts` | Orchestration (tokenizer, block, inline, tree) | Published | +| `tree_builder.ts` | `buildTree(events, { source })` to `WikistRoot`, preserving source-backed positions | Published | +| `stringify.ts` | AST to wikitext (round-trip) | Not yet implemented | +| `filter.ts` | Filter and visit utilities for trees and event streams | Published | +| `session.ts` | Stateful `Session` wrapper for repeated sync access | Published | + +## Parsing and serialization + +| Function | Description | +|----------|-------------| +| `parse(input)` | Parse wikitext into a `WikistRoot` AST using the cheapest default tree lane. | +| `parseWithDiagnostics(input)` | Parse to `{ tree, diagnostics }` with the default tolerant materialization and preserved diagnostics. | +| `parseStrictWithDiagnostics(input)` | Parse to `{ tree, diagnostics }` with the conservative source-strict materialization. | +| `stringify(tree)` | Serialize a wikist tree back to wikitext. Not yet implemented. | +| `events(input)` | Full event stream (block + inline), with source-backed text ranges and optional diagnostics. | +| `outlineEvents(input)` | Block-only event stream over the same source-backed structure. | +| `parseChunked(chunks)` | Progressive completed block nodes. Not yet implemented. | + +## Low-level streams and tree building + +| Function | Description | +|----------|-------------| +| `tokenize(input)` | Raw token generator stream with offset-based token spans. | +| `tokens(input)` | Raw token generator alias for the sync public API. | +| `blockEvents(source, tokens)` | Block-level event stream from tokens. | +| `inlineEvents(source, blockEvents)` | Inline event enrichment over block events, preserving source-backed text ranges where possible. | +| `buildTree(events, { source })` | Build AST from an event iterable plus source. | +| `buildTreeWithDiagnostics(events, { source })` | Build `{ tree, diagnostics }` with the default HTML-like materialization. | +| `buildTreeStrict(events, { source })` | Build `{ tree, diagnostics }` with the conservative source-strict materialization. | +| `buildTreeWithRecovery(events, { source })` | Build `{ tree, diagnostics, recovered }` with the default tree plus an explicit recovery summary. | + +## Tree utilities + +| Function | Description | +|----------|-------------| +| `filter(tree, type)` | Get all nodes of a type recursively. | +| `visit(tree, visitor)` | Pre-order tree walker. | +| `resolveTreePath(tree, path)` | Resolve a root-relative child-index path back to a node. | +| `resolveDiagnosticAnchor(tree, anchor)` | Resolve a diagnostic anchor back to the nearest node in one materialized tree. | +| `locateDiagnostic(tree, diagnostic)` | Resolve a diagnostic's anchor back to the nearest node while the diagnostic's source span stays authoritative. | +| `createSession(source)` | Cached sync wrapper for repeated access to one source input. | + +## Foundation exports + +| Export | Module | Description | +|--------|--------|-------------| +| `TextSource` | `text_source.ts` | Interface for backing text stores with UTF-16 source access. | +| `slice()` | `text_source.ts` | Resolve an offset range to a string when a caller needs materialized text. | +| `TokenType` | `token.ts` | Constant map of all token types. | +| `Token` | `token.ts` | Token interface with type plus UTF-16 start and end offsets. | +| `isToken()` | `token.ts` | Type guard for token validation. | +| `tokenize()` | `tokenizer.ts` | Generator-based `charCodeAt` scanner. | +| `blockEvents()` | `block_parser.ts` | Block-level event generator from tokens. | +| `EnterEvent`, `ExitEvent`, `TextEvent`, `TokenEvent`, `ErrorEvent` | `events.ts` | Concrete event interfaces, including source-backed text and diagnostic events. | +| `WikitextEvent` | `events.ts` | Discriminated union of event kinds. | +| `ErrorEventOptions`, `DiagnosticSeverity` | `events.ts` | Diagnostic support types. | +| `enterEvent()`, `exitEvent()`, ... | `events.ts` | Event constructors. | +| `isEnterEvent()`, `isExitEvent()`, ... | `events.ts` | Event type guards. | +| `WikistNode`, `WikistRoot`, `WikistNodeType` | `ast.ts` | Core AST unions and aliases. | +| `WikistParent`, `WikistLiteral`, `WikistVoid` | `ast.ts` | Structural AST category unions. | +| `root()`, `heading()`, `text()`, ... | `ast.ts` | AST builder functions. | +| `isRoot()`, `isHeading()`, `isParent()`, ... | `ast.ts` | AST type guards. | \ No newline at end of file diff --git a/docs/architecture/api-direction.md b/docs/architecture/api-direction.md new file mode 100644 index 0000000..046497f --- /dev/null +++ b/docs/architecture/api-direction.md @@ -0,0 +1,232 @@ +# API Direction + +This note explains the current public surface and the likely direction of the +next cleanup. + +It is not a promise that the public API will change immediately. It is a way to +write down the target shape clearly enough that the docs, code, and tests can +move toward the same model. + +## Why this note exists + +The parser docs have been trying to explain three things at once: + +- what callers can use today +- what those current wrappers really mean +- where the public surface likely wants to go next + +Those are related, but they are not the same question. + +This note keeps the future-facing public-surface discussion separate from the +user-facing "which result do I want?" explanation in +[choosing-a-parser-result.md](./choosing-a-parser-result.md). + +It also assumes the same range-first base as the rest of the architecture docs: +diagnostics, events, and text-like tree content should stay anchored to source +spans and UTF-16 offsets for as long as possible. + +## Current public shape + +Today the main wrappers are: + +```text +parse() + -> default tree + +parseWithDiagnostics() + -> default tree + diagnostics + +parseWithRecovery() + -> default tree + diagnostics + recovery summary + +parseStrictWithDiagnostics() + -> conservative tree + diagnostics +``` + +That surface is already more honest than the older docs, but it still mixes two +different questions: + +1. do you want diagnostics? +2. how should malformed regions be materialized? + +## The direction the surface is moving toward + +The cleaner long-term public story is: + +```text +diagnostics are one choice +materialization policy is another choice +findings are the replayable bridge between them +``` + +That makes the API easier to reason about. + +Instead of teaching several top-level parser truths, the surface can teach a +smaller set of orthogonal choices: + +- tree now, or findings first +- diagnostics off or on +- package-owned or caller-owned materialization policy + +## One concrete target shape + +One reasonable target would look like this. + +```ts +interface ParseOptions { + readonly diagnostics?: boolean; + readonly materialization?: 'default-html-like' | 'source-strict'; +} + +interface ParseOutput { + readonly tree: WikistRoot; + readonly diagnostics?: readonly ParseDiagnostic[]; +} +``` + +That does not have to be the final naming. The point is the shape: + +- one option controls whether diagnostics are collected +- another option controls the final tree policy + +`diagnostics` should be the public boolean name everywhere. It reads like a +caller-facing lane choice instead of an internal implementation switch. + +Under that surface, the parser should still keep its range-first behavior. +Changing the API shape should not imply a move toward eagerly copied strings or +away from source-backed text interpretations. + +Convenience wrappers could still exist on top of that base surface. + +## Lane 3: concrete `analyze()` proposal + +There is one stronger diagnostics-first lane that the public API does not fully +offer yet. + +```text +diagnostics and recovery data are exposed +but final materialization is delayed or caller-owned +``` + +That would go beyond `parseStrictWithDiagnostics()`. `parseStrictWithDiagnostics()` still chooses a final tree +policy for the caller. A true `analyze()` lane would expose parser findings +first and let later tools decide which repairs to apply. + +That kind of lane is especially compatible with a range-first design because it +lets the parser preserve diagnostics and source spans now, then defer heavier +materialization choices until a caller actually needs them. + +One practical shape could look more like this: + +```ts +interface AnalyzeOptions { + readonly recovery?: boolean; +} + +interface ParseRecovery { + readonly kind: + | 'missing-close' + | 'unterminated-opener' + | 'eof-autoclose' + | 'mismatched-exit'; + readonly anchor: ParseDiagnosticAnchor; + readonly node_type?: WikistNodeType; + readonly policies: readonly TreeMaterializationPolicy[]; +} + +interface ParseFindings { + readonly source: TextSource; + readonly events: readonly WikitextEvent[]; + readonly diagnostics: readonly ParseDiagnostic[]; + readonly recovery?: readonly ParseRecovery[]; +} + +function analyze(source: TextSource, options?: AnalyzeOptions): ParseFindings; + +function materialize( + findings: ParseFindings, + options?: { readonly policy?: TreeMaterializationPolicy }, +): ParseOutput; +``` + +The important part is the ownership model: + +- the parser exposes replayable findings +- the caller chooses whether to materialize them at all +- if the caller does materialize them, the tree policy is explicit + +That is why a plain one-shot generator is probably not enough as the whole +public lane. The event stream should stay the core primitive, but many callers +will need to inspect diagnostics and materialize more than once without +rerunning the parse. + +The main design bet is that findings should stay narrow and factual. A useful +first public version should probably include: + +- replayable events +- diagnostics +- a small, parser-owned recovery vocabulary + +It should not start by exposing a broad plugin or callback system. + +## Lane 4: exploratory policy proposal + +There is one possible lane after that, but it should stay explicitly more +tentative. + +```text +the caller starts from parse findings +the caller chooses some recovery outcomes itself +the final tree still comes from a later materialization step +``` + +One concrete exploration could look like this: + +```ts +type RecoveryDecision = 'keep-structural' | 'collapse-to-text'; + +interface CustomMaterializeOptions { + readonly policy: 'custom'; + readonly resolve_recovery: ( + recovery: ParseRecovery, + ) => RecoveryDecision; +} + +function materialize( + findings: ParseFindings, + options: CustomMaterializeOptions, +): ParseOutput; +``` + +That would be powerful for editors, migration tools, and profile-specific +compatibility layers, but it also hardens the recovery taxonomy and +callback timing much earlier than the package may want. + +That is the core trade-off. + +- Lane 3 makes parser facts explicit. +- Lane 4 starts making parser policy negotiable. + +Lane 3 is easier to keep stable because it mostly exposes what the parser +already knows. Lane 4 is harder because it starts exposing when and how those +findings are turned into structure. + +The likely order is: + +1. make the `analyze()` lane real +2. make the package-owned materializers explicit +3. only then decide whether a caller-owned policy lane belongs in public API + +That is the kind of thing I meant earlier by an "API proposal doc": a short +design note that describes a possible future public surface before code changes +lock it in. + +## What this note is not + +This note is not: + +- a release plan +- a promise that names or wrappers will change now +- a replacement for the user-facing parser-result docs + +It is just a way to make the future public-surface direction legible. \ No newline at end of file diff --git a/docs/architecture/diagnostic-anchors.md b/docs/architecture/diagnostic-anchors.md new file mode 100644 index 0000000..064bc1a --- /dev/null +++ b/docs/architecture/diagnostic-anchors.md @@ -0,0 +1,78 @@ +# Diagnostic Anchors + +Diagnostics live beside the tree, not inside it. + +That means the parser needs a small way to point back from a diagnostic to the +relevant region of the final materialized tree. + +That design works best because diagnostics are also range-first. They start as +parser findings about positions in the original source, then carry a narrow +tree-local anchor so tooling can reconnect those findings to one materialized +view. + +## What an anchor is + +Today each diagnostic carries an `anchor`. + +That anchor is intentionally narrow. It is a tree-path snapshot to the nearest +materialized node around the diagnostic point. + +```text +root +├─ paragraph path [0] +│ └─ bold path [0, 0] +└─ table path [1] +``` + +That shape is enough for tooling to answer practical questions such as: + +- which node is nearest to this diagnostic? +- where in the final tree should I highlight or inspect? + +The important part is what the anchor does not replace. It does not replace the +source range. The diagnostic still fundamentally refers to a location in the +original text. The anchor just helps map that finding into one tree. + +## Why diagnostics are not tree nodes + +The AST is meant to represent document structure and meaning. + +Diagnostics are parser findings about malformed input or continuation events. +They are important, but they are not normal content nodes. + +Keeping them beside the tree instead of inside it makes the AST easier to use +for normal document transforms while still letting tooling recover the relevant +location when needed. + +That separation also keeps the tree from pretending diagnostics are ordinary +content. A malformed span can still be source-backed text or a recovered +structure in the tree while the diagnostic remains a parallel parser finding. + +## Current helper path + +Callers do not have to resolve anchors by hand. + +Use: + +- `resolveDiagnosticAnchor(tree, diagnostic.anchor)` +- `locateDiagnostic(tree, diagnostic)` + +These helpers turn the stored path back into a concrete node, parent, and child +index. + +## What anchors do not promise yet + +Current anchors are tree-local, not edit-stable. + +They resolve against one final materialized tree. They do not yet promise: + +- stable identity across later edits +- slot semantics across reparses +- session-backed cross-edit anchor durability + +That stronger anchor model belongs to later session and edit-tracking work. + +Until then, the safest thing a caller can trust is: + +- the diagnostic's source range is authoritative +- the anchor is a helper for one final materialized tree \ No newline at end of file diff --git a/docs/architecture/diagnostics-first.md b/docs/architecture/diagnostics-first.md new file mode 100644 index 0000000..b2c290a --- /dev/null +++ b/docs/architecture/diagnostics-first.md @@ -0,0 +1,166 @@ +# Diagnostics-First Model + +This note explains the diagnostics-first idea without the full redesign +discussion. + +Use it when you want the architecture model, not the whole design-history note. + +If you want the caller-facing tree choice first, read +[choosing-a-parser-result.md](./choosing-a-parser-result.md). + +If you want the future public-surface direction, read +[api-direction.md](./api-direction.md). + +## The short version + +Diagnostics-first means the parser should be clearest about what it found +before it is clever about how to repair it. + +That does not remove tolerant parsing. It separates three jobs that were too +easy to blur together. + +## The core split + +The parser does three different jobs that used to blur together in the docs. + +```text +diagnostic = factual parser finding +continuation = minimal internal behavior that lets parsing proceed +materialization = caller-visible choice about the final tree shape +``` + +That split matters because `never throw` does not mean "the parser must pick +one canonical repaired tree for everyone." It means the parser keeps going and +preserves what it found. + +In plain English: + +- diagnostics answer "what went wrong here?" +- continuation answers "how did the parser keep going?" +- materialization answers "what final tree should the caller get?" + +When a caller wants to delay that third choice, it needs an `analyze()` lane: +a replayable package of events, diagnostics, and later possible recovery data +that can feed one or more explicit materializers. + +There is also one more exploratory step beyond that: a policy lane where the +caller decides some recoveries itself. That should be treated as a layer on top +of `analyze()`, not as a new parser truth. + +## What callers are really choosing today + +Most callers are choosing between these lanes: + +```text +parse() + -> cheap default tree + +parseWithDiagnostics() + -> default tolerant tree + diagnostics + +parseStrictWithDiagnostics() + -> conservative tree + diagnostics + +parseWithRecovery() + -> default tolerant tree + diagnostics + recovery summary +``` + +The key point is that `parseWithDiagnostics()` and `parseStrictWithDiagnostics()` are not two +different parser truths. They are two materializations of the same parser +findings. + +That is the heart of the model. Malformed input should not force the docs into +teaching several competing parser realities. + +## What diagnostics-first does not mean + +It does not mean the parser stops doing continuation work internally. + +The parser still has to: + +- keep the event stream well-formed +- survive malformed input +- auto-close or normalize enough internal state to stay usable + +The real question is whether those continuation steps become one mandatory +caller-facing tree shape or remain facts plus materialization choices. + +That is why diagnostics-first is not a "strict parser" slogan. The parser can +stay forgiving and still avoid turning one repair policy into the only public +story. + +## Why this matters for malformed input + +Malformed input is where the distinction becomes visible. + +The parser should stay explicit about structural evidence and commitment points. + +```text +before commitment + -> keep text-backed interpretation + +after commitment + -> keep the structural finding + -> let the chosen tree policy decide how much survives in the final tree +``` + +That is why the same malformed input can yield: + +- a tolerant structural node in the default lane +- plain text in the conservative lane +- and the same diagnostics in both + +The parser finding stays the same. What changes is the caller-facing tree +policy. + +## Why this matters for docs and APIs + +If the docs lead with wrapper names alone, they make it sound like the parser +owns several different malformed-input truths. + +The cleaner explanation is: + +1. the parser finds something malformed +2. the parser keeps going +3. the caller chooses whether to ask for a tree now, or keep findings first +4. if the caller does want a tree, the caller chooses how much repair should + show up in the final tree + +That same split should shape the public API over time. + +## The still-open next lane + +There is one stronger diagnostics-first idea that is not fully a public lane +yet: + +```text +parser findings are preserved +recovery data is exposed +final materialization is delayed or caller-owned +``` + +That would go beyond `parseStrictWithDiagnostics()`. `parseStrictWithDiagnostics()` still chooses a final +tree policy for the caller. A true findings lane would expose findings first +and let the caller decide later which recoveries to keep, discard, or replace. + +The event stream should remain the primitive underneath that lane, but a raw +generator is probably not enough as the public shape. A caller may need to +read diagnostics, compare recoveries, and materialize multiple +different tree policies from the same parse. + +The practical target is closer to this: + +```text +source + -> analyze() + events + diagnostics + recovery data + -> materialize(default-html-like) when wanted + -> materialize(source-strict) when wanted + -> materialize(custom policy) only if that later lane proves worth exposing +``` + +That is a real future direction, but the docs should name it as an open design +question rather than implying that the current API already provides it. + +The longer reasoning, rollout questions, and open design questions stay in +[docs/diagnostics-first-redesign.md](../diagnostics-first-redesign.md). \ No newline at end of file diff --git a/docs/architecture/parser-contracts.md b/docs/architecture/parser-contracts.md new file mode 100644 index 0000000..2c41e0e --- /dev/null +++ b/docs/architecture/parser-contracts.md @@ -0,0 +1,156 @@ +# Parser Contracts + +These are the rules every parser API path needs to preserve. + +If a future optimization is faster but breaks one of these contracts, that is a +behavior change, not a harmless refactor. + +## Why this matters + +This repo exposes several views of the same source: + +- tokens +- block-only events +- full events +- tree materializations +- later session and streaming views + +Those outputs are only useful together if they keep telling the same story. + +That shared story is range-first. The parser should preserve the original +source spans and UTF-16 offsets consistently enough that callers can move +between tokens, events, diagnostics, and trees without losing their bearings. + +## Contract 1: event well-formedness + +Every `enter(X)` must have a matching `exit(X)`. + +In plain English, the event stream has to nest like balanced parentheses. + +That matters because tree builders, stream consumers, and direct event walkers +all rely on the same stack discipline. + +## Contract 2: UTF-16 offsets are authoritative + +`position.offset` uses UTF-16 code unit indexing. + +That matches: + +- `string.charCodeAt(i)` +- `string.slice(start, end)` +- the Language Server Protocol's required UTF-16 compatibility path + +When a caller needs exact source fidelity, offsets and source slices win over +derived convenience fields. + +That also means text-like data should stay source-backed for as long as +possible. A later tree node, event, or diagnostic may expose a convenient text +view, but the authoritative grounding is still the original source range. + +## Contract 3: never throw + +The parser must produce a usable result for any input. + +That does not mean every malformed region becomes a clean structure. +It means malformed input should become: + +- text when the source never committed to structure +- tolerant structure plus diagnostics when the source did commit +- or a more conservative tree when the caller asks for it + +But the parser itself does not get to crash. + +The never-throw contract does not authorize silent rewriting of the source. It +authorizes continuation while keeping malformed regions attached to either a +source-backed text interpretation or a clearly signaled structural finding. + +## Contract 4: determinism + +The same input and the same config must produce the same output. + +That includes: + +- token order +- event nesting +- diagnostics +- tree shape for a given materialization policy + +Determinism is what makes corpus regression testing and later incremental work +trustworthy. + +## Contract 5: cross-mode block consistency + +The default structural story must stay aligned across the main parser views. + +```text +outlineEvents() + == block structure from events() + == block structure from parse() +``` + +This matters because `outlineEvents()` is meant to be the cheap structural +overlay. A caller should be able to use it first, then pay for more detail +later, without discovering that the parser changed its mind about which blocks +exist. + +`parseStrictWithDiagnostics()` is the one intentional exception. It may collapse a malformed +committed region back to text, so it is allowed to diverge from that default +overlay. + +Even there, "back to text" should still mean back to a source-backed text span, +not an invented replacement string. + +## Contract 6: commitment points are real boundaries + +A suspicious opener is not enough to create structure. + +The parser should commit structure only when the source crosses the relevant +recognition boundary. + +The clearest current example is HTML-like and extension-like tags: + +```text +before `>` -> keep text-backed interpretation +after `>` -> opener is structurally real +``` + +That is why these two cases behave differently: + +```text + source-backed text in the most conservative interpretation + +body + -> committed structure in the default tolerant lane +``` + +Boundary cases may still be policy questions for more HTML-like recovery. The +contract here is narrower: the parser should use real recognition boundaries +instead of inventing structure from weak evidence. + +## Contract 7: diagnostics are factual parser findings + +Diagnostics should describe what the parser found and what went wrong. + +They should not silently pretend that one recovery materialization is the only +possible meaning of malformed input. + +That is why the repo increasingly treats diagnostics and final tree shape as +separate concerns. + +## What callers can trust after malformed input + +These are the practical trust rules. + +1. Offsets and source slices stay authoritative. +2. The default tolerant lane keeps the main structural overlay stable. +3. `parseWithDiagnostics()` keeps the same default tolerant tree as `parse()`. +4. `parseStrictWithDiagnostics()` is intentionally more conservative and may collapse a + committed malformed region back to text. +5. If a construct never reaches its commitment point, it stays text-backed in + every current tree lane. +6. Diagnostic anchors are tree-local today. They resolve against one final tree + and do not yet promise edit-stable identity across later session changes. + +In all of those cases, "text-backed" means rooted in the original source span, +not eagerly copied into fresh replacement strings. \ No newline at end of file diff --git a/docs/architecture/performance.md b/docs/architecture/performance.md new file mode 100644 index 0000000..3c235a2 --- /dev/null +++ b/docs/architecture/performance.md @@ -0,0 +1,98 @@ +# Performance Model + +The parser's performance work follows one rule: + +remove repeated work without changing the source ranges the parser reports. + +That wording is deliberate. In this repo, source ranges are not an incidental +detail. They are the main way the parser preserves exact source fidelity while +still avoiding unnecessary string allocation. + +## What kinds of changes are valid + +These are the kinds of optimizations that fit the architecture. + +- coalescing adjacent block-parser text events into larger contiguous ranges +- skipping inline rescans when a merged text group contains no possible inline + opener +- avoiding temporary arrays or closures in hot parser paths +- using compact lookup tables for fixed parser vocabularies where repeated + membership checks matter + +These optimizations all fit the same pattern: do less repeated work while still +pointing at the same original text. + +## What kinds of changes are not valid + +An optimization is not acceptable if it: + +- trims real user content that still belongs to a block +- rewrites spacing inside ordinary text ranges +- removes structural boundaries such as table-cell separators +- makes block parsing depend on inline meaning + +The parser can merge ranges and skip redundant work. It cannot change what part +of the source those ranges actually mean. + +It also should not eagerly replace range-first data with copied strings unless a +consumer-facing need clearly justifies the cost. + +## The main current handoff optimization + +The most important recent example is the handoff from `blockEvents()` to +`inlineEvents()`. + +The useful change is: + +```text +old shape + many neighboring text fragments + -> inline parser merges them again + +new shape + one larger contiguous prose range + -> inline parser scans once +``` + +That saves work because the inline parser only needs accurate source coverage. +It does not need old tokenizer-sized boundaries if they no longer carry useful +meaning. + +That is a good example of the repo's performance philosophy. The parser does +not get faster by knowing less about the source. It gets faster by carrying the +same source meaning with fewer duplicated objects and fewer duplicate scans. + +## Why eager positions still matter + +The parser still pays a real cost for eagerly materialized positions. + +Each event currently carries nested start and end points with: + +- line +- column +- offset + +That is valuable for tooling, but it is also a meaningful allocation and +computation cost. That is why benchmark work should keep isolating how much of +the total time comes from: + +- event creation +- position calculation +- nested position-object allocation + +The same kind of scrutiny should apply to eager string materialization. If a +hot path starts producing copied text eagerly where ranges used to be enough, +that is a real performance regression candidate. + +## What this means for future optimization + +If a proposed optimization does not preserve source fidelity and structural +contracts, it is the wrong optimization. + +If it does preserve them, then the next question is simple: + +```text +does this remove repeated work on real inputs? +``` + +That is the bar the performance work should keep using. \ No newline at end of file diff --git a/docs/architecture/sessions-and-streaming.md b/docs/architecture/sessions-and-streaming.md new file mode 100644 index 0000000..26838f0 --- /dev/null +++ b/docs/architecture/sessions-and-streaming.md @@ -0,0 +1,117 @@ +# Sessions and Streaming + +The stateless parser functions are the core API. + +The session API exists for the cases where a caller wants to ask more than one +question about the same source, or where the source changes over time. + +## What a session is + +A session is a cached wrapper around the same parser pipeline. + +```text +createSession(source) + -> cached outline events + -> cached full events + -> cached tree materializations + -> later streaming and incremental edit support +``` + +It is not a different parser. + +It is still the same range-first parser. The session mainly caches views over +the same source and, later, over changes to that source. + +## What sessions buy you + +Without a session, each call starts again from the original source. + +With a session, one caller can ask: + +- `outline()` +- `events()` +- `parse()` +- `parseWithDiagnostics()` +- `parseStrictWithDiagnostics()` + +and reuse earlier work where that reuse is valid. + +That matters most for editors, live previews, and repeated structural queries. + +That reuse works because the underlying outputs stay anchored to the same +source ranges and UTF-16 offsets. If the parser eagerly rewrote text into new +strings at every step, much less of that work would compose cleanly. + +## Cache lanes follow the same cost model + +The session should not force every caller onto the expensive diagnostics path. + +That is why its caches follow the same split as the stateless APIs: + +```text +diagnostics off + -> cheapest event and tree access + +diagnostics on + -> diagnostics-preserving event and tree access +``` + +If a caller only wants the cheap lane, the session should not allocate or keep +diagnostics-enabled state unless some other call already paid for it. + +The same rule applies to text materialization. Sessions should prefer caching +range-first parser products over eagerly caching many duplicate string views of +the same source. + +## Streaming and the stability frontier + +Streaming input creates one extra problem: the parser needs to distinguish what +is already stable from what may still change. + +```text +stable prefix | provisional tail +``` + +The boundary between them is the stability frontier. + +Before the frontier, events are safe to treat as committed. +After the frontier, later source may still close an open construct and change +the parse. + +That is the core idea behind future streaming support such as: + +- `session.write(chunk)` +- stable-event draining +- provisional-tail replacement + +## Incremental editing + +Incremental parsing is the later session feature for edited text, not just +append-only streaming. + +The parser direction here is: + +- find the smallest safe region to reparse +- reuse the rest of the prior work +- return a `PositionMap` so cursors, selections, and comment-like anchors can + be translated across edits + +That only works if positions and source spans stay authoritative enough to map +old work onto new source confidently. + +That is how the parser supports editor-like workflows without turning the core +pipeline into an editor framework. + +## Why this belongs outside the main architecture overview + +Sessions, streaming, and incremental reparsing are important, but they are not +the first thing most readers need. + +Most people first need to know: + +- what the parser outputs are +- which tree lane to choose +- how the pipeline roughly works + +This note exists so those live-use-case details stop crowding the main +architecture entry point. \ No newline at end of file diff --git a/docs/architecture/utility-first.md b/docs/architecture/utility-first.md new file mode 100644 index 0000000..333bc2b --- /dev/null +++ b/docs/architecture/utility-first.md @@ -0,0 +1,101 @@ +# Utility-First Design + +This parser is utility-first. + +That means it tries to expose stable primitives that other tools can build on, +instead of turning the core parser into a giant plugin host. + +## What utility-first means here + +The package exposes building blocks such as: + +- `TextSource` +- tokens +- events +- tree builders +- tree utilities +- session-oriented wrappers + +Those are the surfaces other tools should compose. + +Just as importantly, those surfaces preserve the parser's native data model: +source-backed ranges first, richer interpretations second. + +## Why the parser does not lead with hooks + +The repo does not currently treat deep parser hooks as the main extension +story. + +That is intentional. + +Deep hooks freeze a lot of internal design too early: + +- ambiguity handling +- malformed-input boundaries +- hot-path control flow +- recovery and materialization policy + +If those internals are still evolving, a broad hook system creates long-term +API cost very quickly. + +## What to do instead + +The preferred extension model is to build on the parser's public primitives. + +In practice, that usually means one of these: + +1. consume tokens and emit your own higher-level events +2. consume the event stream and emit your own derived events +3. build a focused mini-parser for one domain-specific construct and merge its + output with the main parser's results +4. walk the tree and add your own interpretation layer later + +This keeps the core parser small and predictable while still giving consumers a +real extension surface. + +It also lets downstream tools choose when they actually need strings. Many +extensions can do their work directly from tokens, source ranges, event slices, +or tree nodes that still point back into the original source. + +## The recommended mental model + +Think of the package less like this: + +```text +one parser you patch from the inside +``` + +and more like this: + +```text +shared parser primitives + -> your domain-specific logic + -> your own events, trees, or transforms +``` + +If you need special behavior for one feature, it is often better to write a +small focused parser that consumes the source ranges or event slices you care +about than to ask the core parser to grow a new generic hook point. + +That recommendation is partly architectural and partly practical: range-first +primitives are cheaper to compose than hook systems that force the core parser +to materialize and expose every intermediate detail eagerly. + +## What stays public versus internal + +Public and intended for downstream use: + +- source abstractions such as `TextSource` +- token, event, and AST interfaces and unions +- builder functions and type guards +- parser stage entry points such as `tokenize()`, `blockEvents()`, and + `inlineEvents()` + +Still internal: + +- scanner-local context records +- one-stage matcher result shapes +- low-level continuation helpers whose contracts are still evolving + +That boundary is what keeps the package utility-first without making every +local implementation detail part of the public API. \ No newline at end of file diff --git a/docs/diagnostics-first-redesign.md b/docs/diagnostics-first-redesign.md new file mode 100644 index 0000000..add43fc --- /dev/null +++ b/docs/diagnostics-first-redesign.md @@ -0,0 +1,550 @@ +# Diagnostics-First Parser Redesign + +Status: the first public realignment has landed. The parser now exposes +`parseStrictWithDiagnostics()` and `buildTreeStrict()`, and event-stream diagnostics are +opt-in by default. The sections below still capture the reasoning and the +remaining architectural direction behind that shift. + +If you want the smaller architecture-facing explanation first, start with: + +- [docs/architecture/diagnostics-first.md](./architecture/diagnostics-first.md) +- [docs/architecture/choosing-a-parser-result.md](./architecture/choosing-a-parser-result.md) +- [docs/architecture/malformed-input.md](./architecture/malformed-input.md) +- [docs/architecture/api-direction.md](./architecture/api-direction.md) +- [docs/architecture/utility-first.md](./architecture/utility-first.md) +- [docs/architecture/performance.md](./architecture/performance.md) +- [docs/architecture/parser-contracts.md](./architecture/parser-contracts.md) + +This note audits the current parser surface against the intended +diagnostics-first model and proposes a concrete redesign. + +Use this note for the deeper design argument, the migration pressure, the +remaining rollout questions, and the open questions. For the shorter +architecture explanation, use +[docs/architecture/diagnostics-first.md](./architecture/diagnostics-first.md). + +The goal is not to remove tolerant defaults. The goal is to separate three +different ideas that have drifted together in the current API: + +- detecting malformed input +- continuing parsing without throwing +- materializing one particular recovered tree shape + +That separation matters because the parser should be explicit about its +defaults without compelling every consumer to accept the parser's own recovery +materialization. + +## Problem + +The current public story treats recovery as a parser-owned output lane. + +```text +current public framing + +source + ├─► parse() -> default tree + ├─► parseWithDiagnostics() -> default tree + diagnostics + └─► parseWithRecovery() -> default tree + diagnostics + recovery summary +``` + +The intended model is: + +- diagnostics are the primitive +- the parser may continue internally so it does not throw +- recovery materialization is an optional consumer response to diagnostics +- downstream consumers may expand parser diagnostics into their own + domain-specific diagnostics + +In other words, `never throw` means the parser keeps going and preserves the +problem. It does not mean the parser must publish one canonical recovered tree +shape as the meaning of malformed input. + +The current code still has useful pieces of that original design. In +[events.ts](events.ts), error events are optional, structured, and already +described as consumer-facing signals that can be logged, surfaced, ignored, or +expanded. The drift happens higher in the stack, where the parser and tree +builder elevate recovery materialization into first-class parser lanes. + +## Goals + +- Keep the parser never-throw contract. +- Stay close to HTML-like parsing with explicit commitment points and tolerant + defaults. +- Make diagnostics the primary public primitive for malformed input. +- Keep diagnostic emission optional so consumers that do not want diagnostics + do not pay for them. +- Surface recovery materialization options explicitly. +- Avoid compelling consumers to accept a recovered tree if they only want + diagnostics, source-backed text, or their own downstream recovery logic. +- Preserve the ability for downstream tools such as renderers and editors to + emit their own diagnostics on top of parser diagnostics. + +## Non-goals + +- Remove all internal continuation heuristics from the parser. +- Make malformed input fatal. +- Remove tolerant default behavior for consumers that do want it. +- Solve edit-stable diagnostic anchors in this redesign. + +## Current Drift + +### The event model is mostly aligned + +The event layer in [events.ts](events.ts) is already close to the intended +shape: + +- diagnostics are optional `error` events +- diagnostic codes are stable parser-owned facts +- diagnostic metadata is open enough for downstream consumers +- docs already describe consumer choice as the point of those diagnostics + +This is the strongest part of the current design and should remain the base +primitive. + +### The parser API still makes materialization feel more parser-owned than it should + +The orchestration API in [parse.ts](parse.ts) currently frames the parser as +several top-level tree lanes: + +- `parse()` returns the default HTML-like tree +- `parseWithDiagnostics()` returns that same default tree plus diagnostics +- `parseStrictWithDiagnostics()` returns the source-strict tree plus diagnostics +- `parseWithRecovery()` returns the default tree plus diagnostics and `recovered` + +That is better than the older framing, because diagnostics are now opt-in by +default and `parseStrictWithDiagnostics()` names the conservative lane more honestly. But it +still makes materialization feel like a parser-owned truth instead of a +consumer policy layered over the same syntax findings. + +### The tree builder currently owns recovery-shape policy + +[tree_builder.ts](tree_builder.ts) currently exposes `TreeBuildMode` as +`'strict' | 'loose'` and documents those modes as recovery-shape policies. + +That creates two problems: + +- tree materialization policy is presented as if it were part of parsing truth +- diagnostics and recovered tree shape are coupled in one API family + +The addition of `buildTreeWithLooseDiagnostics()` helps operationally because +it separates `recovered` from `diagnostics`, but it still keeps the public +story centered on parser-owned recovery shape. + +### The current docs teach the drift as architecture + +[docs/architecture.md](docs/architecture.md) and [readme.md](readme.md) +currently describe `strict` and `loose` as top-level parser choices and teach +the tree lanes as the main way to think about malformed input. + +That is the opposite of the intended mental model. It teaches readers that the +parser owns recovery semantics, when the intended design is that the parser +owns diagnostics and continuation, and consumers own recovery responses. + +### Tests currently lock in parser-owned recovery semantics + +[parse_test.ts](parse_test.ts) and [session_test.ts](session_test.ts) assert +that: + +- `parseWithDiagnostics()` and `parseWithRecovery()` are separate wrappers even + though they now share the same default tree shape +- strict versus loose tree shape is a parser-level distinction +- recovery-specific wrapper nodes are part of the parser's promised result + +Those tests are useful because they show the exact behavioral drift. They will +need to move as the API is realigned. + +## Design Distinction To Restore + +The core distinction should be: + +```text +diagnostic = factual parser finding +continuation = minimal internal behavior that lets parsing proceed +materialization = consumer-visible choice about how to represent malformed input +``` + +That gives a cleaner mental model: + +```text +source + ├─► parser detects malformed input + ├─► parser records a diagnostic when requested + ├─► parser continues internally so it never throws + └─► consumer chooses whether to: + - ignore the diagnostic + - surface the diagnostic + - materialize a tolerant tree + - materialize a conservative tree + - emit richer downstream diagnostics +``` + +The parser still needs continuation heuristics. A block parser cannot keep +streaming without making some local choice when a table never closes or an +event stream ends with open frames. But those choices should be treated as +survivability mechanics, not as the only public recovery meaning. + +## Proposal + +### 1. Make diagnostics the primary malformed-input contract + +Keep parser diagnostics as structured event-layer facts. + +The parser should continue to emit stable `DiagnosticCode` values and optional +details, but the docs should stop describing those diagnostics as evidence of a +parser-owned recovery lane. They are parser findings that consumers can act on +or translate. + +### 2. Rename the internal concept from recovery to continuation where possible + +Inside parser implementation docs and comments, prefer `continuation` for the +minimum logic required to keep scanning or keep the event stream well-formed. + +Use `recovery` for consumer-visible responses to diagnostics or for explicit +materialization helpers that are intentionally opt-in. + +This is especially important in: + +- [block_parser.ts](block_parser.ts) +- [inline_parser.ts](inline_parser.ts) +- [tree_builder.ts](tree_builder.ts) + +### 3. Move tree shape choices out of the parser-success narrative + +Tree shape should be described as a materialization policy, not as the parser's +acceptance or recovery mode. + +A clearer conceptual split is: + +- parser options decide whether diagnostics are emitted +- materializer options decide how malformed regions are represented + +That means the parser surface should stop teaching `strict` and `loose` as the +main mental model for malformed input. + +### 4. Keep diagnostic emission optional and explicit + +This part is already directionally correct in the implementation and should be +preserved. + +If a caller says they do not want diagnostics, block and inline parsers should +not emit diagnostic events at all. + +That cost model should be made explicit in the public API: + +```text +diagnostics off + -> no diagnostic events emitted + -> no diagnostic collection + -> cheapest parse lane + +diagnostics on + -> diagnostic events emitted + -> downstream recovery-aware materializers may consume them + -> caller accepts the added allocation and processing cost +``` + +### 5. Reframe public tree APIs around materialization policy + +Replace the current recovery-centered API story with a diagnostics-first, +materialization-second story. + +One concrete direction: + +```ts +interface ParseOptions { + readonly diagnostics?: boolean; + readonly materialization?: 'default-html-like' | 'source-strict'; +} + +interface ParseOutput { + readonly tree: WikistRoot; + readonly diagnostics?: readonly ParseDiagnostic[]; +} +``` + +With convenience wrappers if desired: + +- `parse(source)` + default HTML-like materialization, no diagnostics +- `parseWithDiagnostics(source)` + default HTML-like materialization plus diagnostics +- `parseStrictWithDiagnostics(source)` + source-strict materialization plus diagnostics by default + +The important change is not the exact names. The important change is that the +API no longer teaches "diagnostics lane versus recovery lane" as if those were +peer parser truths. Instead it teaches: + +- do you want diagnostics? +- if yes, how do you want malformed regions materialized? + +For that public-facing option shape, `diagnostics` should be the public name. +It is shorter, and it describes the caller's choice instead of the +implementation detail of whether diagnostic events are being emitted +internally. + +### 6. Treat explicit recovery materialization as an opt-in helper layer + +If the package wants to preserve strong support for tolerant consumers, keep +that support as explicit helpers rather than as the parser's main malformed- +input identity. + +For example: + +- `buildTree(events, { source })` + default HTML-like materialization +- `buildTreeWithDiagnostics(events, { source })` + same materialization plus anchored diagnostics +- `materializeDiagnostics(events, { source, policy: 'source-strict' })` + conservative materialization for linting or editor diagnostics +- `materializeDiagnostics(events, { source, policy: 'default-html-like' })` + tolerant materialization for rendering + +This keeps recovery materialization surfaced, explicit, and optional. + +### 7. Add materialization hints to diagnostics instead of baking one response into the tree API + +If parser diagnostics need to help consumers choose a response, add stable +metadata that describes the kind of malformed region without forcing one public +repair. + +For example, the diagnostic payload could eventually carry fields such as: + +- `code` +- `source` +- `details` +- `continuation_kind` +- `suggested_materializations` + +That would let a renderer, editor, or linter make an informed choice without +requiring the parser to publish one canonical recovered tree for every case. + +This should stay conservative. The parser should expose facts and hints, not a +large strategy engine. + +### 8. Leave room for a true `analyze()` lane + +There is one stronger diagnostics-first option that the current public API does +not fully provide yet. + +Today the parser offers: + +- a cheap tree with no preserved diagnostics +- the default tolerant tree with diagnostics +- a conservative tree with diagnostics + +What it does not fully offer yet is this: + +```text +diagnostics and possible recoveries are exposed +but the caller chooses the final materialization later +``` + +That would be more diagnostics-first than `parseStrictWithDiagnostics()`. `parseStrictWithDiagnostics()` +still chooses a final conservative tree policy on the caller's behalf. A true +`analyze()` lane would expose findings first and make the final repair policy +more explicitly caller-owned. + +The practical shape matters here. The best public answer is probably not just a +bare one-shot event generator. The parser is already event-stream-first, but a +control-heavy caller often needs more than one pass over the same information. + +That caller may need to: + +- inspect diagnostics before choosing a tree policy +- compare more than one candidate repair path +- materialize both tolerant and conservative trees from the same parse +- cache findings inside a session or incremental workflow + +That points toward a replayable findings utility built on the event stream. +Conceptually: + +```text +source + -> analyze() + -> findings + - replayable events + - diagnostics + - recovery data + -> materialize(findings, tolerant) + -> materialize(findings, conservative) +``` + +This should stay an open design question until the package has a clearer model +for what a "possible recovery" surface actually looks like and how much of that +surface belongs to the parser instead of later consumers. + +## Trust rules and practical contracts + +The practical trust rules now live in the smaller architecture notes so this +design note does not have to reteach the same ground every time. + +- [docs/architecture/parser-contracts.md](./architecture/parser-contracts.md) + covers the invariants the parser should preserve. +- [docs/architecture/malformed-input.md](./architecture/malformed-input.md) + covers commitment points and malformed-input behavior. +- [docs/architecture/diagnostic-anchors.md](./architecture/diagnostic-anchors.md) + covers how diagnostics point back into the final tree. + +## How To Tell Whether The Block/Inline Split Is A Problem + +The existence of two parser stages is not evidence of a design problem. The +useful question is whether the handoff causes correctness drift or measurable +extra work. + +Check it in this order: + +1. Correctness contract: + `outlineEvents()`, `events()`, and `parse()` should agree on block + structure for the same input in the default tolerant lane. +2. Handoff behavior: + merged block text groups should not trim content, invent boundaries, or + force the inline parser to reconstruct information the block parser already + knew. +3. Cost profile: + measure text-event counts, merged text-group counts, inline candidate + scans, allocations, and total throughput on prose-heavy, syntax-heavy, and + malformed corpora. +4. Change threshold: + only collapse the stages or redesign the boundary if the split causes a + real correctness bug or a benchmarked throughput or allocation regression. + +The useful comparison is not "two stages versus one stage" in the abstract. +It is "does this handoff preserve the contracts while reducing repeated work?" + +```text +bad split + block parser fragments text + -> inline parser has to reconstruct the same logical group + -> correctness or cost regresses + +good split + block parser emits one accurate contiguous prose range + -> inline parser scans once for real inline openers + -> contracts stay intact and work drops +``` + +That is why the next verification step should be contract tests plus focused +benchmarks, not a premature rewrite into one monolithic parser. + +## Primitive-first extension model + +The longer architecture-facing explanation now lives in +[docs/architecture/utility-first.md](./architecture/utility-first.md). + +The short version here is still important: the redesign works better if the +parser leads with stable primitives and downstream composition instead of a deep +hook surface that freezes hot-path behavior too early. + +## Ambiguity hotspots worth hardening next + +The narrow handoff and bare-URL follow-up now lives in +[docs/handoff-and-bare-url-notes.md](./handoff-and-bare-url-notes.md). + +That keeps this redesign note focused on the tree and diagnostics model instead +of turning it into a second all-purpose architecture file. + +## HTML-Like Parsing Guidance + +Staying close to HTML-like parsing still fits this redesign. + +The parser should continue to use explicit commitment points. The existing +HTML-like tag rule in [inline_parser.ts](inline_parser.ts) is a good example: + +- before the opener reaches `>`, do not commit a real tag node +- after the opener reaches `>`, the opener is structurally real +- if the close tag never arrives, emit a diagnostic and let materializers + choose how much tolerant structure to preserve + +That is a good default because it is explicit and predictable. What should +change is not the commitment rule. What should change is the public framing of +who owns the final recovery materialization. + +## Concrete API direction + +The future public-surface cleanup now lives in +[docs/architecture/api-direction.md](./architecture/api-direction.md). + +That keeps this redesign note focused on the larger architectural problem while +still preserving one place for the future wrapper and option-shape discussion. + +## Performance Model + +This redesign keeps the fast path clear. + +```text +cheap lane + diagnostics off + -> tokenizer + -> block parser without diagnostic events + -> inline parser without diagnostic events + -> default materialization if requested + +diagnostics lane + diagnostics on + -> tokenizer + -> block parser with diagnostic events + -> inline parser with diagnostic events + -> diagnostics-preserving materialization if requested +``` + +This matches the intended trade-off: + +- consumers that do not want diagnostics do not pay for them +- consumers that want diagnostics and recovery-aware handling accept the cost + +That cost model should be explicit in docs, code comments, and session cache +structure. + +## Rollout Plan + +1. Rewrite docs so diagnostics are the primary malformed-input concept and + tree shape is described as materialization policy. +2. Deprecate recovery-centered names where they teach the wrong model, + especially `parseWithRecovery()` as the main public malformed-input lane. +3. Introduce a findings utility and tree APIs that separate diagnostic emission + from materialization policy. +4. Move tests from "parser owns recovery tree semantics" to + "materializer policy changes representation of the same parser findings". +5. Keep compatibility wrappers temporarily if migration cost matters. +6. Add contract tests for cross-mode block consistency, commitment-point + behavior, and the default-versus-source-strict trust boundaries. +7. Add focused benchmarks or counters around the block-to-inline handoff before + reconsidering the two-stage design. + +This section stays here for now because it is still part of the redesign +discussion, not yet a settled implementation plan. + +## Open Questions + +- Should the default `events()` lane include diagnostics, or should the + cheapest no-diagnostics lane be the default and diagnostics require explicit + opt-in? +- Should the package keep a tolerant convenience wrapper equivalent to the + current `parseWithRecovery()`, or should that become a more clearly named + materialization helper? +- Should the `analyze()` lane expose replayable arrays, a session-backed view, or + another reusable iterable shape? +- How much structured response metadata should diagnostics carry before the + parser starts looking like a strategy engine? +- Should tree-stage findings such as mismatched exits remain public parser + diagnostics, or be treated as internal materializer diagnostics unless the + caller explicitly requests low-level stream integrity facts? +- How narrow should the core bare-URL recognizer stay before profiles or + downstream consumers take over richer URL or IRI handling? + +## Summary + +The current implementation already has the right low-level foundation: +optional structured diagnostics and explicit cost gates for emitting them. + +The drift is mainly in the public framing and the tree APIs. They currently +teach recovery materialization as a parser-owned lane. + +The redesign should restore this rule: + +- the parser owns diagnostics and continuation +- consumers own whether and how malformed input is materially recovered + +That keeps the parser explicit, tolerant, HTML-like, and performant without +forcing downstream consumers into one recovery worldview. \ No newline at end of file diff --git a/docs/examples.md b/docs/examples.md new file mode 100644 index 0000000..25e0d06 --- /dev/null +++ b/docs/examples.md @@ -0,0 +1,163 @@ +# Examples + +These examples show the most common ways to use the parser without reading the +deeper architecture notes first. + +They start with the simplest paths and then move toward more inspection-heavy +flows. + +## Build a tree for normal document work + +Use `parse()` when you want a tree and you do not need diagnostics. + +```ts +import { parse } from '@okikio/wikitext'; + +const tree = parse('== Heading ==\n\nA paragraph with [[Main Page|a link]].'); +console.log(tree.type); // 'root' +``` + +This is the cheapest tree path. It is a good fit when the source is mostly +ordinary wikitext and you do not need the parser's explanation of malformed +regions. + +## Inspect headings without building a tree + +Use `outlineEvents()` when you only care about block structure such as headings +or lists. + +```ts +import { outlineEvents } from '@okikio/wikitext'; + +const headings = []; + +for (const event of outlineEvents('== One ==\n\nText\n\n=== Two ===')) { + if (event.kind === 'enter' && event.node_type === 'heading') { + headings.push(event); + } +} +``` + +## Keep diagnostics while preserving the default tree + +Use `parseWithDiagnostics()` when you want the parser's default tolerant tree +and also want to know where malformed input was detected. + +```ts +import { parseWithDiagnostics } from '@okikio/wikitext'; + +const result = parseWithDiagnostics('Paragraph with note'); + +console.log(result.tree.children[0]?.type); +console.log(result.diagnostics.map((diagnostic) => diagnostic.code)); +``` + +This is the best current fit for the "recover on my behalf, but tell me what +you did" style of usage. + +## Ask for the conservative tree instead + +Use `parseStrictWithDiagnostics()` when diagnostics matter and you want the final tree to be +more conservative about malformed committed structure. + +```ts +import { parseStrictWithDiagnostics } from '@okikio/wikitext'; + +const result = parseStrictWithDiagnostics('Paragraph with note'); + +console.log(result.tree.children[0]?.type); +console.log(result.diagnostics.length > 0); +``` + +This is useful for inspection or linting flows that do not want the final tree +to keep as much repaired structure. + +## Compare the current tree lanes on one malformed input + +The easiest way to understand the tree differences is to run the same input +through more than one lane. + +```ts +import { parse, parseStrictWithDiagnostics, parseWithDiagnostics } from '@okikio/wikitext'; + +const input = 'Paragraph with note'; + +const fast_tree = parse(input); +const recovery_tree = parseWithDiagnostics(input); +const conservative_tree = parseStrictWithDiagnostics(input); + +console.log(fast_tree.children[0]?.type); +console.log(recovery_tree.tree.children[0]?.type); +console.log(conservative_tree.tree.children[0]?.type); +console.log(recovery_tree.diagnostics.map((diagnostic) => diagnostic.code)); +``` + +This kind of comparison is useful when you are deciding whether your tool wants +speed first, tolerant structure first, or conservative diagnostics first. + +## Resolve a diagnostic back to the tree + +Diagnostics live beside the tree, but you can still resolve them back to the +nearest materialized node. + +```ts +import { locateDiagnostic, parseWithDiagnostics } from '@okikio/wikitext'; + +const result = parseWithDiagnostics('{|\n| Cell'); +const location = locateDiagnostic(result.tree, result.diagnostics[0]); + +console.log(location?.node.type); +console.log(location?.parent?.type); +console.log(location?.index); +``` + +This is useful for editor tooling, lint output, or any workflow that wants to +connect a parser finding back to a concrete part of the final tree. + +## Reuse one source through a session + +Use a session when you want to ask more than one question about the same input. + +```ts +import { createSession } from '@okikio/wikitext'; + +const session = createSession('== Heading ==\n\nParagraph with [[Main Page]].'); + +const outline = Array.from(session.outline()); +const full_events = Array.from(session.events()); +const tree = session.parse(); + +console.log(outline.length > 0); +console.log(full_events.length > 0); +console.log(tree.type); +``` + +This avoids redoing the same work from scratch each time you ask a different +question about one source string. + +## Build a focused parser on top of events + +This package is utility-first. If you need domain-specific behavior, prefer +building on the public primitives instead of expecting deep hooks in the core +parser. + +```ts +import { events } from '@okikio/wikitext'; + +function collectLinks(source: string): string[] { + const links: string[] = []; + + for (const event of events(source)) { + if (event.kind === 'enter' && event.node_type === 'wikilink') { + links.push(String(event.props.target)); + } + } + + return links; +} + +console.log(collectLinks('Visit [[Main Page]] and [[Help:Contents]].')); +``` + +If you want the reasoning behind that design, see +[docs/architecture/utility-first.md](./architecture/utility-first.md). \ No newline at end of file diff --git a/docs/future-direction.md b/docs/future-direction.md new file mode 100644 index 0000000..0a338d1 --- /dev/null +++ b/docs/future-direction.md @@ -0,0 +1,168 @@ +# Future Direction + +This note captures the broader direction behind the current parser work. + +The parser is the current product focus, but it is not the full long-term +product story. Writing that down matters because parser architecture choices +look different when the parser is understood as a base layer rather than the +final user-facing system. + +## What this repo is building now + +Right now this repository is building a standards-aligned wikitext source +parser. + +Its job is to turn source text into reusable parser primitives: + +- tokens +- events +- trees +- diagnostics +- session-oriented access patterns later on + +Those primitives need to be good enough for extraction, rendering, +transformation, inspection, and editor-facing tooling. + +## What comes next + +Two higher-level libraries should sit above the parser once the core surface is +ready enough. + +```text +@okikio/wikitext + ├─► @okikio/wikidoc + └─► @okikio/wiki-extract +``` + +### `@okikio/wikidoc` + +This layer should focus on document-oriented work built on top of parser +primitives. + +That likely includes: + +- higher-level document transforms +- structure-aware editing helpers +- richer composition around trees, diagnostics, and sessions +- bridges toward editor and CMS-style workflows + +### `@okikio/wiki-extract` + +This layer should focus on pulling useful structured information out of parsed +content. + +That likely includes: + +- template and infobox extraction +- category and link extraction +- reference extraction +- selective structure queries and summaries + +The key point is that extraction and document composition should not distort the +parser core. They should consume parser primitives instead. + +## Longer-term direction + +Longer term, the same primitives may expand into a broader profile-driven +document engine. + +That direction is bigger than "a better wikitext parser." It points toward a +system that can support: + +- wikitext as the proving ground +- additional markup or rich-text families later +- CMS-style structured blocks +- editor-facing session and incremental workflows +- local-first or collaboration-aware tooling +- LLM-oriented document transforms built on explicit structure rather than raw + text guessing + +The right way to read this is as a direction of travel, not as a promise that +all of those layers belong in this package soon. + +## Why this matters to the parser now + +This longer horizon changes what "good parser architecture" means. + +The parser should optimize for reusable primitives, not just one rendering +pipeline. + +That means: + +- events need to stay first-class +- diagnostics need to stay factual and reusable +- offsets need to stay authoritative +- extension boundaries need to stay narrow and composable +- session and incremental APIs need to build on the same core primitives + +It also means the parser should avoid baking in product-specific behavior that +really belongs one layer up. + +## Non-goals for the core parser + +The parser core should not silently become: + +- a full MediaWiki runtime +- a CMS block model +- an editor framework +- a collaboration engine +- an extraction product + +Those may all become consumers or adjacent layers. They should not become the +hot loop. + +## Design pressure this adds today + +This broader direction puts useful pressure on current decisions. + +### Diagnostics versus materialization + +Different higher-level consumers need different responses to malformed input. +An extractor, a renderer, and an editor may all want the same parser finding +but a different final representation. + +That is why diagnostics should stay parser facts and final tree shape should +stay a materialization choice. + +### Primitive-first extension model + +If the long-term system grows into multiple products, deep parser hooks become +too expensive too early. + +The core should prefer: + +- stable tokens, events, and trees +- narrow feature gates +- explicit tag or profile handlers +- downstream enrichment passes + +### Profile-driven behavior + +The broader engine direction does not require one giant parser that tries to do +everything at once. + +It points instead toward profiles layered over shared primitives: + +```text +shared parser primitives + ├─► syntax profile + ├─► mediawiki-compatible profile + └─► later document-oriented profiles +``` + +That keeps the core maintainable while leaving room for stricter or more +product-specific behavior later. + +## How to use this note + +Use this note when judging parser design trade-offs. + +If two parser designs look equally good for today's tests, prefer the one that: + +- keeps primitives reusable +- keeps product policy out of the hot path +- keeps future extraction and document tooling possible +- preserves optionality for profile-driven evolution + +If a design only makes sense for one consumer and hardens that assumption into +the parser core, it is probably the wrong long-term move. \ No newline at end of file diff --git a/docs/handoff-and-bare-url-notes.md b/docs/handoff-and-bare-url-notes.md new file mode 100644 index 0000000..92abaed --- /dev/null +++ b/docs/handoff-and-bare-url-notes.md @@ -0,0 +1,165 @@ +# Block/Inline Handoff and Bare-URL Notes + +This note records two narrow follow-ups from the diagnostics-first redesign: + +1. how to tell whether the block-to-inline split is helping or hurting +2. which bare-URL rules belong in the core parser instead of downstream tools + +The goal is to keep the work measurable and local. Neither topic needs a broad +parser rewrite to become clearer. + +## What this note is for + +The parser already has a strong default shape: + +```text +TextSource -> tokenizer -> blockEvents() -> inlineEvents() -> tree consumers +``` + +That shape is worth keeping as long as the handoff preserves structure and +reduces repeated work. The right next step is not "merge the stages because two +stages feel suspicious." The right step is to measure the handoff directly and +to tighten the ambiguity rules that still look too ad hoc. + +Bare URLs are the clearest current example of that second category. + +## How to evaluate the block-to-inline handoff + +There are two separate questions. + +First, does the handoff preserve the parser contracts? + +- `outlineEvents()` and `events()` must agree on block structure. +- `parse()` must preserve that same block structure in the default lane. +- contiguous prose ranges must not lose spaces, punctuation, or boundaries + just because they were merged before inline parsing. + +Second, does the handoff actually save work? + +The useful benchmark comparison is: + +```text +fragmented handoff + many neighboring text events + -> inline parser merges them before scanning + +merged handoff + one larger contiguous text event + -> inline parser scans once +``` + +If merged handoff is not faster, or if it causes structural drift, then the +boundary needs work. If it is faster and keeps the contracts intact, then the +split is doing its job. + +That is why the benchmark work here isolates merged versus fragmented text +groups for the same input instead of comparing unrelated documents. + +## What belongs in the core bare-URL matcher + +The core parser does need bare-URL recognition, but it should stay narrow and +cheap. + +The matcher should own only the rules needed to answer this parser question: + +```text +does this source range clearly contain a link-shaped URL token here? +``` + +That means the core matcher should handle: + +- accepted schemes in the core lane +- start boundaries, so URLs do not begin in the middle of ordinary words +- end boundaries, so trailing prose punctuation does not become part of the URL +- lightweight balancing rules for common prose wrappers such as a trailing `)` + +That does not mean the core parser should do full URL normalization, +base-URL-aware resolution, or browser-style canonicalization. Those jobs belong +to downstream consumers. + +## Current acceptance matrix + +The current matcher keeps two different promises. + +- Explicit bracketed external-link syntax is broad. If the author wrote + `[scheme:payload Label]`, the parser should usually trust that they meant a + link and leave scheme-policy decisions to downstream consumers. +- Bare autolinks in prose are narrower. They should catch the common cases and + a strong set of opaque URI shapes, but they should not overmatch ordinary + colon prose. + +That gives this working matrix: + +| Case | Bare prose | Bracketed explicit syntax | Reason | +| --- | --- | --- | --- | +| `https://example.com` | accept | accept | authority form is high-confidence | +| `file:///Users/example/report.txt` | accept | accept | authority form is high-confidence | +| `foo+bar://example.service/path` | accept | accept | authority form is high-confidence | +| `mailto:editor@example.org` | accept | accept | `@` is strong opaque-URI evidence | +| `urn:isbn:0451450523` | accept | accept | repeated `:` plus digits is strong opaque-URI evidence | +| `data:text/plain,hello` | accept | accept | `/` and `,` make the payload structurally URI-like | +| `tel:+12025550123` | accept | accept | leading `+` plus digits is strong opaque-URI evidence | +| `magnet:?xt=urn:btih:...` | accept | accept | `?` and `=` make the payload structurally URI-like | +| `longcustomscheme:alpha` | reject | accept | too little structure for safe bare autolinking | +| `note:abc` | reject | accept | reads like ordinary colon prose in bare text | +| `chapter:one` | reject | accept | reads like ordinary colon prose in bare text | + +The general rule is simple: if a URI is not distinct enough to be recognized +confidently in bare prose, it should stay text and let an end consumer extend +that policy if they care about a niche scheme. + +The stable low-level boundary rules are still: + +- a bare URL must start at a word boundary, not in the middle of an ASCII word +- the scan stops before whitespace, quotes, `]`, `<`, or `>` +- trailing prose punctuation such as `.`, `,`, `;`, `:`, `!`, and `?` is + trimmed from the URL +- a trailing `)` is trimmed only when it is unmatched within the scanned URL + +That gives a practical parser rule: + +```text +Visit https://example.com. + ^^^^^^^^^^^^^^^^^^^ external-link + . plain text +``` + +and: + +```text +(https://example.com/path(test)) + ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ external-link + ) plain text +``` + +The parser keeps the balanced inner parentheses because they can be part of a +real URL path, but it leaves the extra outer prose wrapper behind. + +## What still stays out of scope + +The core matcher still does not try to solve every URI or IRI question. + +That is intentional. + +- RFC 3986 and RFC 3987 are useful references for boundary decisions. +- WHATWG URL behavior is useful for web-facing expectations. +- MediaWiki behavior still matters where it intentionally differs. + +But the hot path should not instantiate `URL` objects or require a base URL +just to recognize a bare link-shaped span in plain text. That would increase +allocation cost and couple parser recognition to downstream normalization. + +## What to watch next + +If future work expands bare-URL support, the next useful checks are: + +1. whether any rejected bare opaque cases deserve promotion into the core + acceptance matrix or should stay consumer-owned extensions +2. whether any IRI behavior is important enough to justify extra scan cost +3. whether MediaWiki-compatible punctuation trimming needs a more exact rule +4. whether URL-heavy prose changes the value of the current block-to-inline + merge strategy + +Those are good follow-ups only if the current tests and benchmarks show a real +gap. Until then, the better strategy is to keep the core matcher small, +predictable, and benchmarked. \ No newline at end of file diff --git a/docs/parser-architecture-comparison.md b/docs/parser-architecture-comparison.md new file mode 100644 index 0000000..a4a3151 --- /dev/null +++ b/docs/parser-architecture-comparison.md @@ -0,0 +1,150 @@ +# Parser Architecture Comparison + +This note compares the parser families that matter most to this repo and names +the practical lessons worth carrying forward. + +The goal is not to pick a winner. The goal is to avoid borrowing the wrong +mental model from the wrong ecosystem. + +## The short answer + +This repo should stay closest to an event-stream-first parser with explicit +commitment points and tolerant defaults. + +That means: + +- not AST-first by default +- not pure Markdown-style fallback-to-text +- not a full HTML-spec tree-construction clone +- not a MediaWiki rendering engine hidden inside a source parser + +The right shape is a hybrid tuned for wikitext. + +## Comparison table + +| Family | Typical strengths | Where it breaks for wikitext | Useful lesson | +| --- | --- | --- | --- | +| Compiler parsers such as SWC, Oxc, Babel | Strong contracts, explicit recovery, precise spans, clear grammar ownership | Wikitext has more legacy ambiguity and more malformed-but-still-intended input | Keep commitment points, diagnostics, and offsets explicit | +| Markdown parsers such as cmark, micromark, goldmark, Comrak | Clean block-inline layering, event pipelines, extension discipline, spec or corpus-driven test culture | Markdown often tolerates unknown syntax by flattening it to text more aggressively than wikitext should | Keep layered parsing and corpus discipline, but do not import Markdown's fallback instinct wholesale | +| HTML parsers and html5lib-style conformance suites | Forgiving parsing after commitment, explicit tokenizer states, rich error categories | HTML's insertion modes and DOM rules are too specific to copy directly into wikitext | Use commitment-driven tolerance and error taxonomy ideas | +| MediaWiki core and Parsoid | Real wikitext correctness pressure, extension boundaries, round-trip expectations | Operationally heavy and too tied to full MediaWiki semantics to use as the core architecture model | Treat them as correctness and corpus references | + +## Why compiler architecture is only partially transferable + +Compiler parsers earn their simplicity from stronger grammars. + +JavaScript, Rust, or TypeScript parsers still have ambiguity, but they usually +know the space of valid forms ahead of time. Wikitext has more legacy syntax, +profile-specific behavior, extension-driven constructs, and malformed content +that users still expect to "work well enough." + +So the lesson from compiler parsers is not "use recursive descent everywhere" +or "turn the grammar into a formal parser generator input." The lesson is to +be explicit about these boundaries: + +- when structure becomes committed +- what counts as a factual parser finding +- what offsets and ranges mean +- which behavior is parser truth versus consumer policy + +That matches the repo's diagnostics-first direction. + +## Why Markdown architecture is useful but incomplete + +Markdown parsers are closer to this repo operationally because they often split +block and inline work and they often support event or token pipelines. + +micromark is especially relevant because it treats events as the primary +interchange layer and lets AST builders sit on top. That is already a strong +fit for this repo. + +But Markdown has one dangerous influence here: many implementations can safely +say "if we cannot prove this syntax, just keep the text." Wikitext cannot +always do that, because users often write malformed markup that is still +clearly trying to be a table, tag, link, or template. + +That is why this repo should keep the pipeline lesson from Markdown, but not +the full malformed-input philosophy. + +## Why HTML-style commitment matters more for malformed input + +HTML parsing gives a better model for the repo's tolerant default lane. + +The useful pattern is: + +```text +before commitment -> keep text-backed interpretation +after commitment -> preserve structural intent +malformed continue -> emit diagnostics and keep going +strict consumer -> optionally collapse back to text later +``` + +This repo should keep that shape for tags and, where it fits, for other +constructs with meaningful commitment points. + +The important part is discipline. A forgiving parser is not a guessing parser. +It still needs explicit rules for when an opener became real. + +## Why MediaWiki and Parsoid are the correctness references + +MediaWiki core and Parsoid are where the ecosystem has already paid for years +of parser mistakes. + +They matter less because their internals are pretty and more because their test +corpora capture: + +- real editorial habits +- malformed inputs that still need usable output +- extension interactions +- round-trip expectations +- regression history + +This repo should not clone their architecture. It should borrow pressure from +their corpora and use that pressure to validate its own simpler architecture. + +## Recommended architecture stance + +This is the working recommendation. + +### Core parser + +- Keep the tokenizer, block parser, and inline parser hand-written. +- Keep events as the primary interchange layer. +- Keep offsets and positions range-first. + +### Malformed-input model + +- Treat diagnostics as parser facts. +- Treat continuation as internal survivability logic. +- Treat final tree shape as materialization policy. + +### Public API philosophy + +- Default lane keeps tolerant structure. +- Strict lane stays conservative. +- Recovery summary is optional metadata, not a separate parser truth. + +### Extension philosophy + +- Prefer primitive-first composition over deep parser plugins. +- Keep extension boundaries explicit, especially for tag-like and parser- + function-like constructs. + +## What to test because of this comparison + +Architecture choices only matter if tests make them real. + +The most important tests implied by this comparison are: + +1. Block structure stays consistent across `outlineEvents()`, `events()`, and + `parse()` in the default tolerant lane. +2. A construct that never reaches its commitment point stays text-backed in all + lanes. +3. A construct that did commit may stay structural in the default lane and + collapse in the strict lane without changing diagnostics. +4. Extension-boundary constructs such as `` and parser functions stay + covered by their own reduced corpora instead of leaking into generic tests. +5. Later round-trip work should use Parsoid-style corpus pressure instead of + synthetic only-happy-path fixtures. + +The corresponding corpus plan lives in [docs/corpus-matrix.md](./corpus-matrix.md). \ No newline at end of file