From 26f638f92f3cc3a4f513ae5bb59aeda4bb9f0e83 Mon Sep 17 00:00:00 2001 From: Okiki Ojo Date: Sat, 25 Apr 2026 21:34:03 -0400 Subject: [PATCH] docs(architecture): update parser-result and API direction notes for clarity Signed-off-by: Okiki Ojo --- docs/architecture/README.md | 9 +- docs/architecture/api-direction.md | 65 +++--- docs/architecture/choosing-a-parser-result.md | 70 ++++--- docs/architecture/malformed-input.md | 191 ++++++++++++------ 4 files changed, 204 insertions(+), 131 deletions(-) diff --git a/docs/architecture/README.md b/docs/architecture/README.md index 45035fd..cae4280 100644 --- a/docs/architecture/README.md +++ b/docs/architecture/README.md @@ -3,7 +3,7 @@ This folder breaks the larger architecture story into smaller documents with one job each. -If you are new to the repo, start with the tree-choice note first. It explains +If you are new to the repo, start with the parser-result note first. It explains the user-facing parser outputs before the deeper pipeline details. This folder is not the package introduction and it is not the symbol lookup @@ -16,9 +16,10 @@ reference. ## Start here -- [choosing-a-tree.md](./choosing-a-tree.md) +- [choosing-a-parser-result.md](./choosing-a-parser-result.md) Explains the practical caller question: which parser result should I ask for, - why, and what trade-off each lane makes. + why, and what trade-off each lane makes, including the planned + `analyze()` findings-first lane and the still-exploratory policy lane. - [pipeline.md](./pipeline.md) Explains the tokenizer -> block parser -> inline parser -> consumer flow in plain English. @@ -29,7 +30,7 @@ reference. Explains the diagnostics-first model without the full redesign note. - [api-direction.md](./api-direction.md) Explains the likely future public-surface cleanup without turning the main - tree-choice note into an API design memo. + parser-result note into an API design memo. - [diagnostic-anchors.md](./diagnostic-anchors.md) Explains how diagnostics point back into the final tree. - [parser-contracts.md](./parser-contracts.md) diff --git a/docs/architecture/api-direction.md b/docs/architecture/api-direction.md index 046497f..3b20fea 100644 --- a/docs/architecture/api-direction.md +++ b/docs/architecture/api-direction.md @@ -41,13 +41,17 @@ parseWithRecovery() parseStrictWithDiagnostics() -> conservative tree + diagnostics -``` -That surface is already more honest than the older docs, but it still mixes two -different questions: +analyze() + -> findings (events + diagnostics + recovery list), no tree + +materialize(findings, { policy? }) + -> tree + diagnostics under one materialization policy +``` -1. do you want diagnostics? -2. how should malformed regions be materialized? +The tree-first wrappers still mix two questions (diagnostics and +materialization), but `analyze()` + `materialize()` now separate them +cleanly for callers who want that split. ## The direction the surface is moving toward @@ -98,25 +102,21 @@ away from source-backed text interpretations. Convenience wrappers could still exist on top of that base surface. -## Lane 3: concrete `analyze()` proposal +## Lane 3: shipped `analyze()` + `materialize()` API -There is one stronger diagnostics-first lane that the public API does not fully -offer yet. +The diagnostics-first lane is now public: ```text diagnostics and recovery data are exposed -but final materialization is delayed or caller-owned +final materialization is delayed or caller-owned via materialize() ``` -That would go beyond `parseStrictWithDiagnostics()`. `parseStrictWithDiagnostics()` still chooses a final tree -policy for the caller. A true `analyze()` lane would expose parser findings -first and let later tools decide which repairs to apply. - -That kind of lane is especially compatible with a range-first design because it -lets the parser preserve diagnostics and source spans now, then defer heavier -materialization choices until a caller actually needs them. +That goes beyond `parseStrictWithDiagnostics()`, which still chooses a final +tree policy for the caller. `analyze()` exposes parser findings first, and +`materialize()` is the explicit step that turns those findings into one +tree under a named policy. -One practical shape could look more like this: +The shipped shape looks like this: ```ts interface AnalyzeOptions { @@ -127,8 +127,12 @@ interface ParseRecovery { readonly kind: | 'missing-close' | 'unterminated-opener' - | 'eof-autoclose' - | 'mismatched-exit'; + | 'unclosed-table' + | 'mismatched-exit' + | 'orphan-exit' + | 'eof-autoclose'; + readonly code: string; + readonly position: Position; readonly anchor: ParseDiagnosticAnchor; readonly node_type?: WikistNodeType; readonly policies: readonly TreeMaterializationPolicy[]; @@ -149,26 +153,12 @@ function materialize( ): ParseOutput; ``` -The important part is the ownership model: +The ownership model is explicit: - the parser exposes replayable findings - the caller chooses whether to materialize them at all - if the caller does materialize them, the tree policy is explicit -That is why a plain one-shot generator is probably not enough as the whole -public lane. The event stream should stay the core primitive, but many callers -will need to inspect diagnostics and materialize more than once without -rerunning the parse. - -The main design bet is that findings should stay narrow and factual. A useful -first public version should probably include: - -- replayable events -- diagnostics -- a small, parser-owned recovery vocabulary - -It should not start by exposing a broad plugin or callback system. - ## Lane 4: exploratory policy proposal There is one possible lane after that, but it should stay explicitly more @@ -211,10 +201,11 @@ Lane 3 is easier to keep stable because it mostly exposes what the parser already knows. Lane 4 is harder because it starts exposing when and how those findings are turned into structure. -The likely order is: +The likely remaining order is: -1. make the `analyze()` lane real -2. make the package-owned materializers explicit +1. ~~make the `analyze()` lane real~~ (shipped) +2. ~~make the package-owned materializers explicit~~ (shipped via + `TreeMaterializationPolicy`) 3. only then decide whether a caller-owned policy lane belongs in public API That is the kind of thing I meant earlier by an "API proposal doc": a short diff --git a/docs/architecture/choosing-a-parser-result.md b/docs/architecture/choosing-a-parser-result.md index 4739a37..590da88 100644 --- a/docs/architecture/choosing-a-parser-result.md +++ b/docs/architecture/choosing-a-parser-result.md @@ -20,8 +20,9 @@ exploratory policy layer, and those should stay clearly separated. 4. A policy path, still exploratory, where the caller wants to make some of those later materialization decisions itself. -The first three paths are the core story. The fourth path is worth exploring, -but it should not be treated as equally settled yet. +The first three paths are the core story, and all three are shipped today. The +fourth path is worth exploring, but it should not be treated as equally +settled yet. ## The simple version @@ -146,21 +147,35 @@ do not apply the recovery for me yet I may want more than one materialization from the same parse ``` -What this path should give: +What this path gives you today: - diagnostics preserved in the result - replayable parser findings, not just a one-shot event generator - recovery data listed explicitly - no recovery applied on the caller's behalf -- enough information for the caller to choose which repairs to keep, - discard, or replace later +- enough information to choose which repairs to keep, discard, or replace + later at materialization time -The event stream is still the right primitive underneath this path, but a bare -generator is probably too weak as the public shape. A control-heavy caller may -need to inspect diagnostics, compare recoveries, and materialize more than one -final tree without reparsing the same source every time. +Current API: -One concrete target shape could look like this. +```ts +import { analyze, materialize, TreeMaterializationPolicy } from '@okikio/wikitext'; + +const findings = analyze(source); + +// inspect without materializing a tree +for (const entry of findings.recovery ?? []) { + console.log(entry.kind, entry.code, entry.policies); +} + +// materialize once or many times from the same findings +const tolerant = materialize(findings); +const strict = materialize(findings, { + policy: TreeMaterializationPolicy.SOURCE_STRICT, +}); +``` + +The shapes live in `parse.ts`: ```ts interface AnalyzeOptions { @@ -171,8 +186,12 @@ interface ParseRecovery { readonly kind: | 'missing-close' | 'unterminated-opener' - | 'eof-autoclose' - | 'mismatched-exit'; + | 'unclosed-table' + | 'mismatched-exit' + | 'orphan-exit' + | 'eof-autoclose'; + readonly code: string; + readonly position: Position; readonly anchor: ParseDiagnosticAnchor; readonly node_type?: WikistNodeType; readonly policies: readonly TreeMaterializationPolicy[]; @@ -196,7 +215,7 @@ function materialize( ): ParseOutput; ``` -That shape keeps the ownership split clear: +The ownership split is explicit: - `analyze()` owns parser facts - `materialize()` owns tree policy @@ -356,27 +375,28 @@ structures. ## Where the current public API fits today -The current wrappers do not map perfectly onto the intended path model yet, but -this is the rough shape: +The current wrappers now map fairly cleanly onto the path model: -- `parse()`, `parseWithDiagnostics()`, and `parseWithRecovery()` are the current +- `parse()`, `parseWithDiagnostics()`, and `parseWithRecovery()` are the default-tree family wrappers -- `parseStrictWithDiagnostics()` is the current conservative tree wrapper -- `analyze()` is the intended third path, but it is not the final public API yet -- `events(source, { diagnostics: true })` and session caches are the closest - current low-level building blocks for a future `analyze()` lane +- `parseStrictWithDiagnostics()` is the conservative tree wrapper +- `analyze()` and `materialize()` are the findings-first lane +- `createSession(source).analyze()` and `session.materialize()` are the same + lane backed by the session cache, so multiple materializations and repeated + findings lookups do not reparse the source -So this page should be read as the intended decision model the docs are trying -to clarify, not as a claim that every current wrapper already matches that -model perfectly. +The policy lane (Path 4) is still exploratory; `materialize()` currently +accepts only the package-owned `DEFAULT_HTML_LIKE` and `SOURCE_STRICT` +policies. ## Which one should most people use? - Use the default tree family when you want good defaults and a usable tree now. - Use the conservative tree when diagnostics matter but you do not want applied recovery to survive in the final tree. -- Use the planned `analyze()` lane when you want to treat recovery as a later - caller decision instead of an already-applied parser policy. +- Use `analyze()` + `materialize()` when you want to treat recovery as a later + caller decision instead of an already-applied parser policy, or when you + want more than one materialization from the same parse. - Treat the policy lane as an advanced follow-on design, not the default next step. diff --git a/docs/architecture/malformed-input.md b/docs/architecture/malformed-input.md index cbecb93..26f5932 100644 --- a/docs/architecture/malformed-input.md +++ b/docs/architecture/malformed-input.md @@ -2,21 +2,21 @@ This note explains the parser's malformed-input model in plain English. -The short version is: +The short version is simple: -the parser never throws, but it also should not treat every malformed thing the +the parser never throws, but it also should not treat every malformed span the same way. The key question is how much evidence the source gave for a real structure, and -then how much repair the chosen tree family should apply. +then which caller-facing path should decide what happens next. One grounding detail matters here too: -when this note says "text-backed" or "plain text," it does not mean the parser -must eagerly copy out a new string value. +when this note says text-backed or plain text, it does not mean the parser must +eagerly copy out a new string value. In this repo, text usually stays range-first for as long as possible. That -means malformed regions can remain attached to the original source by offsets +means malformed regions can stay attached to the original source by offsets instead of being rewritten into fresh strings too early. That matters because the parser is trying to preserve: @@ -34,19 +34,20 @@ The easiest way to reason about malformed input is to split it into two phases. ```text 1. did the source give enough evidence for a real structure? -2. if yes, how much of that structure should survive in the final tree? +2. if yes, which caller-facing path should decide how much of that structure + survives? ``` -That is better than starting with the word "recovery," because recovery only +That is better than starting with the word recovery, because recovery only matters after the parser has decided that a real structure or a real recovery -candidate exists at all. +path exists at all. -## Weak evidence: stay text-backed +## Weak evidence stays text-backed If the source only weakly hints at a construct, the parser should stay conservative. -Here, "text-backed" means the parser keeps the region as a text interpretation +Here, text-backed means the parser keeps the region as a text interpretation rooted in the original source range instead of inventing a stronger structural node too early. @@ -56,19 +57,19 @@ Example: 2 < 3 and ref name="x" ``` -There is not enough real tag structure here to justify recovery. +There is not enough real tag structure here to justify a recovered node. -Current tree-lane result: +Current tree results: - `parse()` keeps a text-backed interpretation of that range - `parseWithDiagnostics()` keeps the same text-backed interpretation and reports the problem -- `parseStrict()` keeps the same text-backed interpretation and reports the +- `parseStrictWithDiagnostics()` keeps the same text-backed interpretation and reports the problem -## Strong evidence: preserve structural intent by default +## Strong evidence preserves structural intent by default -If the source clearly points at a real construct, the default tolerant lane +If the source clearly points at a real construct, the default tree family should usually preserve that structure and attach diagnostics instead of flattening it immediately. @@ -80,20 +81,21 @@ Paragraph with note The opener did reach `>`, so the parser now has a committed structural finding. -Current tree-lane result: +Current tree results: - `parse()` may keep a `reference` node without returning diagnostics -- `parseWithDiagnostics()` keeps the default tolerant `reference` node and - returns diagnostics -- `parseStrict()` may collapse the same region back to text while keeping the +- `parseWithDiagnostics()` keeps the same default `reference` node and returns + diagnostics +- `parseWithRecovery()` keeps that same tree and adds a recovery summary +- `parseStrictWithDiagnostics()` may collapse the same region back to text while keeping the same diagnostics That is the main HTML-like part of the design. ## Boundary cases are policy questions -Some malformed inputs sit right on the border between "text" and "recoverable -structure." +Some malformed inputs sit right on the border between text and recoverable +structure. Example: @@ -104,63 +106,108 @@ Paragraph with ` before the tolerant -lane may keep it structurally real, that is still a valid policy. The important -thing is that this is a recovery-policy decision, not something the docs should -accidentally present as an unquestionable parser fact. +If the project later decides that an opener must reach `>` before the default +lane may keep it structurally real, that is still a valid policy. The +important point is that this is a path-policy decision, not something the docs +should accidentally present as an unquestionable parser fact. That is also why range-first text matters here. A text-backed interpretation is not a dead end. The parser can keep the exact source range intact now, attach a diagnostic to it, and still let a later materializer or caller decide whether that same span should stay text or become a repaired structure. -## The three current caller-facing lanes +## The current tree lanes and the planned `analyze()` lane -### Cheap tree +The docs should not talk as if the third option is another tree. -```ts -parse(source) -``` +The event stream and diagnostics are already the lower-level facts. The third +path should be `analyze()`: a replayable findings layer that can later be +materialized into one tree policy or another. -- cheapest tree lane -- no preserved diagnostics -- still never-throw -- still uses the default tolerant materialization +There is also one more possible layer after that: a caller-owned policy lane +built on top of those same findings. That is more exploratory, because it +would freeze more of the recovery model. -### Tolerant tree with diagnostics +### Default tree family ```ts +parse(source) parseWithDiagnostics(source) parseWithRecovery(source) ``` -- same default tolerant tree family as `parse()` -- diagnostics preserved -- optional `recovered` summary boolean from `parseWithRecovery()` +- applies the parser's default recovery behavior when the source gave strong + enough structural evidence +- keeps good defaults by default +- optionally preserves diagnostics, depending on which wrapper you call +- still never throws + +The split inside that family is narrow: -### Conservative tree with diagnostics +- `parse()` gives the cheapest default tree +- `parseWithDiagnostics()` gives that same tree plus diagnostics +- `parseWithRecovery()` gives that same tree plus diagnostics and a `recovered` + summary boolean + +### Conservative tree ```ts -parseStrict(source) +parseStrictWithDiagnostics(source) ``` - diagnostics preserved +- no applied recovery in the final tree - more malformed committed regions collapsed back to source-backed text -Again, "source-backed text" means the final tree prefers a text node that still +Again, source-backed text means the final tree prefers a text node that still points back to the original source span instead of preserving a more tolerant repaired wrapper. +### `analyze()` findings lane + +```ts +analyze(source) +``` + +- diagnostics preserved +- event-level findings preserved in a replayable form +- recovery data exposed without applying it on the caller's behalf +- final materialization deferred to an explicit later step + +That is more diagnostics-first than `parseStrictWithDiagnostics()`. `parseStrictWithDiagnostics()` still +chooses one conservative tree policy for the caller. + +### Policy lane + +```ts +const findings = analyze(source, { recovery: true }); + +materialize(findings, { + policy: 'custom', + resolve_recovery(recovery) { + return recovery.node_type === 'reference' + ? 'keep-structural' + : 'collapse-to-text'; + }, +}); +``` + +- starts from the same `analyze()` findings lane +- lets the caller choose some recoveries itself +- still produces a final tree through a later materialization step + +That is not the same as inventing a fourth parser truth. It is a more advanced +consumer policy layer. + ## What the parser is and is not doing on the caller's behalf This is where the docs have been too fuzzy in the past. @@ -169,27 +216,26 @@ The parser always does some continuation work internally. It has to, otherwise it could not keep the event stream well-formed or uphold the never-throw contract. -So the real caller choice is not: +So the real caller choice is not this: ```text parser continuation vs no parser continuation ``` -The real caller choice is closer to: +The real caller choice is closer to this: ```text -do I want diagnostics preserved? -and if so, which final tree policy do I want? +do I want the default tree family, the conservative tree, or analyze() first? +and if I want a tree, do I also want diagnostics preserved? ``` -That is why `parse()` is cheap, but it is not a "no parser help at all" mode. +That is why `parse()` is cheap, but it is not a no-parser-help-at-all mode. -It is also why "flatten back to text" should be read carefully in this repo. -Usually that means "prefer the original source span as text material" rather -than "throw away structural knowledge and allocate a brand new replacement -string." +It is also why flatten back to text should be read carefully in this repo. +Usually that means prefer the original source span as text material rather than +throw away structural knowledge and allocate a brand new replacement string. -## The still-open fourth lane +## The still-open public shape for `analyze()` There is one more possibility that the current docs should name explicitly. @@ -197,33 +243,48 @@ It is not fully a public lane yet, but it is the next important design question: ```text -diagnostics and possible recoveries are exposed +diagnostics and recovery data are exposed but the caller chooses the final materialization later ``` -That would be more diagnostics-first than `parseStrict()`. It would expose the -parser's findings without forcing either the default tolerant tree or the +That would be more diagnostics-first than `parseStrictWithDiagnostics()`. It would expose the +parser's findings without forcing either the default recovered tree or the conservative tree as the final answer. +The practical shape should probably be a findings utility or cached session +lane, not a bare generator alone. + +A one-pass event generator is still a useful primitive, but it is weak as the +whole public answer for control-heavy callers because they may need to: + +- inspect diagnostics +- compare recoveries +- materialize more than one tree policy from the same parse +- keep findings around in a session or streaming workflow + That is likely the right future home for the most control-heavy use cases. For now, the closest current tools are: -- `events(source, { include_diagnostics: true })` for event-level findings -- `parseStrict(source)` for the conservative tree lane +- `events(source, { diagnostics: true })` for low-level event facts + today +- `createSession(source).events({ diagnostics: true })` when replay and + caching matter +- `parseStrictWithDiagnostics(source)` for the conservative tree lane today ## Why this is closer to HTML than to Markdown -Many Markdown parsers can say, "if we cannot prove it, keep it as text," and +Many Markdown parsers can say, if we cannot prove it, keep it as text, and still feel natural. Wikitext often needs a more forgiving default. If a user clearly started a real tag, table, or similar construct and only later broke the syntax, flattening it immediately can hide useful structural intent. -That is why the default lane here stays closer to HTML-like tolerant parsing: +That is why the default tree family here stays closer to HTML-like tolerant +parsing: - do not guess too early - do not commit too early - but once the source clearly committed, preserve that intent unless the caller - asks for a more conservative tree \ No newline at end of file + asks for the conservative tree instead \ No newline at end of file -- 2.51.2