From d87ff1497695f8e81348351987dfd1bcf9eafe52 Mon Sep 17 00:00:00 2001 From: Pierre Le Fevre Date: Sun, 17 May 2026 21:48:29 +0200 Subject: [PATCH] Add Phase 23: real-web soak & bug bash MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Random-picker model: a stdlib-only script intersects an upstream popular-sites ranking with upstream SFW and threat-intel blocklists, randomizes the remainder, and emits one URL per iteration. The repo holds only the picker, the per-site scenarios it accumulates, and their snapshots — no vendored domain lists. Co-Authored-By: Claude Opus 4.7 --- PLAN.md | 71 +++++++++++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 71 insertions(+) diff --git a/PLAN.md b/PLAN.md index d0e0ccc..9ac625b 100644 --- a/PLAN.md +++ b/PLAN.md @@ -404,6 +404,76 @@ The execution engine lives in a new `we-wasm` workspace crate that must be 100% **Test milestone:** Pass the WebAssembly spec test suite for the MVP (`address.wast`, `br.wast`, `call.wast`, `const.wast`, `conversions.wast`, `fac.wast`, `forward.wast`, `i32.wast`, `i64.wast`, `if.wast`, `int_literals.wast`, `loop.wast`, `memory.wast`, `names.wast`, `return.wast`, `select.wast`, `switch.wast`, `traps.wast` at minimum) plus the post-MVP proposals listed above. Run a real Rust-compiled wasm module end-to-end: fetch it over HTTP, instantiate it via `WebAssembly.instantiateStreaming`, call exports from JS, and observe host-imported callbacks. Extend WPT coverage for the `WebAssembly.*` JS API surface, extend `we-js` tests for `Promise`/microtask semantics used by `compile`/`instantiate`, and extend `we-e2e` scenarios for the end-to-end page. +## Phase 23: Real-Web Soak & Bug Bash + +**Goal:** Repeatedly point the browser at random, popular, known-safe real websites and fix every defect each one exposes. No new features in this phase — only diagnosis and repair. A picker script chooses one site per iteration by intersecting an upstream popular-sites ranking with upstream SFW and threat-intel blocklists; each defect becomes an `isu` issue and each fix lands as its own merge to `main`. Over many iterations the soak organically covers the entire safe popular web. + +### The picker (the only committed input-side machinery) + +`tests/popular-sites/pick.py` is the single committed script that produces candidate URLs. It is stdlib-only Python (no external deps), and runs each time the implementor needs a new site. + +What it does, per invocation: + +1. Fetch a popular-sites ranking from an upstream source (no local copy retained). +2. Fetch SFW and known-dangerous blocklists from upstream sources (no local copies retained). +3. Take the top N entries of the ranking (default N = 1000). +4. Remove every domain that appears in any blocklist, plus any domain registered within the last 30 days. +5. Randomize the remaining list and print one URL to stdout. + +Determinism is intentionally not a goal of the picker — randomness is the mechanism by which the soak eventually covers everything in the safe top N over many iterations. Reproducibility lives downstream: once a site is chosen and a scenario is committed, that scenario is fully deterministic against its committed snapshot. + +### Upstream sources (fetched at pick time, never committed) + +- Popularity ranking: the [Tranco list](https://tranco-list.eu/) (CC BY 4.0). Cloudflare Radar Top 1000 as fallback. +- SFW blocklists: an existing public SFW DNS-blocklist source such as one of the curated lists at [oisd.nl](https://oisd.nl/) or [StevenBlack/hosts](https://github.com/StevenBlack/hosts) (porn / gambling / fake-news variants). These already encode the categories we care about — adult, gambling, piracy, illegal-marketplace, hate-speech, shock — so we consume them rather than re-classify. +- Known-dangerous blocklists: [URLhaus](https://urlhaus.abuse.ch/) (malware, CC0), [OpenPhish](https://openphish.com/) (phishing), [PhishTank](https://phishtank.org/) (phishing), [Spamhaus DBL](https://www.spamhaus.org/dbl/) (known-bad zones). +- WHOIS / RDAP for the "registered in the last 30 days" heuristic. + +The picker treats every fetch as best-effort: if a feed is unreachable, it fails closed — refusing to emit a URL — rather than admitting candidates that might have been filtered. + +### Soak iteration workflow + +When the implementor finds the issue queue empty during Phase 23, it does not break a new phase down. Instead, it runs one soak iteration: + +1. Invoke `tests/popular-sites/pick.py` to obtain a URL. +2. Snapshot the site under `crates/e2e/real-web/snapshots//` (HTML + CSS + JS + assets fetched once). +3. Write a scenario at `crates/e2e/scenarios/real-web/.we` that loads the snapshot via `goto_as` + `cache_put`, takes a screenshot, dumps DOM and console, and asserts the page's primary content. +4. Run the scenario. For every defect surfaced (panic, layout glitch, missing glyph, JS exception, network misbehaviour, perf cliff), file an `isu` issue with the `real-web` label, the scenario path, and pointers to the captured artifacts. +5. Commit the scenario, snapshot, and golden screenshot — even if defects remain — so the next iteration of the implementor picks up the open defect issues via the normal Priority 1 path. +6. Exit. Subsequent iterations close defects one by one; when all defects for the site are closed, the scenario becomes `passing` and the soak naturally moves on to a new random pick the next time the queue is empty. + +If `pick.py` happens to return a domain that already has a committed scenario, re-snapshot it against the live site: any divergence is either a real-web regression on our end (file an issue) or upstream drift (update the snapshot and golden as its own commit). + +### Methodology + +- One scenario per site, named after the domain. Each scenario does at minimum: `goto_as `, `cache_put` for subresources, `screenshot`, `dump_dom`, `dump_console`, plus targeted `assert_dom_contains` / `assert_console_contains` for the page's key content. +- Snapshots are committed so the scenario suite is offline-runnable and deterministic. Re-snapshot only when intentionally bumping a baseline, in its own commit. +- Golden screenshots live next to each scenario as `.expected.png`. Diff against the rendered output with a per-pixel tolerance helper added to the e2e harness; any diff over threshold fails the scenario. +- Each fix lands in its own merge to `main`, references the `isu` issue, and includes (a) the scenario asserting the bug is gone and (b) any narrower unit test that catches the regression at a lower layer. + +### Bug intake & triage + +- New `isu` label set for this phase: `real-web` (any defect found by the soak), plus existing layer labels (`html`, `css`, `layout`, `render`, `js`, `net`, `text`, `image`, `wasm`, `e2e-smoke`). +- Severity labels: `sev-crash` (panic, hang, segfault, infinite loop), `sev-broken` (page does not display its primary content), `sev-visual` (visible but wrong), `sev-perf` (correct but unusably slow). +- Fix order: `sev-crash` first across all sites, then `sev-broken`, then `sev-visual`, then `sev-perf`. Within a severity, prefer fixes that unblock the most committed scenarios. +- A bug that requires a new feature (e.g. an unimplemented CSS property a site relies on) is split: file the feature request as its own issue and mark the scenario `xfail` with a pointer to the tracking issue until the feature lands. + +### Harness extensions + +- Add a real-web runner: `cargo run -p we-e2e -- --real-web` walks every scenario under `crates/e2e/scenarios/real-web/`, in offline snapshot mode by default, and writes a summary report (`crates/e2e/artifacts/real-web/report.json`) listing per-site pass/fail, screenshot diff %, and any captured panic. +- Add a panic catcher to the e2e harness so a single scenario crashing does not abort the rest of the run; failed scenarios produce a stack trace and a partial artifact set. +- Extend the harness with a perf budget: per-scenario wall-clock load time and peak resident set size, captured into the report; regressions over a configurable threshold fail the scenario under `sev-perf`. + +### Exit criteria + +- Coverage: at least the top 100 filtered popular sites have committed `passing` scenarios under `crates/e2e/scenarios/real-web/`. +- Stability: a sustained streak (e.g. 25 consecutive random picks) of new sites that pass on first scenario run with no `sev-crash` or `sev-broken` defects. +- Every `real-web`-labelled `isu` issue is either closed by a fix or explicitly re-categorised as a feature request scheduled for a later phase. +- `cargo run -p we-e2e -- --real-web` is green in CI and runs as part of the standard pre-merge checklist for any change to parsing, styling, layout, rendering, scripting, networking, or the UA stylesheet. +- A short retrospective in `crates/e2e/real-web/README.md` lists the bugs found, the fixes that landed, and any features the soak motivated promoting to a future phase. + +**Test milestone:** The random soak runs for an extended period without producing new `sev-crash` or `sev-broken` defects; the committed scenarios cover the top 100 filtered popular sites and all pass green offline. + --- ## Test Suites @@ -416,6 +486,7 @@ The execution engine lives in a new `we-wasm` workspace crate that must be 100% | Test262 (core language) | `js` | 10 | | WPT (progressive) | all | 11+ | | WebAssembly spec tests | `wasm` | 22 | +| Real-web soak (offline snapshots + online re-baseline) | `e2e` | 23 | | E2E smoke (headless screenshots + scripted scenarios) | `e2e` | continuous | ### E2E Harness -- 2.51.2