diff --git a/PLAN.md b/PLAN.md index a045eb5..68ca844 100644 --- a/PLAN.md +++ b/PLAN.md @@ -1,7 +1,7 @@ # AArch64 ELF-to-WebAssembly Browser Runtime ## Engineering Build Plan and Agent Handoff -**Status:** Phase 3 complete; Phase 4 hot-block Wasm dispatch implemented for the initial scalar subset +**Status:** Phase 4 complete; Phase 5 robust CPython environment is next **Primary implementation language:** Rust **Initial browser target:** Google Chrome **Guest architecture:** AArch64, little-endian, Linux userspace @@ -1863,7 +1863,9 @@ The first Phase 4 checkpoint is implemented. The interpreter identifies basic-bl The initial basic-block lowering checkpoint is also implemented. Decoder output now crosses into the project-owned `binarrow-execution-ir` crate, where blocks are validated as non-empty, contiguous single-entry sequences with a final terminator and scratch storage no longer exposes Icicle types. `binarrow-wasm-backend` lowers its scalar Tier-1 subset against an explicit imported state-memory ABI, emits deterministic modules and translation metrics, and returns structured unsupported-operation results for interpreter fallback. A checked-in AArch64 scalar-loop fixture gives Chromium a real decoded block to compile dynamically and compare against interpreter state. Hot-dispatch integration, broader operation coverage, and the translation cache followed from this checkpoint. -Hot-dispatch integration is now implemented for the initial scalar subset. The interpreter offers blocks to a backend after a deterministic execution threshold, charges translated instructions against the same process budget, applies replacement architectural state atomically, and remembers unsupported block identities for interpreter fallback. The browser backend compiles and caches core Wasm modules by guest address plus instruction encodings, installs their exported entry functions in a bounded Worker-local WebAssembly table, reuses a module-local imported state memory, and exposes translated-block, translated-instruction, fallback, compilation, cache-hit, and emitted-byte counters. The Chromium infinite-loop regression executes 64 cold iterations in the interpreter and the next 64 through one generated function-table entry with 63 cache hits. Code-identity regressions replace executable bytes at the same guest address and prove that both the fallback set and compiled-module cache select a new identity. The dynamic translation probe initializes interpreter and Wasm execution from identical state, then compares the complete scalar state ABI and next PC across a 17-instruction block covering normal and exceptional signed/unsigned division, shifts, and bitwise operations. Structured forward internal branches lower to Wasm blocks while malformed control retains interpreter fallback. Broader operation coverage, a larger differential corpus, and benchmark evidence remain before Phase 4 is complete. +Hot-dispatch integration is implemented for the initial scalar subset. The interpreter offers blocks to a backend after a deterministic execution threshold, charges translated instructions against the same process budget, applies replacement architectural state atomically, and remembers unsupported block identities for interpreter fallback. The browser backend compiles and caches core Wasm modules by guest address plus instruction encodings, installs their exported entry functions in a bounded Worker-local WebAssembly table, reuses a module-local imported state memory, and exposes translated-block, translated-instruction, fallback, compilation, cache-hit, and emitted-byte counters. The Chromium infinite-loop regression executes 64 cold iterations in the interpreter and the next 64 through one generated function-table entry with 63 cache hits. Code-identity regressions replace executable bytes at the same guest address and prove that both the fallback set and compiled-module cache select a new identity. The dynamic translation probe initializes interpreter and Wasm execution from identical state, then compares the complete scalar state ABI and next PC across a 17-instruction block covering normal and exceptional signed/unsigned division, shifts, and bitwise operations. Structured forward internal branches lower to Wasm blocks while malformed control retains interpreter fallback. + +Phase 4 is complete for the planned Tier-1 scope. An opt-in Chromium benchmark runs the 17-instruction integer-mix loop through fresh interpreter-only and hot-translation sessions, warms both paths, takes five samples, and compares their medians without burdening normal CI with a timing assertion. A July 24, 2026 local run over 278,528 guest instructions measured 36.8 ms interpreted and 24.5 ms translated, a 1.50x end-to-end speedup including scalar state exchange and function-table dispatch. The benchmark requires at least a 1.2x speedup when explicitly enabled. Broader operation coverage and a larger differential corpus remain ongoing quality work, but every Phase 4 deliverable and acceptance criterion now has an implementation and regression or reproducible benchmark. Phase 5 begins with multi-file CPython programs, filesystem-backed imports, exceptions/tracebacks, pure-Python package installation, and compatibility documentation. Do not begin the full web IDE before item 30 passes. diff --git a/README.md b/README.md index 84c9f93..bb7de68 100644 --- a/README.md +++ b/README.md @@ -75,6 +75,18 @@ and packaging details. The browser build generates its Memory64, JSPI, and P-code `.wasm` probes before starting Vite. Generated artifacts are not committed. Select **Uploaded AArch64 ELF** to run an external static executable with a chosen `argv[0]` and one argument per line; the executable is transferred directly to the runtime Worker. +Run the opt-in Phase 4 interpreter/translator benchmark in Chromium with all +temporary output inside the repository: + +```sh +TMPDIR="$PWD/.tmp/rust-tmp" BINARROW_TRANSLATION_BENCHMARK=1 \ + pnpm --dir web test tests/translation-benchmark.spec.ts +``` + +The default benchmark warms both paths and reports the median of five +278,528-instruction samples. Set `BINARROW_TRANSLATION_ITERATIONS` to change the +loop count. + ## Scope The browser controller can start the checked-in freestanding C, musl C, Rust `std`, filesystem, and infinite-loop fixtures or a user-supplied static AArch64 ELF with explicit instruction, syscall, output, committed-memory, and filesystem limits. Execution errors return stable diagnostic codes instead of rejected JavaScript calls. The Stop action terminates the active Worker, so even a guest that never reaches a syscall or yield point can be interrupted and the runtime restarted. Chromium verifies the C/Rust outputs, uploaded-ELF transfer, in-memory file round trip, deterministic counters, resource-limit diagnostic, and manual infinite-loop termination. Phase 4's bounded profiler offers hot blocks to the scalar Wasm backend; supported blocks compile and execute from a session-local module cache while unsupported blocks remain in the interpreter. The UI reports hotness, translated execution, fallback, compilation, cache-hit, and emitted-byte metrics. The interpreter's interim browser memory design is recorded in [ADR-0002](docs/decisions/0002-use-sparse-memory-for-browser-interpreter.md). See [PLAN.md](PLAN.md) for the roadmap and [docs/architecture.md](docs/architecture.md) for the current boundaries. diff --git a/crates/browser-runtime/src/lib.rs b/crates/browser-runtime/src/lib.rs index 077f4fb..7a99b9f 100644 --- a/crates/browser-runtime/src/lib.rs +++ b/crates/browser-runtime/src/lib.rs @@ -586,6 +586,7 @@ pub struct BrowserGuestSession { terminal: CapturedTerminal, system: DeterministicSystem, block_executor: BrowserBlockExecutor, + hot_block_threshold: u64, startup_failure: Option, finished: bool, } @@ -623,7 +624,7 @@ impl BrowserGuestSession { &mut self.system, &mut host_input, &mut self.block_executor, - HOT_BLOCK_THRESHOLD, + self.hot_block_threshold, MAX_TRANSLATED_BLOCK_INSTRUCTIONS, ); let (outcome, diagnostic_code, diagnostic_message, exit_code, input_max_bytes) = @@ -678,6 +679,36 @@ pub fn start_fixture( max_memory_bytes: u64, max_filesystem_bytes: u64, filesystem_snapshot: &[u8], +) -> BrowserGuestSession { + start_fixture_with_hot_threshold( + fixture_name, + instruction_budget, + syscall_budget, + max_output_bytes, + max_memory_bytes, + max_filesystem_bytes, + filesystem_snapshot, + HOT_BLOCK_THRESHOLD, + ) +} + +/// Create a fixture session with an explicit translation threshold. +/// +/// A zero threshold disables translated dispatch. This entry point exists for +/// reproducible interpreter/translator benchmarks; normal execution should use +/// [`start_fixture`]. +#[wasm_bindgen] +#[must_use] +#[allow(clippy::too_many_arguments)] +pub fn start_fixture_with_hot_threshold( + fixture_name: &str, + instruction_budget: u64, + syscall_budget: u64, + max_output_bytes: u64, + max_memory_bytes: u64, + max_filesystem_bytes: u64, + filesystem_snapshot: &[u8], + hot_block_threshold: u64, ) -> BrowserGuestSession { let request = fixture(fixture_name) .map(|(elf, argv0)| (elf, vec![argv0.to_vec()])) @@ -697,6 +728,7 @@ pub fn start_fixture( max_filesystem_bytes, ), filesystem_snapshot, + hot_block_threshold, ) } @@ -728,6 +760,7 @@ pub fn start_program( max_filesystem_bytes, ), filesystem_snapshot, + HOT_BLOCK_THRESHOLD, ) } @@ -766,6 +799,7 @@ fn start_guest_session( request: Result<(&[u8], Vec>), BrowserExecution>, config: ProcessConfig, filesystem_snapshot: &[u8], + hot_block_threshold: u64, ) -> BrowserGuestSession { let max_filesystem_bytes = config.limits.max_filesystem_bytes; let filesystem = if filesystem_snapshot.is_empty() { @@ -810,6 +844,7 @@ fn start_guest_session( terminal: CapturedTerminal::default(), system: host_system(), block_executor: BrowserBlockExecutor::default(), + hot_block_threshold, startup_failure, finished: false, } diff --git a/docs/architecture.md b/docs/architecture.md index 3707946..958e242 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -100,4 +100,4 @@ The Icicle feasibility spike is isolated under `experiments/icicle`; it is not a `experiments/icicle-wasm` separately validates Icicle's lightweight `pcode` crate in raw `wasm32-unknown-unknown`. The production `binarrow-aarch64` crate pins `pcode`, `sleigh-runtime`, and the filesystem-free portion of `sleigh-compile` at that same revision and keeps their types private. The required AHash/getrandom browser backend is selected in the workspace's target configuration; `icicle-cpu` and its native VM/JIT remain outside the production graph. -Phase 3's VFS, OPFS persistence, host services, image packaging, and native/browser CPython checkpoints are complete. Phase 4 now has bounded profiling, scalar basic-block lowering, browser-side dynamic compilation, function-table hot dispatch, interpreter fallback, a session-local module cache, code-identity invalidation tests, and observable translation metrics. The differential probe compares the complete scalar state ABI and next PC after interpreting and translating the same 17-instruction arithmetic/control block. Broader lowering and differential coverage plus performance evidence remain. +Phase 4 is complete for the initial Tier-1 scope: bounded profiling, scalar basic-block lowering, browser-side dynamic compilation, function-table hot dispatch, interpreter fallback, a session-local module cache, code-identity invalidation tests, and observable translation metrics are all integrated. The differential probe compares the complete scalar state ABI and next PC after interpreting and translating the same 17-instruction arithmetic/control block. An opt-in five-sample Chromium benchmark measures warmed interpreter-only and translated sessions; the recorded 278,528-instruction checkpoint is 1.50x faster through translation. Broader lowering and differential coverage remain ongoing work while Phase 5 shifts to the robust CPython environment. diff --git a/web/src/main.ts b/web/src/main.ts index 9300da5..2e88f72 100644 --- a/web/src/main.ts +++ b/web/src/main.ts @@ -218,6 +218,9 @@ function createWorker(): void { setControls(); return; } + if (event.data.kind === "translation-benchmark") { + return; + } if (event.data.kind === "filesystem-image") { if (event.data.requestId !== imageRequestId) { return; diff --git a/web/src/probe.ts b/web/src/probe.ts index de87c01..af42b9b 100644 --- a/web/src/probe.ts +++ b/web/src/probe.ts @@ -68,6 +68,16 @@ export interface ExecutionReport { inputMaxBytes: bigint; } +export interface TranslationBenchmarkReport { + iterations: number; + guestInstructions: bigint; + interpretedMilliseconds: number; + translatedMilliseconds: number; + speedup: number; + translatedBlocks: bigint; + fallbackBlocks: bigint; +} + export type WorkerCommand = | { kind: "run"; @@ -100,6 +110,11 @@ export type WorkerCommand = maxFilesystemBytes: bigint; mode: "replace" | "install"; snapshot: Uint8Array; + } + | { + kind: "translation-benchmark"; + requestId: number; + iterations: number; }; export type WorkerMessage = @@ -116,4 +131,9 @@ export type WorkerMessage = operation: "export" | "replace" | "install"; snapshot: Uint8Array; } + | { + kind: "translation-benchmark"; + requestId: number; + report: TranslationBenchmarkReport; + } | { kind: "error"; message: string }; diff --git a/web/src/probe.worker.ts b/web/src/probe.worker.ts index 17c3190..0e7f1cc 100644 --- a/web/src/probe.worker.ts +++ b/web/src/probe.worker.ts @@ -10,6 +10,7 @@ import initBrowserRuntime, { install_filesystem_snapshot, normalize_filesystem_snapshot, start_fixture, + start_fixture_with_hot_threshold, start_program, translate_fixture_entry, } from "./generated/binarrow_browser_runtime.js"; @@ -21,6 +22,7 @@ import type { FeatureResult, FixtureName, RunTarget, + TranslationBenchmarkReport, WorkerCommand, WorkerMessage, } from "./probe"; @@ -49,6 +51,7 @@ interface JspiApi { } const filesystemSnapshotName = "binarrow-project-v2.snapshot"; +const translatedLoopInstructions = 17n; async function loadFilesystemSnapshot(): Promise { const root = await navigator.storage.getDirectory(); @@ -348,6 +351,89 @@ async function handleFilesystemCommand( self.postMessage(message); } +interface BenchmarkSample { + milliseconds: number; + guestInstructions: bigint; + translatedBlocks: bigint; + fallbackBlocks: bigint; +} + +function runBenchmarkSample( + iterations: number, + hotBlockThreshold: bigint, +): BenchmarkSample { + const instructionBudget = BigInt(iterations) * translatedLoopInstructions; + const session = start_fixture_with_hot_threshold( + "translated-loop", + instructionBudget, + 1_000n, + 1_024n, + 256n * 1_024n * 1_024n, + 1_024n * 1_024n, + new Uint8Array(), + hotBlockThreshold, + ); + const started = performance.now(); + const result = session.resume(new Uint8Array(), false); + const milliseconds = performance.now() - started; + try { + if (result.diagnostic_code !== "resource.instructions") { + throw new Error( + `translation benchmark ended with ${result.diagnostic_code || result.outcome}`, + ); + } + return { + milliseconds, + guestInstructions: result.executed_instructions, + translatedBlocks: result.translated_block_executions, + fallbackBlocks: result.translation_fallback_blocks, + }; + } finally { + result.free(); + session.free(); + } +} + +function median(values: number[]): number { + const ordered = [...values].sort((left, right) => left - right); + return ordered[Math.floor(ordered.length / 2)] ?? Number.NaN; +} + +function runTranslationBenchmark(iterations: number): TranslationBenchmarkReport { + if (!Number.isSafeInteger(iterations) || iterations < 128 || iterations > 100_000) { + throw new Error("translation benchmark iterations must be an integer from 128 to 100000"); + } + const warmupIterations = Math.min(iterations, 256); + runBenchmarkSample(warmupIterations, 0n); + runBenchmarkSample(warmupIterations, 1n); + + const interpreted: BenchmarkSample[] = []; + const translated: BenchmarkSample[] = []; + for (let sample = 0; sample < 5; sample += 1) { + interpreted.push(runBenchmarkSample(iterations, 0n)); + translated.push(runBenchmarkSample(iterations, 1n)); + } + const interpretedMilliseconds = median( + interpreted.map((sample) => sample.milliseconds), + ); + const translatedMilliseconds = median( + translated.map((sample) => sample.milliseconds), + ); + const translatedSample = translated[translated.length - 1]; + if (!translatedSample) { + throw new Error("translation benchmark did not produce a sample"); + } + return { + iterations, + guestInstructions: translatedSample.guestInstructions, + interpretedMilliseconds, + translatedMilliseconds, + speedup: interpretedMilliseconds / translatedMilliseconds, + translatedBlocks: translatedSample.translatedBlocks, + fallbackBlocks: translatedSample.fallbackBlocks, + }; +} + const ready = initializeRuntime().then(createReport); ready @@ -367,6 +453,15 @@ self.addEventListener("message", (event: MessageEvent) => { const command = event.data; void ready .then(async () => { + if (command.kind === "translation-benchmark") { + const message: WorkerMessage = { + kind: "translation-benchmark", + requestId: command.requestId, + report: runTranslationBenchmark(command.iterations), + }; + self.postMessage(message); + return; + } if ( command.kind === "filesystem-export" || command.kind === "filesystem-import" diff --git a/web/tests/translation-benchmark.spec.ts b/web/tests/translation-benchmark.spec.ts new file mode 100644 index 0000000..c3f2432 --- /dev/null +++ b/web/tests/translation-benchmark.spec.ts @@ -0,0 +1,62 @@ +import { expect, test } from "@playwright/test"; + +import type { + TranslationBenchmarkReport, + WorkerCommand, + WorkerMessage, +} from "../src/probe"; + +const enabled = process.env.BINARROW_TRANSLATION_BENCHMARK === "1"; +const iterations = Number(process.env.BINARROW_TRANSLATION_ITERATIONS ?? "16384"); + +test("benchmarks warmed scalar translation against interpretation", async ({ page }) => { + test.skip(!enabled, "set BINARROW_TRANSLATION_BENCHMARK=1"); + test.setTimeout(2 * 60 * 1_000); + + await page.goto("/"); + const report = await page.evaluate( + ({ benchmarkIterations }) => + new Promise((resolve, reject) => { + const worker = new Worker("/src/probe.worker.ts", { type: "module" }); + const timeout = window.setTimeout(() => { + worker.terminate(); + reject(new Error("translation benchmark timed out")); + }, 90_000); + worker.addEventListener("message", (event: MessageEvent) => { + if (event.data.kind === "error") { + window.clearTimeout(timeout); + worker.terminate(); + reject(new Error(event.data.message)); + } else if (event.data.kind === "ready") { + const command: WorkerCommand = { + kind: "translation-benchmark", + requestId: 1, + iterations: benchmarkIterations, + }; + worker.postMessage(command); + } else if (event.data.kind === "translation-benchmark") { + window.clearTimeout(timeout); + worker.terminate(); + resolve(event.data.report); + } + }); + worker.addEventListener("error", (event) => { + window.clearTimeout(timeout); + worker.terminate(); + reject(new Error(event.message)); + }); + }), + { benchmarkIterations: iterations }, + ); + + console.log( + `translation benchmark: ${report.guestInstructions} instructions, ` + + `${report.interpretedMilliseconds.toFixed(2)} ms interpreted, ` + + `${report.translatedMilliseconds.toFixed(2)} ms translated, ` + + `${report.speedup.toFixed(2)}x speedup`, + ); + expect(report.guestInstructions).toBe(BigInt(iterations) * 17n); + expect(report.translatedBlocks).toBe(BigInt(iterations - 1)); + expect(report.fallbackBlocks).toBe(0n); + expect(report.speedup).toBeGreaterThan(1.2); +});