This repository has no description
README.md

Rosetta Stone #

Can agents invent novel formal systems that outstrip oversight — and can other agents reverse-engineer them from artifacts alone?

Motivation #

Current oversight research (e.g. EsoLang-Bench) uses static environments and test sets, but real AI R&D automation will feature agents that invent novel formal systems. An agent that builds a complex abstraction retains the context of how it works; a monitor facing that abstraction — without the builder's weights, reasoning, or context — may not. We hypothesize that this in-context tool building is a primary vector for agents outstripping oversight, enabling RSI-like dynamics even without weight updates.

Method #

Authors. ~10 models each invent ~10 novel esolangs. A model may submit an environment iff it first solves 8 EsoLang-Bench problems — two per difficulty tier, seeded-sampled from the 80 — in its own invented language, with 3 scored submission attempts per problem (local runs are unlimited and unscored).

The Mario Maker test. Every submission must contain a working solution set, which is then redacted for solvers. Execution is a fresh, network-disabled Linux container; a submission fails if it doesn't reproduce there deterministically (N=3 independent re-runs). The contract is fixed in advance: one entry point (run.py, run.ts, or a compiled ./run) taking one argument, a solution file; test input on stdin, output on stdout; verification judges captured output only — never exit codes.

Stages. A verified submission becomes a durable, reusable bundle. v1's deliverable is 80–100 verified stages. Solver experiments come later: a fresh agent sees 4 of a stage's task/solution pairs, 4 are held out, and it must submit a working interpreter/compiler — under ablations (docs / no docs, white-box source / black-box runner, solutions only) all derivable from the stored bundle without re-authoring.

The full spec is docs/BRIEF.md (binding decisions in §6); design history lives in docs/handoffs/.

How to run the experiments #

For the complete command-oriented runbook—including the execute-only MURK React solver spike and its current limitations—see docs/HOW-TO-RUN.md.

Where experiments run #

The persistent Ubuntu server reached through the SSH alias rosetta-runner is the jump box: it holds the checkout, configs, preflight records, and Daytona launch manifests. Actual author campaigns run inside Daytona sandboxes. A laptop or coding agent connects to the jump box to launch and inspect those sandboxes; it does not run the fleet itself. The conventional remote checkout is ~/rosetta-stone.

ssh rosetta-runner
cd ~/rosetta-stone
tmux list-sessions
sed -n '1,240p' logs/shards/ambitious-sweep/daytona_launch_manifest.json

Each laptop needs its own authorized SSH key and a local ~/.ssh/config entry for rosetta-runner. No server address, username, private key, or provider credential belongs in this repository. See the runbook for the one-time Ubuntu setup, Daytona launch and observation commands, artifact-collection limitation, and private Inspect View tunnel. The provider-neutral operating checklist is docs/runbooks/jump-box.md.

Prerequisites #

  • The jump box needs uv, git, and tmux. Daytona sandboxes bootstrap uv, bun, cargo, and Docker for campaign work.
  • A DAYTONA_API_KEY and the per-model provider keys named by the allocation manifest. DAYTONA_API_URL and DAYTONA_TARGET are optional overrides for non-default Daytona deployments.
  • An OpenRouter key (OPENROUTER_API_KEY) or per-provider keys — model routes are set per model in configs/models.yaml.
  • Optional: S3/R2 credentials for artifact sync (see .env.example for every env var).
ssh rosetta-runner
cd ~/rosetta-stone
uv sync
uv run pytest -m "not slow and not network"   # hermetic suite — no Docker, no keys

Configure #

  1. Experiment YAML — copy configs/experiment.example.yaml to e.g. configs/run-2026-07.yaml and fill in: gate_seed, wall_time_budget_hours, the campaign: block (stage_target, max_attempts_per_model, dir, sync), and a models: mapping keyed by model id. Each model block owns its provider route, exact pinned agent scaffold (scaffold + scaffold_version), offered reasoning_effort_options, and selected reasoning_effort. For OpenRouter runs, set an explicit openrouter_min_context_tokens floor and pin exactly one endpoint with openrouter_provider.only plus allow_fallbacks: false. rosetta preflight pulls live OpenRouter pricing/endpoints, validates that exact pin, writes the pricing snapshot, and hard-errors if the pinned route is ineligible. Do not hand-maintain a pricing table in reviewed sweep YAMLs.
  2. Optional shared model registry — older configs may keep models: [id, ...] in the experiment YAML plus reasoning_effort: {id: effort} and put model definitions in configs/models.yaml / --models. New run configs should prefer inline model blocks so each model is edited in one place. To charge a run to a dedicated OpenRouter key, set openrouter_api_key_env: ROSETTA_OPENROUTER_API_KEY and put the secret in that env var; the campaign passes it to Inspect as an explicit api_key, so ambient OPENROUTER_API_KEY does not override it. For black-box-ready production runs, set entry_point_policy: compiled_binary; the author prompt and gate then accept only an actual compiled executable ./run rather than run.py / run.ts or text scripts with an executable bit.
  3. Env — copy .env.example to .env. With sync: true, all ROSETTA_S3_* / AWS_* vars are required (the run refuses to start without them; there is no silent local-only mode).

Run #

Connect to the Ubuntu jump box and create a persistent launch session. Generate and preflight the shards there, then ask Daytona to start one sandbox per shard:

ssh rosetta-runner
cd ~/rosetta-stone
tmux new -s daytona-launch

set -a
source .env
set +a

uv run rosetta shard configs/experiment.ambitious-sweep.yaml \
  --out-dir logs/shards/ambitious-sweep \
  --shards-per-model 2 \
  --key-env-scheme 'ROSETTA_OPENROUTER_{model}_API_KEY'

for shard in logs/shards/ambitious-sweep/*.yaml; do
  uv run rosetta preflight "$shard" --min-context 262144
done

uv run --with daytona python scripts/daytona_launch.py \
  logs/shards/ambitious-sweep/allocation_manifest.yaml \
  --repo-root "$PWD" \
  --output logs/shards/ambitious-sweep/daytona_launch_manifest.json

Detach with Ctrl-b d. The launch command uses tmux so disconnecting during sandbox creation or bootstrap does not interrupt it. After the launcher returns, campaigns keep running under nohup inside Daytona; the jump-box session does not own them. See docs/runbooks/scaling-daytona.md for queued-shard retries and inspection.

Drives author runs per model toward the stage target: author → package → verify. Only a submission that verifies in 3 independent fresh no-network containers counts as a stage. Failed runs (gate failed / timed out / verify failed) are journaled with reasons and retried under a fresh seed up to the per-model cap. The campaign is resumable: re-running the same config skips completed work; state divergence is a hard error, never a silent reconcile.

Observe #

Read the launch manifest on the jump box, list the Daytona sandboxes, then execute rosetta status inside the selected sandbox. A jump-box-local status command sees a campaign only after its artifacts have been copied there. The current launcher does not perform that collection automatically; do not delete a Daytona sandbox before retrieving the artifacts. Exact commands are in the scaling runbook.

Everything needed to re-run or analyze is captured: full LLM I/O in Inspect .eval logs, the exact prompts-file hash and pinned scaffold versions in each stage's manifest.json, and the effective config snapshot in the campaign dir. sync: true pulls the bucket before a run and pushes after — including on failure.

Generate a static run report #

rosetta report creates a disposable, read-only directory for one explicitly named verified stage, author attempt, campaign, or execute-only spike. It never resumes, scores, verifies, uploads, or changes the source run.

uv run rosetta report \
  --scope verified_stage \
  --input-root logs/campaigns/example \
  --identity BUNDLE_CONTENT_HASH \
  --disclosure shareable \
  --inspect-log-root logs/campaigns/example/logs \
  --inspect-view-base https://inspect.example.org \
  --output reports/example-stage

The shareable profile contains only the sanitized Rosetta page and an HTTPS link to an access-controlled Inspect viewer. The internal_evidence profile omits --inspect-view-base and adds an Inspect-owned static bundle for exactly the logs named by the report. That entire output is sensitive and unsuitable for public hosting. Inspect remains authoritative for raw prompts, completions, messages, tool calls, and full model trajectories in both profiles. See docs/reporting.md.

Integration tests (real Docker + real model) #

ROSETTA_RUN_SLOW=1 uv run pytest -m slow             # live author/verify e2e (~20-35 min)
ROSETTA_RUN_NETWORK_TESTS=1 uv run pytest -m network # one real EsoLang-Bench download