Agent objective
- Interact with the supplied stateful environment.
- Produce verifier-checkable actions or artifacts.
- Maximise scalar reward under the package contract.
A cumulative twelve-stage OCaml compiler-engineering environment that grows from an interpreter into a durable development system. Agents repair lexer, parser, evaluator, type inference, bytecode, diagnostics, caching, optimization, dependency planning, snapshots, and replay while preserving every earlier frontier across focused learning, regression recovery, cold-start reconstruction, and seeded-fault diagnosis.
A cumulative twelve-stage OCaml compiler-engineering environment that grows from an interpreter into a durable development system. Agents repair lexer, parser, evaluator, type inference, bytecode, diagnostics, caching, optimization, dependency planning, snapshots, and replay while preserving every earlier frontier across focused learning, regression recovery, cold-start reconstruction, and seeded-fault diagnosis.
Enough detail to understand the intellectual terrain; generated instances, hidden mechanisms, and solution paths remain inside the private package.
| Environment | Mathematical or technical frontier | Adaptive research problem |
|---|---|---|
| Levels 0–4 | Interpreter, types, lists, stable API | Build lexical, parsing, closure, Hindley–Milner, pattern, and compatibility semantics with deterministic errors. |
| Levels 5–6 | Bytecode and diagnostics | Compile to a checked bounded stack machine and produce phase-specific Unicode-aware source diagnostics. |
| Levels 7–9 | Cache, optimizer, dependency graph | Preserve migrations, invalidation, corruption recovery, strict errors, and deterministic planning through cycles and missing edges. |
| Level 10 | Durable snapshots | Implement immutable generations, publication locks, cleanup, fsync, atomic rename, and atomic-current updates. |
| Level 11 | Deterministic session replay | Replay build, status, write, and delete histories while preserving dependency validity across standard through research profiles. |
In a no-network single session, Claude Opus 5 completed every focused level, four cold-start Level-11 profiles, four regression runs, and a mutation run at 1.000 accepted. The final level initially missed one replay case, then recovered after diagnosing a dependency-deletion contract. The run covers 39 judged executions, while references pass all 36 level/track combinations and the corrected mutation matrix kills 15/15. Important caveats include unsandboxed execution and open commercial-certification gates.
We publish aggregate behavior and task structure, while withholding generated instances, hidden labels, exact successful probes, private checks, and solution trajectories.
Shown with its provenance and limitations; it is not a performance guarantee.
Single-session evaluation with no external network. All focused, gauntlet, regression, and mutation runs were accepted; each scored 1.000.
As identified by the supplied artifact.
39 reported runs.
n2_v4_run_notes.md
Machine-readable provenance and the exact displayed metric are available in results.json.
The paid ZIP will live in a private R2 bucket. Vercel authorizes the buyer and issues a 2–5 minute object URL; R2 serves the bytes directly.
Authenticated buyer + entitlement check
+ private R2 object + 2–5 minute signed URL
= direct, auditable download
Package SHA-256
5e299bdd59936aa52ddd6b49888d8dfa103b31583d1095987633780236dc8e5eOne purchase licenses this identified item to one legal organisation for worldwide, perpetual commercial model training, evaluation, research and development. Redistribution and resale of the package are not permitted.