52 items Real launch catalogue

Every environment has a permanent technical page.

Inspect its contract, immutable package hash, task-level price, licence, and supplied model or policy result before checkout.

Mathematics5Math-1

5Math-1

A five-environment mathematics suite in which agents uncover hidden mechanisms through budgeted experiments across optimization dynamics, kernel geometry, attention operators, Boolean Fourier structure, and entropic transport. Noisy compressed observations and sealed transfer checks reward informative measurement choices and calibrated uncertainty rather than static answer recall.

Version1.0.0
Environments5
Result0.8401
Mathematics5Math-2

5Math-2

A five-environment mathematics suite spanning formal proof repair, convex optimization, graph invariants, automata, and strategic systems. Agents must identify a hidden system through carefully chosen experiments, maintain a calibrated posterior, predict sealed holdouts, and stop efficiently under deterministic verification.

Version1.0.0
Environments5
Result80%
Mathematics5Math-3

5Math-3

A five-environment mathematics suite about scientific identification under uncertainty. Agents choose costly, noisy experiments across generative models, symmetry, topology, neural geometry, and conformal uncertainty, then transfer the inferred mechanism to sealed conditions—a compact test of mathematical reasoning, planning, and judgment.

Version1.0.0
Environments5
Result0.8784
Mathematics5Math-4

5Math-4

Five adaptive mathematics environments spanning tensor decomposition, topological invariants, kernel geometry, coding theory, and nonlinear dynamics. Each hides meaningful structure behind deliberately non-identifying summaries, forcing agents to choose complementary probes, manage experimental cost, and make calibrated predictions on sealed cases.

Version1.0.0
Environments5
Result0.9395
Mathematics5Math-5

5Math-5

Five research mathematics environments built around exact nuisance equivalences: familiar summaries are identical across candidate systems, so agents must discover which advanced invariant actually carries information. The suite spans cellular sheaves, Schrödinger bridges, inverse operators, algebraic varieties, and random-matrix spectra with strict budgets and sealed generalization.

Version1.0.0
Environments5
Result0.9958
Mathematics5Math-6

5Math-6

Five research-tier mathematical investigations built around hidden systems with deceptively identical public invariants. Across cellular sheaves, non-normal dynamics, algebraic geometry, quantum channels, and tensor orbits, agents must spend a severe measurement budget, distinguish competing mechanisms, calibrate their posterior, and generalize beyond chosen probes.

Version1.0.0
Environments5
Result0.8213
Mathematics5Math-7

5Math-7

A five-environment mathematics suite built around long-horizon inference under scientific and measurement uncertainty. Each domain has its own mathematical model and action grammar, requiring agents to maintain compact beliefs, choose observation-dependent experiments, manage nonlinear costs and irreversible actions, predict sealed outcomes, and know when to identify, return a candidate set, or abstain.

Version2.0.0
Environments5
Result0.8790
Mathematics5Math-8

5Math-8

A five-environment supercompact mathematics benchmark with thousands of candidate systems, identical starting evidence, and only two noisy measurements before submission. It spans Hodge-theoretic period maps, higher algebra, information geometry, nonlinear convex PDEs, and graph-cover spectra, rewarding complementary experiment planning and calibrated transfer to sealed probes.

Version1.0.0
Environments5
Result0.7073
Mathematics5Math-9

5Math-9

A five-environment research suite where advanced mathematics becomes a sequence of consequential decisions. Agents navigate higher coherence, Stokes geometry, singular learning, generalized symmetry, and mean-field games while managing noisy evidence, nuisance variables, scarce resources, and irreversible interventions; success requires calibrated beliefs and sealed prediction, not merely naming the right concept.

Version1.0.0
Environments5
Result0.6700
Mathematics5Math-10

5Math-10

A five-environment mathematics benchmark for extreme adaptive identification: each expert episode hides one system among 32,768 candidates and permits only two noisy measurements. Its worlds range from thermodynamic dynamics and cluster wall crossing to entropy solutions, planar algebras, and non-Shannon information geometry, demanding complementary experiment planning, sharp posterior inference, and sealed transfer.

Version1.0.0
Environments5
Result0.5007
Mathematics5Math-11

5Math-11

Thirty exact-mathematics environments spanning algebraic geometry, homological algebra, representation theory, arithmetic forms, and real algebraic geometry. Agents build connected chains of verified artifacts rather than isolated answers: accepted work unlocks later steps, revisions invalidate dependencies, and exact certificate gates test transfer across semantically distinct constructions.

Version3.0.0
Environments30
Result50.0%
Mathematics5Math-Hard-1

5Math-Hard-1

Five apex mathematics environments where 64 candidate mechanisms share exactly non-identifying seed evidence and agents get at most three costly probes. The suite spans rough paths, weighted automata, lattice geometry, replicator dynamics, and inverse conductivity, demanding noncommutative reasoning, complementary measurements, calibrated posteriors, and transfer to sealed probes.

Version1.0.0
Environments5
Result18.3%
Mathematics5Math-Hard-2

5Math-Hard-2

Five advanced mathematics environments built around adaptive inference from deliberately non-identifying evidence. Across arithmetic representations, optimal stopping, free probability, oriented matroids, and filtered homological algebra, agents face 128 shuffled hypotheses and at most two noisy measurements, rewarding complementary probes, calibrated uncertainty, and transfer to unseen conditions.

Version1.0.0
Environments5
Result0.7342
Mathematics5Math-Hard-3

5Math-Hard-3

Five advanced mathematical inverse problems compressed into an extreme two-measurement challenge. Across microlocal analysis, Hamiltonian geometry, polynomial certificates, braid topology, and oscillatory microstructure, agents distinguish 256 hidden models from noisy observations; exact seed invariants provide no shortcut, so success requires complementary design, calibrated inference, and sealed transfer.

Version1.0.0
Environments5
Result10.6%
Mathematics5Math-Hard-4

5Math-Hard-4

Five omega-tier mathematics environments with 512 seed-shuffled candidates and only two noisy measurements. Across Stokes phenomena, decoder pseudocodewords, singular learning theory, higher associativity, and flag geometry, agents must escape exact seed equivalence, reason through nonlinear interactions, and submit calibrated posteriors that transfer to unseen mathematical functionals.

Version1.0.0
Environments5
Result0.7047
Mathematics5Math-Hard-5

5Math-Hard-5

Five mathematical system-identification environments with an extreme information bottleneck: two priced measurements must resolve a calibrated posterior over 1,024 hidden systems. From renormalization and p-adic dynamics to noncommutative geometry, spin glasses, and information geometry, obvious low-order invariants are identical and sealed transfer matters as much as identification.

Version1.0.0
Environments5
Result0.2868
Mathematics5Math-Hard-6

5Math-Hard-6

Five mathematical system-identification environments with 2,048 hypotheses and only two noisy measurements. From hyperbolic geometry and graph limits to inverse scattering, conformal bootstrap, and mean-field games, every candidate shares substantial exact invariants, so success depends on a complementary probe pair and a sharply calibrated posterior that transfers to sealed functionals.

Version1.0.0
Environments5
Result0.7619
Mathematics5Math-Hard-7

5Math-Hard-7

Five measurable-tier mathematics environments spanning Bridgeland wall crossing, KAM resonance, random-matrix edges, theta characteristics, and proof-net dynamics. Agents receive invariant starting evidence and only two noisy measurements, then must choose complementary probes that expose hidden geometry while maintaining calibrated uncertainty and predicting disjoint sealed functionals.

Version1.0.0
Environments5
Result0.2443
Mathematics5Math-Hard-8

5Math-Hard-8

Five independent frontier mathematics environments spanning derived deformation theory, knot concordance, renormalization flows, proof complexity, and ergodic joinings. Agents jointly infer scientific structure and apparatus nuisance variables while navigating branch-dependent probes, scarce resources, irreversible transformations, destructive decoys, calibrated terminal decisions, and sealed predictions.

Version1.0.0
Environments5
Result0.7711
Mathematics5Math-Hard-9

5Math-Hard-9

Five deep-frontier mathematical investigations across polyhedral homotopy, symplectic capacities, graphon stability, microlocal propagation, and motivic regulators. Each environment has a distinct scientific state machine: agents must route observation-dependent diagnostics, manage nuisance variables and irreversible choices, preserve future options, and satisfy separate posterior, prediction, decision, and path-quality gates.

Version1.0.0
Environments5
Result0.9716
Mathematics5Math-Hard-10

5Math-Hard-10

A five-environment frontier suite with 16,384 hidden hypotheses and only two compressed measurements. It spans automorphic forms, Morse continuation, symplectic phase retrieval, scattering amplitudes, and convex integral geometry. Every signal sits beyond an exact invariant quotient, forcing complementary experiments, enormous-posterior reasoning, and confident transfer to disjoint sealed mathematics.

Version1.0.0
Environments5
Result0.6827
Mathematics5Math-Hard-11

5Math-Hard-11

Five long-horizon mathematics environments with noisy apparatus, hidden nuisance variables, and irreversible decisions. Across cluster wall crossing, p-adic Hodge theory, rough paths, KAM resonance, and spin-glass landscapes, agents route observations through distinct state machines, allocate scarce resources, and choose justified identification, bounded-set, or abstention decisions with sealed transfer.

Version1.0.0
Environments5
Result0.4923
Machine learning research5ML-1

5ML-1

Five machine-learning research environments for adaptive diagnosis under model misspecification. They span causal mechanisms, scaling behavior, representation geometry, continual learning, and uncertainty under distribution shift, requiring costly diagnostics, compositional mechanism inference, continuous-variation tracking, and explicit forecasts under heavy-tailed, common-mode error.

Version2.0.0
Environments5
Result0.7550
Machine learning research5ML-2

5ML-2

Five machine-learning forensics environments where agents diagnose compositional systems through only five audits. Each expert case combines mechanisms, latent continuous effects, model mismatch, heavy-tailed observations, and twelve sealed forecasts, so success requires crossing diagnostic regimes, calibrated uncertainty, and predicting average and worst-case behavior—not merely naming a plausible stack.

Version3.0.0
Environments5
Result0.2951
Machine learning research5ML-3

5ML-3

Five research environments for investigating complex machine-learning systems rather than merely scoring outputs. Agents intervene across transformer circuits, model editing, graph reasoning, compression, and meta-learning; distinguish compositional mechanisms; estimate continuous properties; and forecast unseen outcomes with calibrated intervals from only five experiments.

Version4.0.0
Environments5
Result0.5411
Machine learning research5ML-4

5ML-4

Five frontier machine-learning environments for diagnosing complex systems under severe model misspecification. Across sparse expert routing, multimodal alignment, neural operators, preference optimization, and world-model planning, agents get only five experiments to identify one of 256 mechanism stacks, estimate latent parameters, and produce calibrated intervals for sealed transfer forecasts.

Version5.0.0
Environments5
Result0.3128
Quantum5Quantum-1

5Quantum-1

Five quantum-information environments that require complete experimental protocols rather than chasing the largest immediate signal. They span non-Markovian processes, open-system generators, many-body phases, fault-tolerant noise, and network nonlocality, combining mechanism identification, continuous parameter recovery, calibration, constrained sequencing, credible intervals, and prediction on sealed experiments.

Version2.0.0
Environments5
ResultPassed
Quantum5Quantum-2

5Quantum-2

Five quantum-information environments for adaptive experimental reasoning under model aliasing, nuisance parameters, laboratory drift, and strict resource limits. Each expert episode requires calibration, complementary probe families, a posterior-selected stress test, and unseen transfer—combining physical-model identification, continuous parameter recovery, calibrated uncertainty, and scientific certification.

Version3.0.0
Environments5
Result0.5666
Quantum5Quantum-3

5Quantum-3

Five research-level quantum-information investigations in which agents distinguish 128 aliased physical hypotheses, calibrate nuisance effects, recover continuous parameters, choose posterior-dependent stress and falsification experiments, and publish uncertainty-aware certificates. Success is tested across interpolation, extrapolation, and counterfactual regimes—not merely by choosing a plausible mechanism.

Version4.0.0
Environments5
Result0.5257
Operations & simulationsA3-OptiOps

A3-OptiOps

A 12-stage proof-carrying optimization-engineering benchmark built around a realistic GPU-cluster scheduling service. Agents extract policy semantics, build scalable anytime solvers, react to failures and uncertainty, diagnose infeasibility, preserve production state, implement repository-scale protocol changes, resist adversarial evaluation, transfer to unseen constraints, and demonstrate optimizer improvements.

Version2.0.0
Environments12
Result0.7303
Operations & simulationsCloudFinOps

CloudFinOps

A production-remediation benchmark where agents reduce a generated cloud estate's spend without sacrificing performance, resilience, retention, governance, or reproducibility. It reconciles provider-shaped billing, telemetry, inventory, dependencies, incidents, commitments, and Terraform state, then tests open-ended multi-resource changes through staged rollout, recovery, approval, and evidence-backed savings verification.

Version2.0.0
Environments1
Result0.9975
Operations & simulationsCrisisGrid

CrisisGrid

A long-horizon wildfire-command environment for planning under partial observability, uncertain weather, failing roads, limited resources, and competing public-safety goals. Centralized or cooperative policies coordinate engines, bulldozers, helicopters, scouts, and evacuation buses across procedural maps while balancing containment, evacuation, assets, ecology, responder safety, equity, cost, and forecast shifts.

Version0.1.0-beta.1
Environments1
Result0.4291
Data engineeringDATA-0188

DATA-0188

A time-boxed PostgreSQL optimization environment where agents rewrite one analytical cohort query and select up to three constrained indexes. Hidden evaluation requires exact business semantics before scoring p95 speedup across shifted data regimes alongside worst-case latency, index footprint, build cost, and write overhead—rewarding robust query-and-index co-design rather than narrow benchmark speed.

Version2.0.0
Environments1
Result0.3623
Data engineeringDATA-0211

DATA-0211

A demanding data-engineering repair task for incremental event-time reconciliation across thousands of interleaved streams. Timestamps can expand into timezone and daylight-saving candidates, while exact elapsed constraints, cross-chunk state, ambiguity witnesses, deterministic output, and bounded memory must remain correct at up to 250,000 events without enumerating complete histories.

Version4.0.0
Environments1
Result55 min
Data engineeringData-Pipeline

Data-Pipeline

A private-judge environment for diagnosing and repairing procedural, multi-fault analytics incidents across SQL, restricted Python, contracts, orchestration, warehouse state, lineage, logs, and dashboards. Agents must gather evidence, patch durable artifacts, run targeted recovery, and submit one structured diagnosis that survives hidden rebuild, idempotency, partition-isolation, and future-replay checks.

Version2.0.0
Environments1
Result0.9248
MathematicsErdosRL-v2

ErdosRL-v2

Seventeen mathematics environments that turn finite extremal problems and audited reasoning failures into stateful, verifier-backed work. Thirteen demand exact combinatorial witnesses and sealed optima; four require a concrete obstruction, a repaired argument, preserved valid claims, and a consistent claim ledger under deterministic dual verification.

Version0.2.0
Environments17
Result13.20 / 17 (77.65%)
Operations & simulationsIncident-Commander

Incident-Commander

A compositional incident-response environment for diagnosing and mitigating outages in a synthetic multi-region microservice system. Episodes combine initiating mechanisms, dependency edges, regional or tenant scope, propagation, observability failures, recovery blockers, convergence stages, and plausible unrelated changes; agents must gather evidence, apply narrow controls, verify recovery independently, communicate accurately, and submit an evidence-linked postmortem.

Version3.0.0
Environments1
Result0.9926
Software engineeringM1-Lineage

M1-Lineage

An ML-systems benchmark for repairing stateful training infrastructure without corrupting continuation, optimizer lineage, distributed state, schema history, or durable checkpoints. Eight private stages contain multi-module incidents with semantic defects across resume, shards, surgery, schemas, compression, recovery, transactions, and incident response under attestation, corruption, concurrency, crash, determinism, and resource checks.

Version5.0.0
Environments8
Result0.5779
Operations & simulationsMercantile-Commons

Mercantile-Commons

A procedural, partially observed, general-sum multi-agent environment where policies operate firms inside evolving supply networks. Agents produce, trade, bargain, share or challenge forecasts, manage credit and trust, and survive correlated shocks with unfamiliar counterparties while balancing profit, service, solvency, contribution, relationship integrity, sustainability, fairness, and lower-tail resilience.

Version0.1.0-beta.1
Environments1
Result0.2413
Software engineeringN2-Ocaml

N2-Ocaml

A cumulative twelve-stage OCaml compiler-engineering environment that grows from an interpreter into a durable development system. Agents repair lexer, parser, evaluator, type inference, bytecode, diagnostics, caching, optimization, dependency planning, snapshots, and replay while preserving every earlier frontier across focused learning, regression recovery, cold-start reconstruction, and seeded-fault diagnosis.

Version4.0.0
Environments12
Result39/39
Software engineeringN4-ForthForge

N4-ForthForge

A stateful coding-agent environment for engineering and operating a Forth virtual machine. Its twelve stages progress from arithmetic, stacks, compilation, and control flow through persistent images, migration, forensic recovery, multi-session services, portability, adversarial robustness, namespace transfer, and optimizer discovery, with private procedural families and independent replay discouraging fixture-specific patches.

Version2.0.0
Environments12
Result0.9998
Operations & simulationsOPS-0082

OPS-0082

A time-boxed production-operations incident in a deterministic Kubernetes-behavior simulator with continuous traffic. Each episode combines latent fault classes and may require rollback, canary-template repair, multi-resource configuration repair, or an evidence-supported no-write hold. Agents must diagnose safely, preserve four stable plus one canary endpoint, and stay within a two-mutation budget.

Version4.0.0
Environments1
Result1.000
MathematicsProofRepair

ProofRepair

Five environments for auditing and repairing flawed mathematical arguments under evidence and tool constraints. Agents inspect interleaved branches, search a frozen corpus, test theorem applicability, construct executable counterexamples, recover nonlocal dependencies, locate the earliest fatal gap, and propose minimal repairs that survive controlled hidden variants.

Version0.3.0
Environments5
Result0.9472
Quantitative financeQFIN-1

QFIN-1

Fifteen quantitative-finance control environments across execution, statistical arbitrage, market making, portfolio allocation, and option hedging at Core, Research, and Challenge tiers. Agents face richer costs, constraints, delays, hidden liquidity, and stress, with paired exogenous tapes, observation-only policies, native financial metrics, and deterministic replay.

Version3.0.0
Environments15
Result7/10
Operations & simulationsRedQueen-1

RedQueen-1

Four non-pausing environments for agents whose world worsens while they deliberate. Every action incurs exogenous latency before application, and commitment costs extra. Across evolving circuits, online job queues, volatile repositories, and locally observed navigation, policies must allocate compute, recover from stale plans, gather bounded information, checkpoint reversible progress, and know when further reasoning is no longer worth its delay.

Version0.3.0
Environments4
Result41.8332
Operations & simulationsRedQueen-2

RedQueen-2

Five metareasoning environments where the world worsens while an agent thinks. Every reasoning, inspection, hold, checkpoint, or commitment advances hidden state first, so computation can improve a plan while making it infeasible. Exact simulators test information gathering, option preservation, stopping, and action across negotiation, infrastructure, auctions, forensics, and distributed recovery.

Version0.1.0
Environments5
Result55.2503
Operations & simulationsRedQueen-3

RedQueen-3

Five adaptive-computation environments that continue changing while an agent deliberates. Every action consumes unavoidable latency before interpretation; observations arrive late, plans stale, and commitments cost deployment time. Across revision-sensitive proofs, causal drift, schema migration, orbital hazards, and mutating protocols, the suite measures stopping, value of computation, option preservation, stale-plan recovery, active information gathering, and compute allocation.

Version0.1.0
Environments5
Result7.6814
Operations & simulationsRedQueen-4

RedQueen-4

Five adaptive test-time-computation environments in worlds that never pause. Every action consumes latency before interpretation, allowing evidence, rules, utilities, network conditions, or identities to change while the agent reasons. Across memory, planning, sensor fusion, network coding, and exact-cover assembly, success depends on what to inspect, stabilize, validate, and commit before a solution becomes stale.

Version0.1.0
Environments5
Result27.4083
Operations & simulationsSOC-Defender

SOC-Defender

A fully synthetic, offline defensive-security environment where agents triage alerts, investigate identity, endpoint, network, cloud, and change telemetry, preserve evidence, apply targeted containment, update an incident ticket, and submit a structured report. Eight procedural families include genuine attacks and benign lookalikes, with deterministic scoring for containment, evidence, scope, continuity, escalation, and investigation quality.

Version1.0.0
Environments1
Result0.9021 · failed
Software engineeringSWE-0147

SWE-0147

A TypeScript software-engineering benchmark for credential epochs under hostile JavaScript re-entry. The visible suite starts green while private cases expose interacting authority, liveness, retry, boundary, validation, and resource failures. Agents must repair stale-token rotation, refresh ownership, cross-context races, logout, and exact public declarations under isolated, mutation-audited verification.

Version6.0.0
Environments1
ResultReference patch
Software engineeringSWE-0194

SWE-0194

A production-style Go benchmark for reconnect ownership, protocol recovery, event integrity, concurrency, and resource safety in a resilient WebSocket client. Agents must repair a state machine that behaves plausibly in simple tests but fails race-enabled fixed cases and held-out deterministic traces under a confidential operator verifier.

Version2.0.0
Environments1
Result1.000
Software engineeringSWE-0240

SWE-0240

A long-horizon software-engineering incident that migrates compromised browser token custody to a server-mediated OAuth/OIDC architecture. Across an 8k-line workspace, agents must diagnose cross-service failures, reason about concurrency, implement Authorization Code with PKCE, preserve migration compatibility, and recover safely under sealed generated traces and mutation-sensitive verification.

Version3.0.0
Environments1
Result0.000