Agent objective
- Interact with the supplied stateful environment.
- Produce verifier-checkable actions or artifacts.
- Maximise scalar reward under the package contract.
A 12-stage proof-carrying optimization-engineering benchmark built around a realistic GPU-cluster scheduling service. Agents extract policy semantics, build scalable anytime solvers, react to failures and uncertainty, diagnose infeasibility, preserve production state, implement repository-scale protocol changes, resist adversarial evaluation, transfer to unseen constraints, and demonstrate optimizer improvements.
A 12-stage proof-carrying optimization-engineering benchmark built around a realistic GPU-cluster scheduling service. Agents extract policy semantics, build scalable anytime solvers, react to failures and uncertainty, diagnose infeasibility, preserve production state, implement repository-scale protocol changes, resist adversarial evaluation, transfer to unseen constraints, and demonstrate optimizer improvements.
Enough detail to understand the intellectual terrain; generated instances, hidden mechanisms, and solution paths remain inside the private package.
| Environment | Mathematical or technical frontier | Adaptive research problem |
|---|---|---|
| Stages 0–2 | Model extraction, scalable optimization, dispatch | Compile business policy into a coherent model, solve at scale, and select algorithms according to structure rather than hard-coded size. |
| Stages 3–5 | Anytime, online, and robust scheduling | Improve feasible schedules across budgets, replan after failures and arrivals, and balance expected cost with tail risk. |
| Stages 6–7 | Diagnosis and production state | Prove minimally harmful feasibility repairs and support concurrent sessions, snapshots, restoration, migration, and deterministic replay. |
| Stages 8–10 | Repository change, adversarial robustness, transfer | Extend accelerator protocols without regressions, resist decoys and tampering, and handle unseen constraint families through plugins. |
| Stage 11 | Algorithm discovery | Improve an already strong portfolio solver and preserve the speed–quality gain under hidden families and scale shifts. |
The evaluated patch produced valid submissions on all 12 stages and averaged 0.7303 privately, passing five. Public performance averaged 0.8425, but several stages fell sharply on hidden cases—model extraction dropped from 0.7243 to 0.1858. That public/private gap makes the suite useful for studying overfitting, operational shifts, persistent consequences, and real transfer. Validation reports 25 package tests and 84 workspace contract tests.
We publish aggregate behavior and task structure, while withholding generated instances, hidden labels, exact successful probes, private checks, and solution trajectories.
Shown with its provenance and limitations; it is not a performance guarantee.
One fresh hidden micro-tier evaluation per stage; 5/12 stages passed and all 12 submissions were valid.
As identified by the supplied artifact.
12 reported runs.
score_report.json
Machine-readable provenance and the exact displayed metric are available in results.json.
The paid ZIP will live in a private R2 bucket. Vercel authorizes the buyer and issues a 2–5 minute object URL; R2 serves the bytes directly.
Authenticated buyer + entitlement check
+ private R2 object + 2–5 minute signed URL
= direct, auditable download
Package SHA-256
7f329daea1473507277e27ea34bf4e743f4d2dd1ffc02b43a380af156184456dOne purchase licenses this identified item to one legal organisation for worldwide, perpetual commercial model training, evaluation, research and development. Redistribution and resale of the package are not permitted.