Agent objective
- Interact with the supplied stateful environment.
- Produce verifier-checkable actions or artifacts.
- Maximise scalar reward under the package contract.
A five-environment mathematics suite spanning formal proof repair, convex optimization, graph invariants, automata, and strategic systems. Agents must identify a hidden system through carefully chosen experiments, maintain a calibrated posterior, predict sealed holdouts, and stop efficiently under deterministic verification.
A five-environment mathematics suite spanning formal proof repair, convex optimization, graph invariants, automata, and strategic systems. Agents must identify a hidden system through carefully chosen experiments, maintain a calibrated posterior, predict sealed holdouts, and stop efficiently under deterministic verification.
Enough detail to understand the intellectual terrain; generated instances, hidden mechanisms, and solution paths remain inside the private package.
| Environment | Mathematical or technical frontier | Adaptive research problem |
|---|---|---|
| ProofAudit | Formal methods, finite-model finding, proof repair | Select informative proof obligations and checking strategies to discriminate among plausible repairs without spending the budget on redundant searches. |
| ConvexOracle | Convex analysis and optimization | Infer hidden optimization geometry by deciding where to probe and which value, gradient, curvature, or proximal evidence is worth its cost. |
| GraphInvariant | Spectral graph theory, WL methods, GNN expressivity | Combine spectral, rooted, higher-order, and anchored probes to distinguish regular graph structures that defeat ordinary refinement. |
| AutomataOracle | Active learning and automata theory | Choose discriminating words among low-information decoys while accounting for noisy membership and coarse equivalence feedback. |
| EquilibriumOracle | Game theory and multi-agent learning | Use off-equilibrium payoff, regret, response, and learning-dynamics probes when every candidate agrees at the obvious equilibrium. |
A sealed GPT-5.6 Pro evaluation identified all five hidden systems but passed only four, with 0.8712 mean reward. GraphInvariant missed the 0.80 gate despite correct identification and 0.9938 held-out prediction because its experiment selection and calibration were weaker. That gap makes the suite useful: it distinguishes reaching an answer from conducting an efficient, well-calibrated investigation. The release passed 71 tests and all hard/expert fixture replays.
We publish aggregate behavior and task structure, while withholding generated instances, hidden labels, exact successful probes, private checks, and solution trajectories.
Shown with its provenance and limitations; it is not a performance guarantee.
Reported result from the evaluation artifact supplied with this package.
As identified by the supplied artifact.
See the source methodology.
5Math_2_gpt56pro_hard_report.md
Machine-readable provenance and the exact displayed metric are available in results.json.
The paid ZIP will live in a private R2 bucket. Vercel authorizes the buyer and issues a 2–5 minute object URL; R2 serves the bytes directly.
Authenticated buyer + entitlement check
+ private R2 object + 2–5 minute signed URL
= direct, auditable download
Package SHA-256
d8a11f7841efba2a3079f8e4760479721a9d21e50c10957bde300353f39b30efOne purchase licenses this identified item to one legal organisation for worldwide, perpetual commercial model training, evaluation, research and development. Redistribution and resale of the package are not permitted.