Agent objective
- Interact with the supplied stateful environment.
- Produce verifier-checkable actions or artifacts.
- Maximise scalar reward under the package contract.
Five frontier machine-learning environments for diagnosing complex systems under severe model misspecification. Across sparse expert routing, multimodal alignment, neural operators, preference optimization, and world-model planning, agents get only five experiments to identify one of 256 mechanism stacks, estimate latent parameters, and produce calibrated intervals for sealed transfer forecasts.
Five frontier machine-learning environments for diagnosing complex systems under severe model misspecification. Across sparse expert routing, multimodal alignment, neural operators, preference optimization, and world-model planning, agents get only five experiments to identify one of 256 mechanism stacks, estimate latent parameters, and produce calibrated intervals for sealed transfer forecasts.
Enough detail to understand the intellectual terrain; generated instances, hidden mechanisms, and solution paths remain inside the private package.
| Environment | Mathematical or technical frontier | Adaptive research problem |
|---|---|---|
| SparseExpertRoutingLab | Mixture-of-experts routing and systems | Separate routing mechanisms that match average utilization but diverge under overflow, tail-token, shift, and communication stress. |
| MultimodalAlignmentStressLab | Cross-modal alignment and grounding | Use corruption, conflict, missing-modality, and localization tests to distinguish stacks near modality-dominance transitions. |
| NeuralOperatorInductiveBiasLab | Neural operators and physical constraints | Reverse-engineer architecture and constraints across resolution, boundaries, stiffness, geometry, and long rollouts. |
| PreferenceOptimizationForensics | Preference learning and safety | Probe noisy, shifted, adversarial, and safety-conflicting comparisons to expose pipelines that look alike on ordinary accuracy. |
| WorldModelPlanningAutopsy | Model-based RL and temporal abstraction | Test long-horizon planning under stochasticity, aliasing, delay, bias, and shift where one-step prediction hides compounding errors. |
In a direct sealed evaluation, a GPT-5.6 Pro chat instance made legal experiment calls on one fresh expert case per environment. All submissions were valid, but it selected the wrong 256-way mode every time, averaging 0.3128 with no passes. The five-call sample is small, yet it cleanly exposes forecasting, interval, and latent-estimation difficulty beyond competent experiment selection. The package passed 121 source and wheel tests plus ten replays.
We publish aggregate behavior and task structure, while withholding generated instances, hidden labels, exact successful probes, private checks, and solution trajectories.
Shown with its provenance and limitations; it is not a performance guarantee.
Reported result from the evaluation artifact supplied with this package.
As identified by the supplied artifact.
1 reported runs.
gpt56_blind_report.md
Machine-readable provenance and the exact displayed metric are available in results.json.
The paid ZIP will live in a private R2 bucket. Vercel authorizes the buyer and issues a 2–5 minute object URL; R2 serves the bytes directly.
Authenticated buyer + entitlement check
+ private R2 object + 2–5 minute signed URL
= direct, auditable download
Package SHA-256
14448e1c7a42dd2692ef3ea9bd9e39158e5ba306ca23472431da19ba787d5a0cOne purchase licenses this identified item to one legal organisation for worldwide, perpetual commercial model training, evaluation, research and development. Redistribution and resale of the package are not permitted.