Agent objective
- Interact with the supplied stateful environment.
- Produce verifier-checkable actions or artifacts.
- Maximise scalar reward under the package contract.
Five machine-learning forensics environments where agents diagnose compositional systems through only five audits. Each expert case combines mechanisms, latent continuous effects, model mismatch, heavy-tailed observations, and twelve sealed forecasts, so success requires crossing diagnostic regimes, calibrated uncertainty, and predicting average and worst-case behavior—not merely naming a plausible stack.
Five machine-learning forensics environments where agents diagnose compositional systems through only five audits. Each expert case combines mechanisms, latent continuous effects, model mismatch, heavy-tailed observations, and twelve sealed forecasts, so success requires crossing diagnostic regimes, calibrated uncertainty, and predicting average and worst-case behavior—not merely naming a plausible stack.
Enough detail to understand the intellectual terrain; generated instances, hidden mechanisms, and solution paths remain inside the private package.
| Environment | Mathematical or technical frontier | Adaptive research problem |
|---|---|---|
| DatasetProvenanceForensics | Data leakage, quality, subgroup robustness | Select audits that separate overlapping pipeline pathologies and forecast unseen deduplication, temporal, and rare-slice conditions. |
| OptimizerPhasePortrait | Optimization dynamics and stability | Probe curvature, noise, shocks, schedules, and delayed evaluation to distinguish mechanisms hidden by ordinary loss curves. |
| FederatedProtocolAutopsy | Federated learning, robustness, privacy | Audit heterogeneity, attacks, participation, clipping, communication, and revisits to recover a hidden protocol. |
| RetrievalGroundingForensics | Retrieval-augmented generation | Separate retrieval improvements from grounding and abstention across stale, duplicated, ambiguous, and evidence-poor corpora. |
| DiffusionSamplerInterrogatory | Diffusion training and sampling | Stress denoising, guidance, SNR mismatch, outliers, prompts, and seeds to infer the hidden sampling stack. |
GPT-5.6 Pro ran one public-interface expert episode per environment and averaged 0.2951 with no passes. It recovered 17/30 individual mechanism bits but missed every complete system; experiment design was stronger than sealed forecasting and calibration. The result exposes the difference between recognizing pieces of an ML stack and building a causal diagnosis that transfers. The package passed 85/85 tests and all five fixtures replayed.
We publish aggregate behavior and task structure, while withholding generated instances, hidden labels, exact successful probes, private checks, and solution trajectories.
Shown with its provenance and limitations; it is not a performance guarantee.
Reported result from the evaluation artifact supplied with this package.
As identified by the supplied artifact.
5 reported runs.
rlvr_gpt56_results.zip:rlvr_gpt56_results/summary.json
Machine-readable provenance and the exact displayed metric are available in results.json.
The paid ZIP will live in a private R2 bucket. Vercel authorizes the buyer and issues a 2–5 minute object URL; R2 serves the bytes directly.
Authenticated buyer + entitlement check
+ private R2 object + 2–5 minute signed URL
= direct, auditable download
Package SHA-256
0407133862171930c1704aeef25d8a7396ec6dfb389db6926f10433d2a59e399One purchase licenses this identified item to one legal organisation for worldwide, perpetual commercial model training, evaluation, research and development. Redistribution and resale of the package are not permitted.