Agent objective
- Interact with the supplied stateful environment.
- Produce verifier-checkable actions or artifacts.
- Maximise scalar reward under the package contract.
Five machine-learning research environments for adaptive diagnosis under model misspecification. They span causal mechanisms, scaling behavior, representation geometry, continual learning, and uncertainty under distribution shift, requiring costly diagnostics, compositional mechanism inference, continuous-variation tracking, and explicit forecasts under heavy-tailed, common-mode error.
Five machine-learning research environments for adaptive diagnosis under model misspecification. They span causal mechanisms, scaling behavior, representation geometry, continual learning, and uncertainty under distribution shift, requiring costly diagnostics, compositional mechanism inference, continuous-variation tracking, and explicit forecasts under heavy-tailed, common-mode error.
Enough detail to understand the intellectual terrain; generated instances, hidden mechanisms, and solution paths remain inside the private package.
| Environment | Mathematical or technical frontier | Adaptive research problem |
|---|---|---|
| CausalMechanismLab | Causal discovery and intervention design | Combine interventions with mean and covariance diagnostics to separate direct, mediated, and confounded effects under uncertain coefficients. |
| ScalingLawCartographer | Neural scaling laws and compute allocation | Choose training configurations that separate coupled scaling mechanisms while supporting forecasts beyond the observed regime. |
| RepresentationXRay | Representation geometry, shortcuts, invariance | Audit layers and counterfactual batches through compressed probes when mechanisms share rank or separation signatures. |
| ContinualLearningAutopsy | Forgetting, replay, adapters, path dependence | Design curricula and retention/transfer measurements revealing which mechanisms act across conflicting and revisited tasks. |
| UncertaintyShiftDoctor | Calibration, OOD detection, selective prediction | Pair coverage, risk, calibration, entropy, and OOD diagnostics when pipeline components cancel in aggregate metrics. |
A nominal Bayesian reference was evaluated on 100 fresh expert episodes outside the development seed range. It reached 0.7550 mean reward but only 40 passes and 51 correct identifications, with pass rates ranging from 11/20 to 3/20 across domains. This makes the suite useful for studying where one-step information gain breaks under extrapolation, path dependence, and nuisance mismatch. The package passed 66/66 tests.
We publish aggregate behavior and task structure, while withholding generated instances, hidden labels, exact successful probes, private checks, and solution trajectories.
Shown with its provenance and limitations; it is not a performance guarantee.
Fresh 100-episode expert calibration; 40% pass rate.
As identified by the supplied artifact.
100 reported runs.
greedy_bayes_expert_summary.csv
Machine-readable provenance and the exact displayed metric are available in results.json.
The paid ZIP will live in a private R2 bucket. Vercel authorizes the buyer and issues a 2–5 minute object URL; R2 serves the bytes directly.
Authenticated buyer + entitlement check
+ private R2 object + 2–5 minute signed URL
= direct, auditable download
Package SHA-256
e439aad667706019a88761a145f2e952b2f40ce011e43077b7992581458357a9One purchase licenses this identified item to one legal organisation for worldwide, perpetual commercial model training, evaluation, research and development. Redistribution and resale of the package are not permitted.