Agent objective
- Interact with the supplied stateful environment.
- Produce verifier-checkable actions or artifacts.
- Maximise scalar reward under the package contract.
Five research environments for investigating complex machine-learning systems rather than merely scoring outputs. Agents intervene across transformer circuits, model editing, graph reasoning, compression, and meta-learning; distinguish compositional mechanisms; estimate continuous properties; and forecast unseen outcomes with calibrated intervals from only five experiments.
Five research environments for investigating complex machine-learning systems rather than merely scoring outputs. Agents intervene across transformer circuits, model editing, graph reasoning, compression, and meta-learning; distinguish compositional mechanisms; estimate continuous properties; and forecast unseen outcomes with calibrated intervals from only five experiments.
Enough detail to understand the intellectual terrain; generated instances, hidden mechanisms, and solution paths remain inside the private package.
| Environment | Mathematical or technical frontier | Adaptive research problem |
|---|---|---|
| Mechanistic Circuit Surgery | Transformer interpretability and causal tracing | Combine ablations, patches, steering, normalization, and positional tests to separate primary paths from backup and mediation. |
| Model Editing Interference Lab | Model editing and locality | Test paraphrase transfer, semantic locality, interference, reversibility, persistence, and representation geometry. |
| Graph Reasoning Topology Lab | GNNs, topology, heterophily | Vary distance, bottlenecks, symmetry, depth, and corruption, then transfer to adversarial topology combinations. |
| Compression Stack Forensics | Quantization, pruning, distillation | Use quality, robustness, hardware, outlier, sparsity, and recovery probes to reverse-engineer similar-looking stacks. |
| Meta-Learning Adaptation Chamber | Few-shot and inner-loop adaptation | Manipulate task distance, label symmetry, support corruption, sequence, shots, and update depth to diagnose adaptation. |
The supplied executable, non-oracle policy averaged 0.5411 across 15 expert episodes and passed none. Experiment design was strong at 0.8866 and sealed forecasts reached 0.7335, yet only one hidden system was identified and every run failed uncertainty and continuous-latent gates. The benchmark therefore separates nominally informative experiments from a complete, calibrated ML investigation. The release reports 100/100 tests.
We publish aggregate behavior and task structure, while withholding generated instances, hidden labels, exact successful probes, private checks, and solution trajectories.
Shown with its provenance and limitations; it is not a performance guarantee.
Reproducible scores from the included non-oracle executable policy; the report says this was not a pure model rollout.
As identified by the supplied artifact.
15 reported runs.
notes.md
Machine-readable provenance and the exact displayed metric are available in results.json.
The paid ZIP will live in a private R2 bucket. Vercel authorizes the buyer and issues a 2–5 minute object URL; R2 serves the bytes directly.
Authenticated buyer + entitlement check
+ private R2 object + 2–5 minute signed URL
= direct, auditable download
Package SHA-256
29c903ff480a9d9e57dcefca99aaf9ce7d4f7bd63d4fd68cfe17f3678186cb3dOne purchase licenses this identified item to one legal organisation for worldwide, perpetual commercial model training, evaluation, research and development. Redistribution and resale of the package are not permitted.