Agent objective
- Interact with the supplied stateful environment.
- Produce verifier-checkable actions or artifacts.
- Maximise scalar reward under the package contract.
Five independent frontier mathematics environments spanning derived deformation theory, knot concordance, renormalization flows, proof complexity, and ergodic joinings. Agents jointly infer scientific structure and apparatus nuisance variables while navigating branch-dependent probes, scarce resources, irreversible transformations, destructive decoys, calibrated terminal decisions, and sealed predictions.
Five independent frontier mathematics environments spanning derived deformation theory, knot concordance, renormalization flows, proof complexity, and ergodic joinings. Agents jointly infer scientific structure and apparatus nuisance variables while navigating branch-dependent probes, scarce resources, irreversible transformations, destructive decoys, calibrated terminal decisions, and sealed predictions.
Enough detail to understand the intellectual terrain; generated instances, hidden mechanisms, and solution paths remain inside the private package.
| Environment | Mathematical or technical frontier | Adaptive research problem |
|---|---|---|
| DerivedDeformationLab | Derived deformation and obstruction theory | Route tangent, obstruction, higher-operation, gauge, and miniversal evidence while managing irreversible base changes. |
| KnotConcordanceLab | Floer/Khovanov concordance and surgery | Coordinate coefficients, satellites, covers, mutation, and surgery while consuming distinct topological resources. |
| RenormalizationFlowLab | Nonlinear PDE blow-up and dynamic rescaling | Infer scaling and measurement clock jointly through profile, modulation, spectrum, virial, and defect diagnostics. |
| ProofComplexityLab | Proof lower bounds and restrictions | Preserve valid width, degree, rank, communication, and pseudoexpectation routes through irreversible transformations. |
| ErgodicJoiningLab | Ergodic systems and joinings | Distinguish intrinsic structure from recoding and induction before choosing identification, a bounded set, or abstention. |
On ten hard/expert calibration episodes, a bounded continuation-aware exact-likelihood control passed five, while static planning passed one and myopic planning none. The package's task narratives explain why: branch choices change which later measurements remain legal, and attractive decoys can destroy the fine diagnostic route. These are evaluator controls rather than model scores, but they provide direct evidence that observation-contingent planning matters.
We publish aggregate behavior and task structure, while withholding generated instances, hidden labels, exact successful probes, private checks, and solution trajectories.
Shown with its provenance and limitations; it is not a performance guarantee.
Reported result from the evaluation artifact supplied with this package.
As identified by the supplied artifact.
40 reported runs.
rl_env_scores.json
Machine-readable provenance and the exact displayed metric are available in results.json.
The paid ZIP will live in a private R2 bucket. Vercel authorizes the buyer and issues a 2–5 minute object URL; R2 serves the bytes directly.
Authenticated buyer + entitlement check
+ private R2 object + 2–5 minute signed URL
= direct, auditable download
Package SHA-256
a48e7822099a0b491fbd3f0cacf2fc018e77c238fdfe90b79246377dbe17e171One purchase licenses this identified item to one legal organisation for worldwide, perpetual commercial model training, evaluation, research and development. Redistribution and resale of the package are not permitted.