Agent objective
- Interact with the supplied stateful environment.
- Produce verifier-checkable actions or artifacts.
- Maximise scalar reward under the package contract.
Five measurable-tier mathematics environments spanning Bridgeland wall crossing, KAM resonance, random-matrix edges, theta characteristics, and proof-net dynamics. Agents receive invariant starting evidence and only two noisy measurements, then must choose complementary probes that expose hidden geometry while maintaining calibrated uncertainty and predicting disjoint sealed functionals.
Five measurable-tier mathematics environments spanning Bridgeland wall crossing, KAM resonance, random-matrix edges, theta characteristics, and proof-net dynamics. Agents receive invariant starting evidence and only two noisy measurements, then must choose complementary probes that expose hidden geometry while maintaining calibrated uncertainty and predicting disjoint sealed functionals.
Enough detail to understand the intellectual terrain; generated instances, hidden mechanisms, and solution paths remain inside the private package.
| Environment | Mathematical or technical frontier | Adaptive research problem |
|---|---|---|
| BridgelandWallCipher | Derived categories and stability conditions | Select wall- and phase-sensitive probes to recover hidden stability geometry under exact categorical invariants. |
| KAMResonanceAtlas | KAM theory and resonant normal forms | Combine resonance, splitting, and transport evidence across complementary dynamical regimes. |
| RandomMatrixEdgeOracle | Random matrices and edge asymptotics | Distinguish latent edge mechanisms through bulk, tail, deformation, and finite-size measurements. |
| ThetaCharacteristicXRay | Theta functions and abelian varieties | Use phase- and characteristic-sensitive analytic identities beyond shared coarse geometric summaries. |
| ProofNetGoIProbe | Linear logic and Geometry of Interaction | Probe execution and path geometry to separate proof structures that share ordinary correctness summaries. |
The packaged 80-episode control sweep found no strict passes. Its strongest exact-likelihood greedy policy still achieved 0.7336 mean reward, 0.9297 experiment design, and 0.9115 held-out prediction, but only 30% identification. This is useful evidence of an intentionally unsaturated frontier: locally excellent two-probe plans leave substantial posterior ambiguity. The scores are controls, not external model results.
We publish aggregate behavior and task structure, while withholding generated instances, hidden labels, exact successful probes, private checks, and solution trajectories.
Shown with its provenance and limitations; it is not a performance guarantee.
Reported result from the evaluation artifact supplied with this package.
As identified by the supplied artifact.
20 reported runs.
ulam_fresh_scores.json
Machine-readable provenance and the exact displayed metric are available in results.json.
The paid ZIP will live in a private R2 bucket. Vercel authorizes the buyer and issues a 2–5 minute object URL; R2 serves the bytes directly.
Authenticated buyer + entitlement check
+ private R2 object + 2–5 minute signed URL
= direct, auditable download
Package SHA-256
041b2324813be053d9fe0c6fcdbc71c26f7c74144862191d5a3a421510b66df0One purchase licenses this identified item to one legal organisation for worldwide, perpetual commercial model training, evaluation, research and development. Redistribution and resale of the package are not permitted.