Agent objective
- Interact with the supplied stateful environment.
- Produce verifier-checkable actions or artifacts.
- Maximise scalar reward under the package contract.
Seventeen mathematics environments that turn finite extremal problems and audited reasoning failures into stateful, verifier-backed work. Thirteen demand exact combinatorial witnesses and sealed optima; four require a concrete obstruction, a repaired argument, preserved valid claims, and a consistent claim ledger under deterministic dual verification.
Seventeen mathematics environments that turn finite extremal problems and audited reasoning failures into stateful, verifier-backed work. Thirteen demand exact combinatorial witnesses and sealed optima; four require a concrete obstruction, a repaired argument, preserved valid claims, and a consistent claim ledger under deterministic dual verification.
Enough detail to understand the intellectual terrain; generated instances, hidden mechanisms, and solution paths remain inside the private package.
| Environment | Mathematical or technical frontier | Adaptive research problem |
|---|---|---|
| 13 finite-certificate tasks | Extremal set systems, additive and multiplicative combinatorics | Construct exact witnesses and prove sealed optima across finite geometry, divisor constraints, intersections, graph coloring, and modular closure. |
| Cayley parallelogram repair | Abelian Cayley graphs | Find a concrete cycle obstruction, repair the girth claim, and reconcile every dependent claim in the ledger. |
| Directed-corner repair | Additive constructions | Diagnose an orientation-sensitive failure and preserve only the claims supported by exact finite checks. |
| Periodic-square repair | Periodic combinatorics | Locate a delayed counterexample and replace the failed general claim with a calibrated finite-scope repair. |
| CRT-overlap repair | Chinese remainder constructions | Expose residue-cover multiplicity and update the argument, evidence, verdict, and ledger consistently. |
A fresh isolated symbolic agent earned 13.20/17 terminal reward and full credit on all 13 finite-certificate tasks, but none of four repair tasks. It found the mathematical obstructions and repairs, yet each repair was capped because its claims did not reconcile exactly with the sealed ledger. That gap makes the suite unusually effective at testing rigorous stateful argument maintenance, not just mathematical insight. Validation passed 36/36 tests.
We publish aggregate behavior and task structure, while withholding generated instances, hidden labels, exact successful probes, private checks, and solution trajectories.
Shown with its provenance and limitations; it is not a performance guarantee.
Fresh authenticated online single-pass evaluation of all 17 tasks at seed 0. The deterministic symbolic/exhaustive-search agent was not an LLM; it earned full credit on 13/17 episodes without replay, snapshots, forks, or post-grade retries.
As identified by the supplied artifact.
17 reported runs.
ErdosRL_real_agent_artifacts.zip · real_agent_score_report.json
Machine-readable provenance and the exact displayed metric are available in results.json.
The paid ZIP will live in a private R2 bucket. Vercel authorizes the buyer and issues a 2–5 minute object URL; R2 serves the bytes directly.
Authenticated buyer + entitlement check
+ private R2 object + 2–5 minute signed URL
= direct, auditable download
Package SHA-256
9a4c63e554cb3809d39eefe3630c43bf195b5c55dfb45d528033177345fc3b45One purchase licenses this identified item to one legal organisation for worldwide, perpetual commercial model training, evaluation, research and development. Redistribution and resale of the package are not permitted.