MathematicsProofRepair

ProofRepair

Five environments for auditing and repairing flawed mathematical arguments under evidence and tool constraints. Agents inspect interleaved branches, search a frozen corpus, test theorem applicability, construct executable counterexamples, recover nonlocal dependencies, locate the earliest fatal gap, and propose minimal repairs that survive controlled hidden variants.

MathematicsAdaptive reasoningRLVR
Version0.3.0
Environments5
RewardScalar · 0–1
DeliveryPrivate ZIP
01 Task contract

A real, versioned RL environment.

Five environments for auditing and repairing flawed mathematical arguments under evidence and tool constraints. Agents inspect interleaved branches, search a frozen corpus, test theorem applicability, construct executable counterexamples, recover nonlocal dependencies, locate the earliest fatal gap, and propose minimal repairs that survive controlled hidden variants.

Agent objective

  • Interact with the supplied stateful environment.
  • Produce verifier-checkable actions or artifacts.
  • Maximise scalar reward under the package contract.

Evaluation

  • 5 verifier-backed environments.
  • Reported reward range 0–1.
  • Package-specific public and private checks.

Delivery boundary

  • Private object stored in Cloudflare R2.
  • Authenticated entitlement required.
  • Short-lived signed URL per download.
02 What you will work on

Distinct environments, one demanding research contract.

Enough detail to understand the intellectual terrain; generated instances, hidden mechanisms, and solution paths remain inside the private package.

EnvironmentMathematical or technical frontierAdaptive research problem
AlgebraicFaultlineAlgebra and number theoryAudit divisibility, cancellation, inverse, sign, and polynomial arguments, then repair missing hypotheses across hidden variants.
LinearAlgebraSurgeryLinear algebra and matrix theoryDiagnose invalid claims about positivity, products, spectra, rank, nilpotence, and inverses, then rebuild dependencies.
AnalysisQuantifierLabReal and functional analysisTrack quantifiers and global hypotheses through convergence, completeness, differentiation, integration, and local-to-global steps.
ProbabilitySetAuditProbability and conditioningTest dependence structure and set corrections, then repair the earliest inference without overclaiming convergence or independence.
CombinatorialInductionLabCombinatorics and graph argumentsAudit bases, multiplicities, recurrences, assumptions, and inclusion–exclusion dependencies, then cover every relevant case.
03 Why it is interesting

What the supplied evaluation reveals.

The supplied five-episode hard test rollout averaged 0.9472 and passed four tasks. It localized every earliest gap and complete defect set, yet AlgebraicFaultline failed because one repair did not survive hidden variants. That distinction is the benchmark's core value: diagnosing a flaw is not enough unless the repair is robust. The package reports 190 tests, 1,200 stress cases, and 1,200 exact replays.

We publish aggregate behavior and task structure, while withholding generated instances, hidden labels, exact successful probes, private checks, and solution trajectories.

04 Supplied evaluation

Observed evaluation result.

Shown with its provenance and limitations; it is not a performance guarantee.

i
Methodology matters

Reported result from the evaluation artifact supplied with this package.

Evaluated system / policyGPT-5.6 Pro-assisted evaluation

As identified by the supplied artifact.

Mean reward0.9472

5 reported runs.

Result artifactIncluded

ulam_run_bundle.zip:ulam_score_report.json

Public result record

Machine-readable provenance and the exact displayed metric are available in results.json.

Open result JSON
05 Private delivery

The package stays off the public website.

The paid ZIP will live in a private R2 bucket. Vercel authorizes the buyer and issues a 2–5 minute object URL; R2 serves the bytes directly.

Included with purchase

  • Exact package version 0.3.0
  • Environment and task contracts
  • Verifier or scoring interface
  • Supplied reference/evaluation artifacts
  • Purchase record and licence v1.1
delivery flow
Authenticated buyer + entitlement check
+ private R2 object + 2–5 minute signed URL
= direct, auditable download

Package SHA-256
2a051f0db17564364d0d29ef6b16806af11c197bd9823d6f5eb6cd17c574cb3e
06 Licence v1.1

Commercial use, without exclusivity.

One purchase licenses this identified item to one legal organisation for worldwide, perpetual commercial model training, evaluation, research and development. Redistribution and resale of the package are not permitted.

Read the full licenceYotta Content LTD · business customers only
ProofRepair

Ready to add this environment?

Back to marketplace
Stripe checkout

Business purchase confirmation

Sign in or create an account, then complete secure Stripe Checkout. Access is granted only by the verified payment webhook.