Agent objective
- Interact with the supplied stateful environment.
- Produce verifier-checkable actions or artifacts.
- Maximise scalar reward under the package contract.
A private-judge environment for diagnosing and repairing procedural, multi-fault analytics incidents across SQL, restricted Python, contracts, orchestration, warehouse state, lineage, logs, and dashboards. Agents must gather evidence, patch durable artifacts, run targeted recovery, and submit one structured diagnosis that survives hidden rebuild, idempotency, partition-isolation, and future-replay checks.
A private-judge environment for diagnosing and repairing procedural, multi-fault analytics incidents across SQL, restricted Python, contracts, orchestration, warehouse state, lineage, logs, and dashboards. Agents must gather evidence, patch durable artifacts, run targeted recovery, and submit one structured diagnosis that survives hidden rebuild, idempotency, partition-isolation, and future-replay checks.
Enough detail to understand the intellectual terrain; generated instances, hidden mechanisms, and solution paths remain inside the private package.
| Environment | Mathematical or technical frontier | Adaptive research problem |
|---|---|---|
| Schema and contract drift | Mixed producer versions and metadata | Repair SQL, restricted Python, YAML contracts, and semantic definitions so fixes survive new schemas and future replay. |
| Incremental correctness | Watermarks, deduplication, tombstones | Recover late arrivals, authoritative versions, retries, and deletions without corrupting unaffected partitions. |
| Business-time semantics | SCD joins, FX, refunds, settlement dates | Preserve effective-dated dimensions, daily conversion, partial reversals, and dashboard date meaning. |
| Operational recovery | Lineage, asynchronous jobs, cache and publication | Queue minimal dependency-ordered repairs, validate canaries, refresh stale views, and avoid an over-budget full rebuild. |
| Incident diagnosis | Three to five interacting faults | Separate true causes from decoys and submit exact assets, ranges, evidence, verification, and residual risk. |
A frozen task-aware policy developed on seeds 20–39 scored 100/100 on all 50 disjoint held-out incidents, covering 45 distinct fault combinations, while using about 58% of compute budget. This establishes broad procedural solvability, not average model performance. The package passed 14/14 tests. A separate HTTP stress test found six sessions with SQLite thread-affinity failures, so the core direct judge is strong but the HTTP deployment still needs concurrency hardening.
We publish aggregate behavior and task structure, while withholding generated instances, hidden labels, exact successful probes, private checks, and solution trajectories.
Shown with its provenance and limitations; it is not a performance guarantee.
Frozen after development seeds 20–39, then evaluated on 50 disjoint held-out seeds with 50/50 successes.
As identified by the supplied artifact.
50 reported runs.
data_pipeline_lab_v2_score_summary.json
Machine-readable provenance and the exact displayed metric are available in results.json.
The paid ZIP will live in a private R2 bucket. Vercel authorizes the buyer and issues a 2–5 minute object URL; R2 serves the bytes directly.
Authenticated buyer + entitlement check
+ private R2 object + 2–5 minute signed URL
= direct, auditable download
Package SHA-256
0ae856b31d79a59cbaaa254087b8118a34177521c263ca0407e493d438c54508One purchase licenses this identified item to one legal organisation for worldwide, perpetual commercial model training, evaluation, research and development. Redistribution and resale of the package are not permitted.