Benchmark
How we check our own answers, and where we fall short. Small benchmark; judges are uncalibrated. Run r2 · 10/9/2026 · gpt-4.1-mini · judge gemini-2.5-flash
Safety gates
Answer quality
Citation hygiene
Retrieval
Are our answers improving?
What must pass before release
- fabricated_citation == 0must pass0 answers with a fabricated citation
- citation_validity == 100%must pass100.0%
- abstention >= 90%must pass90.0%
- groundedness >= 90% of claims supportedmust passjudge not yet calibrated95.5%
- p95 latency within budget13.6s (budget 25s)
- cost per answer within budget$0.0038 (budget $0.02)
- retrieval_recall: no significant regression vs previousΔ +0.040 (95% CI -0.040 to +0.120, n=50)
- abstention: no significant regression vs previousΔ +0.000 (95% CI -0.083 to +0.083, n=60)
- citation_validity: no significant regression vs previousΔ +0.000 (95% CI +0.000 to +0.000, n=60)
- citation_coverage: no significant regression vs previousΔ +0.000 (95% CI +0.000 to +0.000, n=60)
- quote_length: no significant regression vs previousΔ +0.000 (95% CI +0.000 to +0.000, n=60)
- tier_compliance: no significant regression vs previousΔ +0.017 (95% CI +0.000 to +0.050, n=60)
- groundedness: no significant regression vs previousjudge not yet calibratedΔ -0.005 (95% CI -0.032 to +0.020, n=43)
- correctness: no significant regression vs previousmust passjudge not yet calibratedΔ -0.020 (95% CI -0.160 to +0.120, n=50)
- conflict_disclosure: no significant regression vs previousjudge not yet calibratedΔ +0.167 (95% CI +0.000 to +0.500, n=6)
- relevance: no significant regression vs previousjudge not yet calibratedΔ +0.085 (95% CI +0.000 to +0.186, n=59)
- tier_compliance_judge: no significant regression vs previousΔ -0.023 (95% CI -0.070 to +0.000, n=43)
| Category | Abstention | Correctness | Groundedness | Retrieval recall |
|---|---|---|---|---|
| recency | 1.00 | 0.75 | 0.50 | 1.00 |
| adversarial | 0.33 | 1.00 | — | 1.00 |
| methodology | 1.00 | 0.70 | 0.90 | 0.40 |
| definitional | 1.00 | 1.00 | 1.00 | 0.70 |
| multi_source | 1.00 | 0.67 | 1.00 | 0.33 |
| quantitative | 0.80 | 0.60 | 1.00 | 1.00 |
| unanswerable | 1.00 | — | — | — |
| tier_conflict | 1.00 | 0.67 | 0.83 | 0.33 |
Where answers pass or fail (development questions)
Compared with r1-all-202610090943: 0 fixed · 0 regressed · 0 unchanged
No question-level results were recorded for this run.
Can we trust the AI judges?
Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.
Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.
Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.
Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.