Benchmark
How we check our own answers, and where we fall short. Small benchmark; judges are uncalibrated. Run r3 · 10/10/2026 · gpt-4.1-mini · judge gemini-2.5-flash
Safety gates
Answer quality
Citation hygiene
Retrieval
Are our answers improving?
What must pass before release
- fabricated_citation == 0must pass0 answers with a fabricated citation
- citation_validity == 100%must pass100.0%
- abstention >= 90%must pass88.3%
- groundedness >= 90% of claims supportedmust passjudge not yet calibrated92.2%
- p95 latency within budget11.6s (budget 25s)
- cost per answer within budget$0.0037 (budget $0.02)
- retrieval_recall: no significant regression vs previousΔ +0.080 (95% CI +0.020 to +0.160, n=50)
- abstention: no significant regression vs previousΔ -0.017 (95% CI -0.083 to +0.050, n=60)
- citation_validity: no significant regression vs previousΔ +0.000 (95% CI +0.000 to +0.000, n=60)
- citation_coverage: no significant regression vs previousΔ +0.000 (95% CI +0.000 to +0.000, n=60)
- quote_length: no significant regression vs previousΔ +0.000 (95% CI +0.000 to +0.000, n=60)
- tier_compliance: no significant regression vs previousΔ -0.050 (95% CI -0.133 to +0.017, n=60)
- groundedness: no significant regression vs previousjudge not yet calibratedΔ -0.035 (95% CI -0.071 to -0.005, n=43)
- correctness: no significant regression vs previousmust passjudge not yet calibratedΔ +0.000 (95% CI -0.140 to +0.140, n=50)
- conflict_disclosure: no significant regression vs previousjudge not yet calibratedΔ -0.167 (95% CI -0.667 to +0.333, n=6)
- relevance: no significant regression vs previousjudge not yet calibratedΔ -0.017 (95% CI -0.068 to +0.034, n=59)
- tier_compliance_judge: no significant regression vs previousΔ -0.093 (95% CI -0.209 to +0.000, n=43)
| Category | Abstention | Correctness | Groundedness | Retrieval recall |
|---|---|---|---|---|
| recency | 1.00 | 1.00 | 0.25 | 1.00 |
| adversarial | 0.33 | 1.00 | — | 1.00 |
| methodology | 1.00 | 0.80 | 0.90 | 0.40 |
| definitional | 1.00 | 0.90 | 0.90 | 1.00 |
| multi_source | 0.83 | 0.50 | 0.80 | 0.50 |
| quantitative | 1.00 | 0.70 | 1.00 | 1.00 |
| unanswerable | 0.75 | — | — | — |
| tier_conflict | 1.00 | 0.50 | 0.50 | 0.33 |
Where answers pass or fail (development questions)
Compared with r2-all-202610090951: 0 fixed · 0 regressed · 0 unchanged
No question-level results were recorded for this run.
Can we trust the AI judges?
Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.
Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.
Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.
Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.