Benchmark
How we check our own answers, and where we fall short. Small benchmark; judges are uncalibrated. Run r4 · 10/10/2026 · gpt-4.1-mini · judge gemini-2.5-flash
Safety gates
Answer quality
Citation hygiene
Retrieval
Are our answers improving?
What must pass before release
- fabricated_citation == 0must pass0 answers with a fabricated citation
- citation_validity == 100%must pass100.0%
- abstention >= 90%must pass93.3%
- groundedness >= 90% of claims supportedmust passjudge not yet calibrated94.5%
- p95 latency within budget13.0s (budget 25s)
- cost per answer within budget$0.0038 (budget $0.02)
- retrieval_recall: no significant regression vs previousΔ -0.040 (95% CI -0.100 to +0.000, n=50)
- abstention: no significant regression vs previousΔ +0.050 (95% CI +0.000 to +0.100, n=60)
- citation_validity: no significant regression vs previousΔ +0.000 (95% CI +0.000 to +0.000, n=60)
- citation_coverage: no significant regression vs previousΔ +0.000 (95% CI +0.000 to +0.000, n=60)
- quote_length: no significant regression vs previousΔ +0.000 (95% CI +0.000 to +0.000, n=60)
- tier_compliance: no significant regression vs previousΔ +0.033 (95% CI +0.000 to +0.083, n=60)
- tier_wording: no significant regression vs previousΔ -0.017 (95% CI -0.100 to +0.067, n=60)
- citation_density: no significant regression vs previousΔ -0.033 (95% CI -0.083 to +0.000, n=60)
- groundedness: no significant regression vs previousjudge not yet calibratedΔ +0.021 (95% CI -0.009 to +0.053, n=45)
- correctness: no significant regression vs previousmust passjudge not yet calibratedΔ +0.000 (95% CI -0.120 to +0.140, n=50)
- conflict_disclosure: no significant regression vs previousjudge not yet calibratedΔ +0.000 (95% CI -0.500 to +0.500, n=6)
- relevance: no significant regression vs previousjudge not yet calibratedΔ +0.000 (95% CI -0.068 to +0.068, n=59)
- tier_compliance_judge: no significant regression vs previousΔ +0.022 (95% CI -0.067 to +0.111, n=45)
| Category | Abstention | Correctness | Groundedness | Retrieval recall |
|---|---|---|---|---|
| recency | 1.00 | 1.00 | 0.50 | 1.00 |
| adversarial | 0.33 | 1.00 | — | 1.00 |
| methodology | 1.00 | 0.50 | 0.90 | 0.30 |
| definitional | 1.00 | 1.00 | 1.00 | 1.00 |
| multi_source | 1.00 | 0.67 | 0.83 | 0.33 |
| quantitative | 1.00 | 0.70 | 1.00 | 1.00 |
| unanswerable | 1.00 | — | — | — |
| tier_conflict | 1.00 | 0.67 | 0.67 | 0.33 |
Where answers pass or fail (development questions)
Compared with r3-all-202610100730: 0 fixed · 0 regressed · 0 unchanged
No question-level results were recorded for this run.
Can we trust the AI judges?
Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.
Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.
Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.
Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.