Benchmark
How we check our own answers, and where we fall short. Small benchmark; judges are uncalibrated. Run r1 · 10/9/2026 · gpt-4.1-mini · judge gemini-2.5-flash
Safety gates
Answer quality
Citation hygiene
Retrieval
Are our answers improving?
What must pass before release
- fabricated_citation == 0must pass0 answers with a fabricated citation
- citation_validity == 100%must pass100.0%
- abstention >= 90%must pass90.0%
- groundedness >= 90% of claims supportedmust passjudge not yet calibrated98.4%
- p95 latency within budget12.1s (budget 25s)
- cost per answer within budget$0.0040 (budget $0.02)
- no previous release to compare withfirst release: regression gates start with the next run
| Category | Abstention | Correctness | Groundedness | Retrieval recall |
|---|---|---|---|---|
| recency | 0.75 | 0.75 | 1.00 | 1.00 |
| adversarial | 0.50 | 0.75 | 1.00 | 0.75 |
| methodology | 1.00 | 0.80 | 0.80 | 0.30 |
| definitional | 1.00 | 1.00 | 0.90 | 0.70 |
| multi_source | 1.00 | 0.50 | 1.00 | 0.33 |
| quantitative | 1.00 | 0.80 | 0.90 | 1.00 |
| unanswerable | 0.75 | — | — | — |
| tier_conflict | 1.00 | 0.67 | 1.00 | 0.33 |
Where answers pass or fail (development questions)
No previous release yet. We cannot say what improved or got worse.
No question-level results were recorded for this run.
Can we trust the AI judges?
Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.
Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.
Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.
Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.