Benchmark

How we check our own answers, and where we fall short. Small benchmark; judges are uncalibrated. Run r2 · 10/9/2026 · gpt-4.1-mini · judge gemini-2.5-flash

split alln=60p50 8.1 s · p95 13.6 s$0.0038/answer · total $0.23judge prompt judges-v2LangSmith experimentr1r2r3r4

Safety gates

Abstention
0.90
95% CI [0.82, 0.97] · n=60
Citation validity
1.00
95% CI [1.00, 1.00] · n=60
Fabricated citations
0none found

Answer quality

Answer relevance
uncalibrated
0.98
95% CI [0.95, 1.00] · n=59
Correctness
uncalibrated
0.76
95% CI [0.64, 0.88] · n=50
Groundedness
uncalibrated
0.96
95% CI [0.90, 0.99] · n=44
Conflict disclosure
uncalibrated
0.67
95% CI [0.33, 1.00] · n=6

Citation hygiene

Quote length
1.00
95% CI [1.00, 1.00] · n=60
Tier compliance
0.98
95% CI [0.95, 1.00] · n=60
Citation coverage
1.00
95% CI [1.00, 1.00] · n=60
Tier compliance (judge)
uncalibrated
0.95
95% CI [0.89, 1.00] · n=44

Retrieval

Retrieval recall
0.66
95% CI [0.52, 0.78] · n=50

Are our answers improving?

What must pass before release

  • fabricated_citation == 0must pass
    0 answers with a fabricated citation
  • citation_validity == 100%must pass
    100.0%
  • abstention >= 90%must pass
    90.0%
  • groundedness >= 90% of claims supportedmust passjudge not yet calibrated
    95.5%
  • p95 latency within budget
    13.6s (budget 25s)
  • cost per answer within budget
    $0.0038 (budget $0.02)
  • retrieval_recall: no significant regression vs previous
    Δ +0.040 (95% CI -0.040 to +0.120, n=50)
  • abstention: no significant regression vs previous
    Δ +0.000 (95% CI -0.083 to +0.083, n=60)
  • citation_validity: no significant regression vs previous
    Δ +0.000 (95% CI +0.000 to +0.000, n=60)
  • citation_coverage: no significant regression vs previous
    Δ +0.000 (95% CI +0.000 to +0.000, n=60)
  • quote_length: no significant regression vs previous
    Δ +0.000 (95% CI +0.000 to +0.000, n=60)
  • tier_compliance: no significant regression vs previous
    Δ +0.017 (95% CI +0.000 to +0.050, n=60)
  • groundedness: no significant regression vs previousjudge not yet calibrated
    Δ -0.005 (95% CI -0.032 to +0.020, n=43)
  • correctness: no significant regression vs previousmust passjudge not yet calibrated
    Δ -0.020 (95% CI -0.160 to +0.120, n=50)
  • conflict_disclosure: no significant regression vs previousjudge not yet calibrated
    Δ +0.167 (95% CI +0.000 to +0.500, n=6)
  • relevance: no significant regression vs previousjudge not yet calibrated
    Δ +0.085 (95% CI +0.000 to +0.186, n=59)
  • tier_compliance_judge: no significant regression vs previous
    Δ -0.023 (95% CI -0.070 to +0.000, n=43)
CategoryAbstention Correctness Groundedness Retrieval recall
recency1.000.750.501.00
adversarial0.331.00—1.00
methodology1.000.700.900.40
definitional1.001.001.000.70
multi_source1.000.671.000.33
quantitative0.800.601.001.00
unanswerable1.00———
tier_conflict1.000.670.830.33

Where answers pass or fail (development questions)

Compared with r1-all-202610090943: 0 fixed · 0 regressed · 0 unchanged

No question-level results were recorded for this run.

Can we trust the AI judges?

Groundedness
gemini-2.5-flash · judges-v2

Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.

Correctness
gemini-2.5-flash · judges-v2

Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.

Conflict disclosure
gemini-2.5-flash · judges-v2

Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.

Answer relevance
gemini-2.5-flash · judges-v2

Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.