Benchmark

How we check our own answers, and where we fall short. Small benchmark; judges are uncalibrated. Run r3 · 10/10/2026 · gpt-4.1-mini · judge gemini-2.5-flash

split alln=60p50 8.2 s · p95 11.6 s$0.0037/answer · total $0.22judge prompt judges-v2LangSmith experimentr1r2r3r4

Safety gates

Abstention
0.88
95% CI [0.80, 0.95] · n=60
Citation validity
1.00
95% CI [1.00, 1.00] · n=60
Fabricated citations
0none found

Answer quality

Answer relevance
uncalibrated
0.97
95% CI [0.92, 1.00] · n=59
Correctness
uncalibrated
0.76
95% CI [0.64, 0.88] · n=50
Groundedness
uncalibrated
0.92
95% CI [0.87, 0.97] · n=45
Conflict disclosure
uncalibrated
0.50
95% CI [0.17, 0.83] · n=6

Citation hygiene

Quote length
1.00
95% CI [1.00, 1.00] · n=60
Tier wording matches
0.93
95% CI [0.87, 0.98] · n=60
Tier compliance
0.93
95% CI [0.87, 0.98] · n=60
Citation density (max 2 per sentence)
1.00
95% CI [1.00, 1.00] · n=60
Citation coverage
1.00
95% CI [1.00, 1.00] · n=60
Tier compliance (judge)
uncalibrated
0.87
95% CI [0.76, 0.96] · n=45

Retrieval

Retrieval recall
0.74
95% CI [0.62, 0.86] · n=50

Are our answers improving?

What must pass before release

  • fabricated_citation == 0must pass
    0 answers with a fabricated citation
  • citation_validity == 100%must pass
    100.0%
  • abstention >= 90%must pass
    88.3%
  • groundedness >= 90% of claims supportedmust passjudge not yet calibrated
    92.2%
  • p95 latency within budget
    11.6s (budget 25s)
  • cost per answer within budget
    $0.0037 (budget $0.02)
  • retrieval_recall: no significant regression vs previous
    Δ +0.080 (95% CI +0.020 to +0.160, n=50)
  • abstention: no significant regression vs previous
    Δ -0.017 (95% CI -0.083 to +0.050, n=60)
  • citation_validity: no significant regression vs previous
    Δ +0.000 (95% CI +0.000 to +0.000, n=60)
  • citation_coverage: no significant regression vs previous
    Δ +0.000 (95% CI +0.000 to +0.000, n=60)
  • quote_length: no significant regression vs previous
    Δ +0.000 (95% CI +0.000 to +0.000, n=60)
  • tier_compliance: no significant regression vs previous
    Δ -0.050 (95% CI -0.133 to +0.017, n=60)
  • groundedness: no significant regression vs previousjudge not yet calibrated
    Δ -0.035 (95% CI -0.071 to -0.005, n=43)
  • correctness: no significant regression vs previousmust passjudge not yet calibrated
    Δ +0.000 (95% CI -0.140 to +0.140, n=50)
  • conflict_disclosure: no significant regression vs previousjudge not yet calibrated
    Δ -0.167 (95% CI -0.667 to +0.333, n=6)
  • relevance: no significant regression vs previousjudge not yet calibrated
    Δ -0.017 (95% CI -0.068 to +0.034, n=59)
  • tier_compliance_judge: no significant regression vs previous
    Δ -0.093 (95% CI -0.209 to +0.000, n=43)
CategoryAbstention Correctness Groundedness Retrieval recall
recency1.001.000.251.00
adversarial0.331.00—1.00
methodology1.000.800.900.40
definitional1.000.900.901.00
multi_source0.830.500.800.50
quantitative1.000.701.001.00
unanswerable0.75———
tier_conflict1.000.500.500.33

Where answers pass or fail (development questions)

Compared with r2-all-202610090951: 0 fixed · 0 regressed · 0 unchanged

No question-level results were recorded for this run.

Can we trust the AI judges?

Groundedness
gemini-2.5-flash · judges-v2

Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.

Correctness
gemini-2.5-flash · judges-v2

Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.

Conflict disclosure
gemini-2.5-flash · judges-v2

Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.

Answer relevance
gemini-2.5-flash · judges-v2

Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.