Benchmark

How we check our own answers, and where we fall short. Small benchmark; judges are uncalibrated. Run r4 · 10/10/2026 · gpt-4.1-mini · judge gemini-2.5-flash

split alln=60p50 8.4 s · p95 13.0 s$0.0038/answer · total $0.23judge prompt judges-v2LangSmith experimentr1r2r3r4

Safety gates

Abstention
0.93
95% CI [0.87, 0.98] · n=60
Citation validity
1.00
95% CI [1.00, 1.00] · n=60
Fabricated citations
0none found

Answer quality

Answer relevance
uncalibrated
0.97
95% CI [0.92, 1.00] · n=59
Correctness
uncalibrated
0.76
95% CI [0.64, 0.88] · n=50
Groundedness
uncalibrated
0.94
95% CI [0.89, 0.99] · n=46
Conflict disclosure
uncalibrated
0.50
95% CI [0.17, 0.83] · n=6

Citation hygiene

Quote length
1.00
95% CI [1.00, 1.00] · n=60
Tier wording matches
0.92
95% CI [0.85, 0.98] · n=60
Tier compliance
0.97
95% CI [0.92, 1.00] · n=60
Citation density (max 2 per sentence)
0.97
95% CI [0.92, 1.00] · n=60
Citation coverage
1.00
95% CI [1.00, 1.00] · n=60
Tier compliance (judge)
uncalibrated
0.89
95% CI [0.80, 0.98] · n=46

Retrieval

Retrieval recall
0.70
95% CI [0.58, 0.82] · n=50

Are our answers improving?

What must pass before release

  • fabricated_citation == 0must pass
    0 answers with a fabricated citation
  • citation_validity == 100%must pass
    100.0%
  • abstention >= 90%must pass
    93.3%
  • groundedness >= 90% of claims supportedmust passjudge not yet calibrated
    94.5%
  • p95 latency within budget
    13.0s (budget 25s)
  • cost per answer within budget
    $0.0038 (budget $0.02)
  • retrieval_recall: no significant regression vs previous
    Δ -0.040 (95% CI -0.100 to +0.000, n=50)
  • abstention: no significant regression vs previous
    Δ +0.050 (95% CI +0.000 to +0.100, n=60)
  • citation_validity: no significant regression vs previous
    Δ +0.000 (95% CI +0.000 to +0.000, n=60)
  • citation_coverage: no significant regression vs previous
    Δ +0.000 (95% CI +0.000 to +0.000, n=60)
  • quote_length: no significant regression vs previous
    Δ +0.000 (95% CI +0.000 to +0.000, n=60)
  • tier_compliance: no significant regression vs previous
    Δ +0.033 (95% CI +0.000 to +0.083, n=60)
  • tier_wording: no significant regression vs previous
    Δ -0.017 (95% CI -0.100 to +0.067, n=60)
  • citation_density: no significant regression vs previous
    Δ -0.033 (95% CI -0.083 to +0.000, n=60)
  • groundedness: no significant regression vs previousjudge not yet calibrated
    Δ +0.021 (95% CI -0.009 to +0.053, n=45)
  • correctness: no significant regression vs previousmust passjudge not yet calibrated
    Δ +0.000 (95% CI -0.120 to +0.140, n=50)
  • conflict_disclosure: no significant regression vs previousjudge not yet calibrated
    Δ +0.000 (95% CI -0.500 to +0.500, n=6)
  • relevance: no significant regression vs previousjudge not yet calibrated
    Δ +0.000 (95% CI -0.068 to +0.068, n=59)
  • tier_compliance_judge: no significant regression vs previous
    Δ +0.022 (95% CI -0.067 to +0.111, n=45)
CategoryAbstention Correctness Groundedness Retrieval recall
recency1.001.000.501.00
adversarial0.331.00—1.00
methodology1.000.500.900.30
definitional1.001.001.001.00
multi_source1.000.670.830.33
quantitative1.000.701.001.00
unanswerable1.00———
tier_conflict1.000.670.670.33

Where answers pass or fail (development questions)

Compared with r3-all-202610100730: 0 fixed · 0 regressed · 0 unchanged

No question-level results were recorded for this run.

Can we trust the AI judges?

Groundedness
gemini-2.5-flash · judges-v2

Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.

Correctness
gemini-2.5-flash · judges-v2

Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.

Conflict disclosure
gemini-2.5-flash · judges-v2

Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.

Answer relevance
gemini-2.5-flash · judges-v2

Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.