Benchmark

How we check our own answers, and where we fall short. Small benchmark; judges are uncalibrated. Run r1 · 10/9/2026 · gpt-4.1-mini · judge gemini-2.5-flash

split alln=60p50 7.4 s · p95 12.1 s$0.0040/answer · total $0.24judge prompt judges-v1LangSmith experimentr1r2r3r4

Safety gates

Abstention
0.90
95% CI [0.82, 0.97] · n=60
Citation validity
1.00
95% CI [1.00, 1.00] · n=60
Fabricated citations
0none found

Answer quality

Answer relevance
uncalibrated
0.90
95% CI [0.81, 0.97] · n=59
Correctness
uncalibrated
0.78
95% CI [0.66, 0.88] · n=50
Groundedness
uncalibrated
0.98
95% CI [0.97, 1.00] · n=46
Conflict disclosure
uncalibrated
0.50
95% CI [0.17, 0.83] · n=6

Citation hygiene

Quote length
1.00
95% CI [1.00, 1.00] · n=60
Tier compliance
0.97
95% CI [0.92, 1.00] · n=60
Citation coverage
1.00
95% CI [1.00, 1.00] · n=60
Tier compliance (judge)
uncalibrated
0.98
95% CI [0.93, 1.00] · n=46

Retrieval

Retrieval recall
0.62
95% CI [0.48, 0.74] · n=50

Are our answers improving?

What must pass before release

  • fabricated_citation == 0must pass
    0 answers with a fabricated citation
  • citation_validity == 100%must pass
    100.0%
  • abstention >= 90%must pass
    90.0%
  • groundedness >= 90% of claims supportedmust passjudge not yet calibrated
    98.4%
  • p95 latency within budget
    12.1s (budget 25s)
  • cost per answer within budget
    $0.0040 (budget $0.02)
  • no previous release to compare with
    first release: regression gates start with the next run
CategoryAbstention Correctness Groundedness Retrieval recall
recency0.750.751.001.00
adversarial0.500.751.000.75
methodology1.000.800.800.30
definitional1.001.000.900.70
multi_source1.000.501.000.33
quantitative1.000.800.901.00
unanswerable0.75———
tier_conflict1.000.671.000.33

Where answers pass or fail (development questions)

No previous release yet. We cannot say what improved or got worse.

No question-level results were recorded for this run.

Can we trust the AI judges?

Groundedness
gemini-2.5-flash · judges-v2

Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.

Correctness
gemini-2.5-flash · judges-v2

Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.

Conflict disclosure
gemini-2.5-flash · judges-v2

Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.

Answer relevance
gemini-2.5-flash · judges-v2

Uncalibrated — awaiting owner labels. Uncalibrated: needs ~100 owner labels (tune on 60, measure on 40); target Cohen's kappa >= 0.7 before this judge counts in release gates.