Methodology

You need evidence you can check, not just a confident answer. Here is why each safeguard exists, how it works, and where it falls short.

Evidence before an answer

  1. 01
    Search more than one way: rewrite the question into up to 3 search queries
  2. 02
    Find relevant evidence: vector search (pgvector) plus keyword search, merged
  3. 03
    Put stronger evidence first: small bonus for higher tiers and peer review, penalty for superseded sources; up to 6 sources, max 2 passages each, 8 passages total
  4. 04
    Keep claims checkable: answer with gpt-4.1-mini under the answer policy
  5. 05
    Catch weak answers: rule-based checks on every answer; AI judges (Gemini 2.5 Flash, a different provider) on a sample
  6. 06
    Learn from mistakes: log for review

Release v1 facts

Answer model
gpt-4.1-mini (OpenAI)
Judge model
gemini-2.5-flash (Google — a different provider from the answer model)
Embeddings
text-embedding-3-small, 768 dimensions
Retrieval
Up to 3 search queries; vector + keyword search merged; 8 passages from up to 6 sources (max 2 each)
Abstain threshold
Decline to answer below 0.36 retrieval similarity
Judge calibration
UNCALIBRATED — needs ~100 owner labels; target Cohen's κ ≥ 0.7
Tracing
LangSmith for sampled traces (20%) and benchmark experiments
Score storage
All scores live in the EvalMaestro database

Ten safeguards for answers you can check

  1. 1So you can check a claim, every factual sentence carries at least one citation to an approved source.
  2. 2A convincing-looking reference is not evidence. Citations must point to real corpus sources; fabricated citations block the answer.
  3. 3To avoid overstating the evidence, wording follows tier: T1 “Research shows”, T2 “Official guidance recommends”, T3 “Practitioners argue”.
  4. 4Disagreement matters to your decision. When sources of different tiers conflict, we show both sides and their tiers.
  5. 5Summaries should not invent evidence. They may combine cited claims, but must not add new facts.
  6. 6To keep the original source one click away without reproducing it, we show metadata, quotes under 25 words and links only, never full source text.
  7. 7A guess is not a useful answer. We decline to answer when no retrieved passage reaches 0.36 retrieval similarity.
  8. 8To keep each claim easy to check, a sentence carries at most 2 citations.
  9. 9So the strength of a claim is honest, the wording (“Research shows” / “Official guidance” / “Practitioners”) must match the tier of the source it cites.
  10. 10So every claim can be traced, we state only what a cited source says — no background knowledge from the model.

Checks that catch weak answers

MetricMethodPrompt versionReference
Fabricated citations Rule-based checkv1EvalMaestro answer policy, rule 2
Citation validity Rule-based checkv1Gao et al. 2023 (ALCE)
Citation coverage Rule-based checkv1Gao et al. 2023 (ALCE)
Quote length Rule-based checkv1EvalMaestro answer policy, rule 6
Tier compliance Rule-based checkv1EvalMaestro answer policy, rule 3
Retrieval confidence Rule-based checkv1Saad-Falcon et al. 2023 (ARES)
Groundedness AI judge (gemini-2.5-flash)v1Es et al. 2023 (RAGAS faithfulness)
Tier compliance (judge) AI judge (gemini-2.5-flash)v1EvalMaestro answer policy, rule 3
Answer relevance AI judge (gemini-2.5-flash)v1Liu et al. 2023 (G-Eval)

Checks that protect a release

  • To block invented sources: zero fabricated citations.
  • To keep sources checkable: citation validity 100%.
  • To avoid guessing: abstention (declining when evidence is missing) at least 90%.
  • To avoid unsupported claims: groundedness — at least 90% of claims supported (judge not yet calibrated).
  • To prevent backward steps: correctness must not significantly drop vs the previous release (paired bootstrap 95% interval entirely below zero blocks).
  • Latency and cost are tracked but do not block a release.

Research behind the safeguards

Every source in the corpus (33)

Where we fall short

  • AI judges can share the answer model's biases: answer order, wordiness and preference for their own style. A judge score is not proof.
  • The benchmark is small and confidence intervals are wide. These scores do not establish how well we handle every question.
  • The owner approves every source before it enters the corpus. Evidence tiers are editorial judgments and can be wrong — please report mistakes.
  • We can only answer from our approved sources. Missing evidence here does not mean evidence does not exist elsewhere.