Methodology
You need evidence you can check, not just a confident answer. Here is why each safeguard exists, how it works, and where it falls short.
Evidence before an answer
- 01Search more than one way: rewrite the question into up to 3 search queries
- 02Find relevant evidence: vector search (pgvector) plus keyword search, merged
- 03Put stronger evidence first: small bonus for higher tiers and peer review, penalty for superseded sources; up to 6 sources, max 2 passages each, 8 passages total
- 04Keep claims checkable: answer with gpt-4.1-mini under the answer policy
- 05Catch weak answers: rule-based checks on every answer; AI judges (Gemini 2.5 Flash, a different provider) on a sample
- 06Learn from mistakes: log for review
Release v1 facts
- Answer model
- gpt-4.1-mini (OpenAI)
- Judge model
- gemini-2.5-flash (Google — a different provider from the answer model)
- Embeddings
- text-embedding-3-small, 768 dimensions
- Retrieval
- Up to 3 search queries; vector + keyword search merged; 8 passages from up to 6 sources (max 2 each)
- Abstain threshold
- Decline to answer below 0.36 retrieval similarity
- Judge calibration
- UNCALIBRATED — needs ~100 owner labels; target Cohen's κ ≥ 0.7
- Tracing
- LangSmith for sampled traces (20%) and benchmark experiments
- Score storage
- All scores live in the EvalMaestro database
Ten safeguards for answers you can check
- 1So you can check a claim, every factual sentence carries at least one citation to an approved source.
- 2A convincing-looking reference is not evidence. Citations must point to real corpus sources; fabricated citations block the answer.
- 3To avoid overstating the evidence, wording follows tier: T1 “Research shows”, T2 “Official guidance recommends”, T3 “Practitioners argue”.
- 4Disagreement matters to your decision. When sources of different tiers conflict, we show both sides and their tiers.
- 5Summaries should not invent evidence. They may combine cited claims, but must not add new facts.
- 6To keep the original source one click away without reproducing it, we show metadata, quotes under 25 words and links only, never full source text.
- 7A guess is not a useful answer. We decline to answer when no retrieved passage reaches 0.36 retrieval similarity.
- 8To keep each claim easy to check, a sentence carries at most 2 citations.
- 9So the strength of a claim is honest, the wording (“Research shows” / “Official guidance” / “Practitioners”) must match the tier of the source it cites.
- 10So every claim can be traced, we state only what a cited source says — no background knowledge from the model.
Checks that catch weak answers
| Metric | Method | Prompt version | Reference |
|---|---|---|---|
| Fabricated citations | Rule-based check | v1 | EvalMaestro answer policy, rule 2 |
| Citation validity | Rule-based check | v1 | Gao et al. 2023 (ALCE) |
| Citation coverage | Rule-based check | v1 | Gao et al. 2023 (ALCE) |
| Quote length | Rule-based check | v1 | EvalMaestro answer policy, rule 6 |
| Tier compliance | Rule-based check | v1 | EvalMaestro answer policy, rule 3 |
| Retrieval confidence | Rule-based check | v1 | Saad-Falcon et al. 2023 (ARES) |
| Groundedness | AI judge (gemini-2.5-flash) | v1 | Es et al. 2023 (RAGAS faithfulness) |
| Tier compliance (judge) | AI judge (gemini-2.5-flash) | v1 | EvalMaestro answer policy, rule 3 |
| Answer relevance | AI judge (gemini-2.5-flash) | v1 | Liu et al. 2023 (G-Eval) |
Checks that protect a release
- To block invented sources: zero fabricated citations.
- To keep sources checkable: citation validity 100%.
- To avoid guessing: abstention (declining when evidence is missing) at least 90%.
- To avoid unsupported claims: groundedness — at least 90% of claims supported (judge not yet calibrated).
- To prevent backward steps: correctness must not significantly drop vs the previous release (paired bootstrap 95% interval entirely below zero blocks).
- Latency and cost are tracked but do not block a release.
Research behind the safeguards
Every source in the corpus (33)
- T1JudgeBench: A Benchmark for Evaluating LLM-based Judges (2024)
- T1LLM Evaluators Recognize and Favor Their Own Generations (2024)
- T1Self-Preference Bias in LLM-as-a-Judge (2024)
- T1Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge (2024)
- T1A Survey on LLM-as-a-Judge (2024)
- T1Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences (2024)
- T1Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges (2024)
- T1Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations (2024)
- T1tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains (2024)
- T1Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (2023)
- T1Large Language Models are not Fair Evaluators (2023)
- T1RAGAS: Automated Evaluation of Retrieval Augmented Generation (2023)
- T1G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment (2023)
- T1FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation (2023)
- T1Enabling Large Language Models to Generate Text with Citations (ALCE) (2023)
- T1Prometheus: Inducing Fine-grained Evaluation Capability in Language Models (2023)
- T1Evaluating Verifiability in Generative Search Engines (2023)
- T1Benchmarking Large Language Models in Retrieval-Augmented Generation (2023)
- T1SWE-bench: Can Language Models Resolve Real-World GitHub Issues? (2023)
- T1AgentBench: Evaluating LLMs as Agents (2023)
- T1ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems (2023)
- T1Holistic Evaluation of Language Models (HELM) (2022)
- T2Evaluation best practices (2025)
- T2Evaluation concepts (LangSmith) (2025)
- T2Create strong empirical evaluations (2025)
- T2Gen AI evaluation service overview (Vertex AI) (2025)
- T2Inspect: a framework for large language model evaluations (2025)
- T2Artificial Intelligence Risk Management Framework: Generative AI Profile (NIST AI 600-1) (2024)
- T3A Field Guide to Rapidly Improving AI Products (2025)
- T3Creating a LLM-as-a-Judge That Drives Business Results (2024)
- T3Task-Specific LLM Evals that Do & Don't Work (2024)
- T3Evaluating the Effectiveness of LLM-Evaluators (aka LLM-as-Judge) (2024)
- T3Your AI Product Needs Evals (2024)
Where we fall short
- AI judges can share the answer model's biases: answer order, wordiness and preference for their own style. A judge score is not proof.
- The benchmark is small and confidence intervals are wide. These scores do not establish how well we handle every question.
- The owner approves every source before it enters the corpus. Evidence tiers are editorial judgments and can be wrong — please report mistakes.
- We can only answer from our approved sources. Missing evidence here does not mean evidence does not exist elsewhere.