An LLM-as-Judge Won't Save The Product—Fixing Your Process Will
T3Eugene Yan · 2025
This piece argues that while LLM-as-judge has its place, robust evaluation processes are more critical for product success.
It advocates for scientific methods, eval-driven development, and continuous monitoring of AI outputs.
Why review it: blog · relevance 0.90 · The item directly addresses the utility of LLM-as-judge and emphasizes process-driven evaluation, which is central to AI quality.
llm-as-judgepracticetooling
Evaluating Long-Context Question & Answer Systems
T3Eugene Yan · 2025
This item covers evaluation metrics, dataset construction, and methodologies for long-context Q&A systems.
It also reviews several existing benchmarks in this domain.
Why review it: blog · relevance 0.95 · This item directly addresses evaluation metrics, dataset creation, methodology, and benchmarks for Q&A systems, which are highly relevant to LLM evaluation.
benchmarkspracticerag-eval
Product Evals in Three Simple Steps
T3Eugene Yan · 2025
This item outlines a three-step process for product evaluations: data labeling, aligning LLM-evaluators, and running an eval harness.
It focuses on practical implementation of evaluation methodologies.
Why review it: blog · relevance 0.90 · This item directly discusses practical steps for evaluating LLMs, aligning with the 'eval practice' and 'tooling' aspects of the corpus.
practicetoolingllm-as-judge
Patterns for Building Cybersecurity Evals
T3Eugene Yan · 2026
This item outlines key components for building effective evaluation systems, including sandboxed targets, variable inputs, tools, and a grader.
It provides a structural framework for designing evaluations, applicable across various domains.
Why review it: blog · relevance 0.85 · This item directly discusses patterns for building evaluations, which is highly relevant to LLM evaluation practices, even if the specific domain is cybersecurity.
practicetoolingbenchmarks
Selecting The Right AI Evals Tool
T3Hamel Husain · 2025
This article discusses the process of selecting AI evaluation tools, emphasizing that the 'best' tool depends on team specifics rather than a universal solution.
It presents a comparative assessment of Langsmith, Braintrust, and Arize Phoenix by data scientists tackling the same evaluation challenge.
Why review it: blog · relevance 0.95 · This item directly addresses the practical aspects of selecting and using tools for evaluating LLMs, which is central to the corpus's focus.
toolingpractice
Evals Skills for Coding Agents
T3Hamel Husain · 2026
This post introduces 'evals skills,' a set of practical skills designed to improve AI product evaluations, particularly for coding agents.
It provides guidance and tools to avoid common evaluation mistakes and offers specific skills like 'eval-audit' and 'error-discovery'.
Why review it: blog · relevance 0.90 · This item directly addresses practical evaluation skills and tools for AI agents, aligning perfectly with the corpus's focus on LLM evaluation and AI quality.
agent-evalpracticetooling
The Revenge of the Data Scientist
T3Hamel Husain · 2026
This article discusses the evolving role of data scientists in the age of LLMs, arguing that their core value remains in experimental design, debugging, and metric development.
It suggests that integrating AI via foundation model APIs doesn't eliminate the need for these skills, particularly for evaluating how AI generalizes.
Why review it: blog · relevance 0.80 · While broadly about the role of data scientists, the text emphasizes the enduring importance of experimental design, debugging stochastic systems, and designing good metrics in the context of LLMs, which are core to evaluation.
practicestatistics
“It’s Hard to Eval” Is a Product Smell
T3Hamel Husain · 2026
This post argues that difficulty in evaluating an AI product is a 'product smell' and suggests designing products for ease of verification.
It provides examples and design principles, particularly for AI data agents, to make outputs more verifiable.
Why review it: blog · relevance 0.90 · This item directly addresses the practical challenges and design principles for evaluating AI systems, particularly agents, which is central to the corpus's focus.
agent-evalpracticetooling
Do Automated Evals Work?
T3Hamel Husain · 2026
This item investigates the effectiveness of automated evaluation systems by comparing their outputs against human annotations.
The findings will shed light on the reliability and validity of current automated evaluation practices.
Why review it: blog · relevance 0.90 · The title and text directly address the efficacy of automated evaluation methods, a core topic for LLM evaluation and AI quality.
llm-as-judgebenchmarkspracticetooling
AI Evals: Everything You Need to Know
T3Hamel Husain · 2026
This FAQ addresses common questions and challenges faced by engineers and PMs regarding AI evaluations.
It covers fundamental concepts, practical implementation, and troubleshooting eval scores.
Why review it: blog · relevance 0.95 · This document is a comprehensive FAQ specifically about AI evaluations, directly addressing common challenges and questions in the field.
practicetoolingbenchmarks
Claude’s new auto eval tool
T3Hamel Husain · 2026
This article reviews Anthropic's new auto-evaluation tool for Claude, focusing on its practical application and identifying areas for improvement in its workflow.
The authors livestreamed their experience using the tool to evaluate an apartment leasing assistant, highlighting issues like premature eval creation and insufficient context for judgment validation.
Why review it: blog · relevance 0.90 · This item directly discusses a new evaluation tool for LLMs, its features, and a practical review of its usage.
toolingpracticeagent-eval