Evaluating an agent honestly
Relevance, retrieval recall, groundedness and latency each catch a different failure. Here is how I use them as evidence, and why the report is not a deploy gate.
A single score cannot tell you whether an agent is useful. I work on a production assistant that answers questions by searching a body of documents and writing a reply from what it finds, a pattern usually called retrieval-augmented generation, or RAG. It’s agentic, meaning a model decides which steps to take, so there are more places for it to go wrong than in a fixed pipeline. This post is about how I evaluate it, and about one decision that surprises people: the evaluation produces a report for engineers, not an automatic pass or fail on deployment.
Why “it looks better” stops working
In the first weeks of building an agent you can iterate by feel. You try a handful of questions, read the answers, change a prompt, try again. That’s a perfectly good way to start.
It breaks down once the system has some reach. The assistant I work on serves several markets, takes questions in any language and mixes global and region-specific knowledge. A change that improves one language can quietly degrade another. Nobody can read enough answers by hand to notice. And because the output is fluent, a worse answer doesn’t look worse.
So I treat evaluation as engineering evidence. Compare a change with a baseline, look at the failures one by one, and only then decide what to change next. This is what people mean by evaluation-driven development: the eval isn’t a check at the end, it’s what steers the design.
Four metrics, four different failures
I measure four things: answer relevance, retrieval recall, groundedness and latency. I want to be specific about why those four, because each one exists to catch something the others miss.
- Answer relevance: does the answer address what the user asked?
- Retrieval recall: did the search find the documents that contain the answer? This one needs a reference, the source a good answer should come from, which is why my golden examples record one.
- Groundedness: is every claim in the answer supported by what was retrieved? An answer can be relevant and still invent a detail.
- Latency: how long did the user wait? A correct answer that arrives too late is, for most purposes, not a good one.
Two rows in that matrix deserve a comment. “Faithful to sources, but off-topic” shows why groundedness alone is not enough: an answer can quote its sources perfectly and still not answer the question. And the last row shows the limit of all of them. If the source document is out of date, the assistant can be relevant, well retrieved and faithfully grounded in something that is no longer true. No automated metric flags that. A person reading the failures does, or a content review does.
This is why I don’t collapse the four into one number. An average lets a good latency score hide a poor recall score. Reading them side by side tells you which part of the system to look at.
Two kinds of evaluation, two jobs
There are two evaluations in the loop, and they answer different questions.
Offline evaluation runs on demand against a fixed set of questions. In my case it runs whenever there’s an architecture change or a significant code change, and it’s also what I use to check that a new market or language is ready. Because the test set is fixed, two runs are comparable. It answers: is this change better than what we have?
Online evaluation runs against live traffic. Its job is to detect regressions and bugs in production, the things nobody thought to put in the test set. It answers: what did we miss?
The loop closes when a failure found online becomes a new offline case. Over time the test set stops being what you thought users would ask and becomes a record of what actually broke. Neither kind is a substitute for the other, and neither is a substitute for understanding the task: someone has to read the failures.
Readiness reviews for a new market combine the offline results with something else: the history of issues that came through support tickets. The eval tells you how the system does on the test questions, and the ticket history tells you what people have actually run into. In practice, the findings from that combination have mostly led to improvements in the source content, not in the code.
Why the report is not a gate
The default instinct, and it’s a reasonable one, is to wire the eval into CI and block a release if the score falls below a threshold. It’s what we do with unit tests. I chose not to, and the pipeline I built produces a comprehensive report that compares the current run with a baseline or with agreed criteria, and leaves the decision to an engineer.
My reasoning comes down to three points.
The scores are noisy. Metrics like relevance and groundedness are commonly scored by a language model acting as a judge, and the agent itself is not deterministic. Run the same system twice and you’ll get slightly different numbers. A hard threshold turns that noise into a coin flip:
A single threshold hides the trade-offs. A change that lifts recall while costing a little relevance on a narrow slice might be exactly the change you want. A gate sees one number and says no. A report shows the per-metric differences, the slices that moved, and the actual examples that got worse, so a person can judge whether the trade is worth it.
A red cross ends the conversation. When a gate fails, the natural response is to make it pass: tune until the number moves, or lower the bar. When a report arrives, the natural response is to read it. I’d like the team to argue about specific failing examples, not about the threshold.
Where a gate does belong
I’m not against gates in general. A hard gate suits checks that are cheap, deterministic and clearly binary. A latency budget is a good example: either the response time is inside the budget or it isn’t, and there’s little to interpret. The same goes for a safety filter that must never be bypassed, or a schema check on tool output. Those I’d happily block on.
The metrics I’d keep as reports are the ones that need reading: relevance, and the finer judgements about groundedness. The rule of thumb I use is that if a person would want to look at the case before deciding, it shouldn’t be an automatic block.
What changes on a team
Making the evaluation a report changes how people use it. It becomes something you read before a review, alongside the diff, and not a gate that you try to get past. Reports that people actually want to read need to be legible: baseline next to current, differences by metric and slice, and the worst examples one click away. If nobody reads the report, you’ve built a gate with extra steps, so it’s worth spending effort on the reading experience.
The practical lesson is to make evaluation part of the development loop, keep the reports readable by the people making the next decision, and keep a human in the seat where the decision is a judgement call.
References
- Lewis, P. et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS, 2020. https://arxiv.org/abs/2005.11401
- Xia, B. et al. Evaluation-Driven Development and Operations of LLM Agents: A Process Model and Reference Architecture. arXiv, 2024. https://arxiv.org/abs/2411.13768
- Es, S. et al. Ragas: Automated Evaluation of Retrieval Augmented Generation. arXiv, 2023. https://arxiv.org/abs/2309.15217 — relevance, context and faithfulness metrics.
- Saad-Falcon, J. et al. ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems. NAACL, 2024. https://arxiv.org/abs/2311.09476
- Barnett, S. et al. Seven Failure Points When Engineering a Retrieval Augmented Generation System. arXiv, 2024. https://arxiv.org/abs/2401.05856
- Zheng, L. et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS Datasets and Benchmarks, 2023. https://arxiv.org/abs/2306.05685
- Miller, E. Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations. arXiv, 2024. https://arxiv.org/abs/2411.00640 — quantifying run-to-run noise.
- Jones, C., Wilkes, J. and Murphy, N. Service Level Objectives. In Site Reliability Engineering, Google, O’Reilly, 2016. https://sre.google/sre-book/service-level-objectives/ — latency budgets as SLOs.