← All posts

Evaluating an agent honestly

Relevance, retrieval recall, groundedness and latency each catch a different failure. Here is how I use them as evidence, and why the report is not a deploy gate.

A single score cannot tell you whether an agent is useful. I work on a production assistant that answers questions by searching a body of documents and writing a reply from what it finds, a pattern usually called retrieval-augmented generation, or RAG. It’s agentic, meaning a model decides which steps to take, so there are more places for it to go wrong than in a fixed pipeline. This post is about how I evaluate it, and about one decision that surprises people: the evaluation produces a report for engineers, not an automatic pass or fail on deployment.

Why “it looks better” stops working

In the first weeks of building an agent you can iterate by feel. You try a handful of questions, read the answers, change a prompt, try again. That’s a perfectly good way to start.

It breaks down once the system has some reach. The assistant I work on serves several markets, takes questions in any language and mixes global and region-specific knowledge. A change that improves one language can quietly degrade another. Nobody can read enough answers by hand to notice. And because the output is fluent, a worse answer doesn’t look worse.

So I treat evaluation as engineering evidence. Compare a change with a baseline, look at the failures one by one, and only then decide what to change next. This is what people mean by evaluation-driven development: the eval isn’t a check at the end, it’s what steers the design.

Four metrics, four different failures

I measure four things: answer relevance, retrieval recall, groundedness and latency. I want to be specific about why those four, because each one exists to catch something the others miss.

  • Answer relevance: does the answer address what the user asked?
  • Retrieval recall: did the search find the documents that contain the answer? This one needs a reference, the source a good answer should come from, which is why my golden examples record one.
  • Groundedness: is every claim in the answer supported by what was retrieved? An answer can be relevant and still invent a detail.
  • Latency: how long did the user wait? A correct answer that arrives too late is, for most purposes, not a good one.
Which metric catches which failure A matrix of six failure cases against four metrics. Ignoring the question and off-topic answers are caught by relevance. A missing document is caught by retrieval recall. Unsupported claims are caught by groundedness. A slow answer is caught by latency. A fluent, grounded answer built on a stale source is caught by none of the four. FAILURE CASE RELEVANCE RECALL GROUNDED LATENCY Answer ignores the question Answer ignores the question: caught by relevance Right document never retrieved Right document never retrieved: caught by retrieval recall Answer claims more than sources say Answer claims more than sources say: caught by groundedness Faithful to sources, but off-topic Faithful to sources, but off-topic: caught by relevance Good answer, arrives far too late Good answer, arrives far too late: caught by latency Fluent and grounded, but source is stale Fluent and grounded, but source is stale: none of the four metrics flags itnone of the four
Each metric covers a different failure, and each can look healthy while another one is failing. The last row is the reminder that a passing dashboard is not the same as a good answer.

Two rows in that matrix deserve a comment. “Faithful to sources, but off-topic” shows why groundedness alone is not enough: an answer can quote its sources perfectly and still not answer the question. And the last row shows the limit of all of them. If the source document is out of date, the assistant can be relevant, well retrieved and faithfully grounded in something that is no longer true. No automated metric flags that. A person reading the failures does, or a content review does.

This is why I don’t collapse the four into one number. An average lets a good latency score hide a poor recall score. Reading them side by side tells you which part of the system to look at.

Two kinds of evaluation, two jobs

There are two evaluations in the loop, and they answer different questions.

Offline evaluation runs on demand against a fixed set of questions. In my case it runs whenever there’s an architecture change or a significant code change, and it’s also what I use to check that a new market or language is ready. Because the test set is fixed, two runs are comparable. It answers: is this change better than what we have?

Online evaluation runs against live traffic. Its job is to detect regressions and bugs in production, the things nobody thought to put in the test set. It answers: what did we miss?

Offline evaluation and online monitoring in one loop A change goes through an on-demand offline evaluation run, which produces a report against a baseline. An engineer reads the report and decides whether to release. After release, online evaluation watches live traffic. A regression or new failure found there is turned into a new test case, which feeds the next change. BEFORE RELEASE · OFFLINE AFTER RELEASE · ONLINE Changecode or design Offline runon demand Reportvs baseline Engineerreads, decides Releaseship it Online evallive traffic Regressionor new failure Test setadd the case
The two kinds of evaluation do different jobs. Offline runs answer "is this change better?" before release. Online evaluation answers "what did we miss?" after it, and feeds the offline set.

The loop closes when a failure found online becomes a new offline case. Over time the test set stops being what you thought users would ask and becomes a record of what actually broke. Neither kind is a substitute for the other, and neither is a substitute for understanding the task: someone has to read the failures.

Readiness reviews for a new market combine the offline results with something else: the history of issues that came through support tickets. The eval tells you how the system does on the test questions, and the ticket history tells you what people have actually run into. In practice, the findings from that combination have mostly led to improvements in the source content, not in the code.

Why the report is not a gate

The default instinct, and it’s a reasonable one, is to wire the eval into CI and block a release if the score falls below a threshold. It’s what we do with unit tests. I chose not to, and the pipeline I built produces a comprehensive report that compares the current run with a baseline or with agreed criteria, and leaves the decision to an engineer.

My reasoning comes down to three points.

The scores are noisy. Metrics like relevance and groundedness are commonly scored by a language model acting as a judge, and the agent itself is not deterministic. Run the same system twice and you’ll get slightly different numbers. A hard threshold turns that noise into a coin flip:

Twelve runs of an unchanged system against a fixed threshold Simulated. Twelve repeated evaluation runs of the same system score between 0.78 and 0.82. A pass mark at 0.80 sits in the middle of that spread, so 5 of the twelve runs fail with no change to the system. 0.74 0.78 0.82 0.86 pass mark 0.80 Run 1: score 0.795, fails the 0.80 pass mark 1 Run 2: score 0.810, passes the 0.80 pass mark 2 Run 3: score 0.795, fails the 0.80 pass mark 3 Run 4: score 0.794, fails the 0.80 pass mark 4 Run 5: score 0.781, fails the 0.80 pass mark 5 Run 6: score 0.796, fails the 0.80 pass mark 6 Run 7: score 0.822, passes the 0.80 pass mark 7 Run 8: score 0.808, passes the 0.80 pass mark 8 Run 9: score 0.821, passes the 0.80 pass mark 9 Run 10: score 0.805, passes the 0.80 pass mark 10 Run 11: score 0.808, passes the 0.80 pass mark 11 Run 12: score 0.804, passes the 0.80 pass mark 12 run number, same code and same test set each time answer score BLUE PASSES · ORANGE FAILS
Simulated data. Nothing changed between runs, yet 5 of 12 fail. A hard threshold turns run-to-run noise into a red cross.

A single threshold hides the trade-offs. A change that lifts recall while costing a little relevance on a narrow slice might be exactly the change you want. A gate sees one number and says no. A report shows the per-metric differences, the slices that moved, and the actual examples that got worse, so a person can judge whether the trade is worth it.

A red cross ends the conversation. When a gate fails, the natural response is to make it pass: tune until the number moves, or lower the bar. When a report arrives, the natural response is to read it. I’d like the team to argue about specific failing examples, not about the threshold.

Where a gate does belong

I’m not against gates in general. A hard gate suits checks that are cheap, deterministic and clearly binary. A latency budget is a good example: either the response time is inside the budget or it isn’t, and there’s little to interpret. The same goes for a safety filter that must never be bypassed, or a schema check on tool output. Those I’d happily block on.

The metrics I’d keep as reports are the ones that need reading: relevance, and the finer judgements about groundedness. The rule of thumb I use is that if a person would want to look at the case before deciding, it shouldn’t be an automatic block.

What changes on a team

Making the evaluation a report changes how people use it. It becomes something you read before a review, alongside the diff, and not a gate that you try to get past. Reports that people actually want to read need to be legible: baseline next to current, differences by metric and slice, and the worst examples one click away. If nobody reads the report, you’ve built a gate with extra steps, so it’s worth spending effort on the reading experience.

The practical lesson is to make evaluation part of the development loop, keep the reports readable by the people making the next decision, and keep a human in the seat where the decision is a judgement call.

References

  1. Lewis, P. et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS, 2020. https://arxiv.org/abs/2005.11401
  2. Xia, B. et al. Evaluation-Driven Development and Operations of LLM Agents: A Process Model and Reference Architecture. arXiv, 2024. https://arxiv.org/abs/2411.13768
  3. Es, S. et al. Ragas: Automated Evaluation of Retrieval Augmented Generation. arXiv, 2023. https://arxiv.org/abs/2309.15217 — relevance, context and faithfulness metrics.
  4. Saad-Falcon, J. et al. ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems. NAACL, 2024. https://arxiv.org/abs/2311.09476
  5. Barnett, S. et al. Seven Failure Points When Engineering a Retrieval Augmented Generation System. arXiv, 2024. https://arxiv.org/abs/2401.05856
  6. Zheng, L. et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS Datasets and Benchmarks, 2023. https://arxiv.org/abs/2306.05685
  7. Miller, E. Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations. arXiv, 2024. https://arxiv.org/abs/2411.00640 — quantifying run-to-run noise.
  8. Jones, C., Wilkes, J. and Murphy, N. Service Level Objectives. In Site Reliability Engineering, Google, O’Reilly, 2016. https://sre.google/sre-book/service-level-objectives/ — latency budgets as SLOs.