vault / agents / evals.md

Evaluation-driven development

Agents & LLMs 54 words 4 outgoing 4 backlinks

Compare relevance, retrieval recall, groundedness and latency against a baseline. Inspect examples as well as scores. In my current workflow, evaluation produces a report for engineers rather than an automatic deployment gate. That is a description of this system, not a rule that every team should avoid gates.

Read the full example: related post.