Skip to content
F Fadhil Mochammad ML engineer · Stockholm
About Posts Projects Notes
EXPLORER 24 notes
Agents & LLMs9
  • Agents
  • Evaluation-driven development
  • Golden datasets
  • Groundedness
  • Human review
  • Online evaluation
  • RAG
  • Retrieval recall
  • Feedback triage
Experimentation5
  • A/B testing
  • Multi-armed bandits
  • Metric semantic layer
  • Self-service tooling
  • Thompson sampling
ML platform4
  • Kubernetes
  • SDK design
  • Model serving
  • Distributed traces
System design3
  • Go worker pools
  • P95 latency
  • Observability
Photography1
  • Exposure triangle
Music1
  • FM synthesis
Meta1
  • Start here

No notes match that search.

vault / agents / rag.md

RAG

Agents & LLMs 48 words 3 outgoing 3 backlinks

Retrieval-augmented generation searches a document collection, then writes an answer using the retrieved context. Keep retrieval and answer quality separate: finding the right document does not guarantee the model uses it, and a plausible answer does not prove the document was found.

Read the full example: related post.

LINKS IN THIS NOTE
[[ Retrieval recall ]] Ask whether the expected sources reached the retrieved context. When the answer is wrong, this helps distinguish missing evidence from poor use of available evidence. The expected source set must itself be checked; a score against an incomplete reference can... [[ Groundedness ]] Check whether the answer’s claims are supported by the retrieved sources. This differs from usefulness: an answer can stay within its evidence and still fail to answer the question. Read groundedness alongside relevance and retrieval recall, rather than reducing all... [[ Evaluation-driven development ]] Compare relevance, retrieval recall, groundedness and latency against a baseline. Inspect examples as well as scores. In my current workflow, evaluation produces a report for engineers rather than an automatic deployment gate. That is a description of this system, not...
GRAPHdrag · scroll · click
BACKLINKS Agents In the assistant described in my posts, a model chooses steps while retrieval supplies the evidence. That creates several ways to fail: a poor tool choice, missing documents, or an answer that goes beyond them. A fluent reply alone does... Groundedness Check whether the answer’s claims are supported by the retrieved sources. This differs from usefulness: an answer can stay within its evidence and still fail to answer the question. Read groundedness alongside relevance and retrieval recall, rather than reducing all... Retrieval recall Ask whether the expected sources reached the retrieved context. When the answer is wrong, this helps distinguish missing evidence from poor use of available evidence. The expected source set must itself be checked; a score against an incomplete reference can...
OUTGOING
Retrieval recall Groundedness Evaluation-driven development
F
© 2026 Fadhil Mochammad · Stockholm ● agents ● experimentation ● data ● ml-platform ● systems ● creative