Skip to content
F Fadhil Mochammad ML engineer · Stockholm
About Posts Projects Notes
EXPLORER 24 notes
Agents & LLMs9
  • Agents
  • Evaluation-driven development
  • Golden datasets
  • Groundedness
  • Human review
  • Online evaluation
  • RAG
  • Retrieval recall
  • Feedback triage
Experimentation5
  • A/B testing
  • Multi-armed bandits
  • Metric semantic layer
  • Self-service tooling
  • Thompson sampling
ML platform4
  • Kubernetes
  • SDK design
  • Model serving
  • Distributed traces
System design3
  • Go worker pools
  • P95 latency
  • Observability
Photography1
  • Exposure triangle
Music1
  • FM synthesis
Meta1
  • Start here

No notes match that search.

vault / agents / agents.md

Agents

Agents & LLMs 52 words 3 outgoing 1 backlinks

In the assistant described in my posts, a model chooses steps while retrieval supplies the evidence. That creates several ways to fail: a poor tool choice, missing documents, or an answer that goes beyond them. A fluent reply alone does not show that the loop worked.

Read the full example: related post.

LINKS IN THIS NOTE
[[ RAG ]] Retrieval-augmented generation searches a document collection, then writes an answer using the retrieved context. Keep retrieval and answer quality separate: finding the right document does not guarantee the model uses it, and a plausible answer does not prove the document... [[ Evaluation-driven development ]] Compare relevance, retrieval recall, groundedness and latency against a baseline. Inspect examples as well as scores. In my current workflow, evaluation produces a report for engineers rather than an automatic deployment gate. That is a description of this system, not... [[ Feedback triage ]] A thumbs-down starts an investigation, not a diagnosis. The workflow I built replays the question, collects evidence, distinguishes a content gap from a configuration issue, and routes the case to an owner. Routing is ordinary code; the model proposes the...
GRAPHdrag · scroll · click
BACKLINKS Start here These notes are short companions to the posts, not a separate set of claims. Start with a concept, follow its related notes, then open the article for the example and the limits. For code and worked projects, see [Projects]({{ '/projects/'...
OUTGOING
RAG Evaluation-driven development Feedback triage
F
© 2026 Fadhil Mochammad · Stockholm ● agents ● experimentation ● data ● ml-platform ● systems ● creative