Skip to content
F Fadhil Mochammad ML engineer · Stockholm
About Posts Projects Notes
EXPLORER 24 notes
Agents & LLMs9
  • Agents
  • Evaluation-driven development
  • Golden datasets
  • Groundedness
  • Human review
  • Online evaluation
  • RAG
  • Retrieval recall
  • Feedback triage
Experimentation5
  • A/B testing
  • Multi-armed bandits
  • Metric semantic layer
  • Self-service tooling
  • Thompson sampling
ML platform4
  • Kubernetes
  • SDK design
  • Model serving
  • Distributed traces
System design3
  • Go worker pools
  • P95 latency
  • Observability
Photography1
  • Exposure triangle
Music1
  • FM synthesis
Meta1
  • Start here

No notes match that search.

vault / agents / online-eval.md

Online evaluation

Agents & LLMs 51 words 3 outgoing 1 backlinks

Offline examples help compare changes under controlled conditions. Production feedback shows questions and failures the reference set missed. The two complement each other. Online signals need investigation: a thumbs-down tells us that something went wrong, not whether retrieval, content, configuration or the answer caused it.

Read the full example: related post.

LINKS IN THIS NOTE
[[ Evaluation-driven development ]] Compare relevance, retrieval recall, groundedness and latency against a baseline. Inspect examples as well as scores. In my current workflow, evaluation produces a report for engineers rather than an automatic deployment gate. That is a description of this system, not... [[ Observability ]] For feedback investigation, record enough of the request to reconstruct what happened: the question, retrieval, response and relevant configuration. A trace helps connect those steps. Dashboards can show a pattern, but a concrete replay is often what tells us which... [[ Feedback triage ]] A thumbs-down starts an investigation, not a diagnosis. The workflow I built replays the question, collects evidence, distinguishes a content gap from a configuration issue, and routes the case to an owner. Routing is ordinary code; the model proposes the...
GRAPHdrag · scroll · click
BACKLINKS Evaluation-driven development Compare relevance, retrieval recall, groundedness and latency against a baseline. Inspect examples as well as scores. In my current workflow, evaluation produces a report for engineers rather than an automatic deployment gate. That is a description of this system, not...
OUTGOING
Evaluation-driven development Observability Feedback triage
F
© 2026 Fadhil Mochammad · Stockholm ● agents ● experimentation ● data ● ml-platform ● systems ● creative