Posts

21 posts
2026
· 12 min read What I learned tuning retrieval for a multilingual RAG chatbot Multilingual questions against a multilingual knowledge base: English questions were fine all along, while Swedish, Polish and Indonesian ones quietly lost their documents until each question was translated into the knowledge base's languages. · 11 min read Deconstructing PlanOut: a coin toss you can repeat Follow one shopper from a button experiment to the actual SHA-1 bytes, then look at salts, namespaces and the gap between assignment and exposure. · 7 min read Evaluating an agent honestly Relevance, retrieval recall, groundedness and latency each catch a different failure. Here is how I use them as evidence, and why the report is not a deploy gate. · 7 min read An agent that triages its own feedback Replay the flagged question, check the evidence, name the likely cause, route it to an owner, and send anything doubtful to a person. · 6 min read Golden datasets without the gold rush Simulated personas propose evaluation queries anchored to real documents, and people approve every one. Generate liberally, approve conservatively. · 6 min read Bandits in production: Thompson sampling, explained with coupons What Thompson sampling does, why we ran it as a batch job that edits a traffic split, and what it saves compared with an A/B test. · 7 min read One definition of conversion What building a metric semantic layer taught me about teams. The schema was the easy part. Ownership, and making the shared definition easier than a copy, decided whether it got used.
2025
2024
2023
2022
2021
2020