Posts
21 posts
What I learned tuning retrieval for a multilingual RAG chatbot
Multilingual questions against a multilingual knowledge base: English questions were fine all along, while Swedish, Polish and Indonesian ones quietly lost their documents until each question was translated into the knowledge base's languages.
Deconstructing PlanOut: a coin toss you can repeat
Follow one shopper from a button experiment to the actual SHA-1 bytes, then look at salts, namespaces and the gap between assignment and exposure.
Evaluating an agent honestly
Relevance, retrieval recall, groundedness and latency each catch a different failure. Here is how I use them as evidence, and why the report is not a deploy gate.
An agent that triages its own feedback
Replay the flagged question, check the evidence, name the likely cause, route it to an owner, and send anything doubtful to a person.
Golden datasets without the gold rush
Simulated personas propose evaluation queries anchored to real documents, and people approve every one. Generate liberally, approve conservatively.
Bandits in production: Thompson sampling, explained with coupons
What Thompson sampling does, why we ran it as a batch job that edits a traffic split, and what it saves compared with an A/B test.
One definition of conversion
What building a metric semantic layer taught me about teams. The schema was the easy part. Ownership, and making the shared definition easier than a copy, decided whether it got used.
The exposure triangle, but for learning rates
Aperture, shutter and ISO behave a lot like learning rate, batch size and warmup.
A serving SDK data scientists actually use
An SDK is an interface, so design it like one. What we hid, what we kept visible, and why a narrow path to production gets used.
SSE on Kubernetes: what a service mesh fixes, and what it doesn't
I wanted our experimentation service to push changes instead of being asked for them. What server-sent events are, what a service mesh like Istio does, and which problems it solves for long-lived connections.
FM synthesis and explore–exploit
Where my initials come from, and why modulation feels like bandits.
From polling to push: getting experiment config to every service
The architecture we had, the one we moved to, and the designs I'd weigh today, for the unglamorous job of keeping experiment config fresh everywhere.
From a click to a database: how a request crosses the internet
Every step from typing an address to a row in a database, each piece explained with an everyday analogy, then how the same pieces shape regional and global deployments, and the infra and ops practices that follow.
Making sense of Kubernetes
My first cluster felt like a pile of YAML, charts and unfamiliar names. The one idea underneath it, the parts, namespaces, Helm, and how traffic finds a pod.
Real or luck? P-values, significance and statistical tests
How to tell a real difference from random noise. We start with ten coin flips you can count by hand, then use exactly the same idea on an A/B test.
Probability and distributions, one step at a time
Probability from scratch, with a die and a coin. What a chance really promises, what a distribution is, and why a bell curve is read by area.
Image colorization: learning what grayscale leaves out
How local detail and scene context help a neural network predict colour, and why plausible is not the same as correct.
Learning data streaming by building one
A student pipeline that ranked trending artists from live tweets, the ideas behind streaming, and how I'd run it on a real cluster.