<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator><link href="https://fadhilmch.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://fadhilmch.github.io/" rel="alternate" type="text/html" /><updated>2026-10-10T11:47:58+02:00</updated><id>https://fadhilmch.github.io/feed.xml</id><title type="html">Fadhil Mochammad</title><subtitle>Machine learning platforms, experimentation, and agent systems in production.</subtitle><author><name>Fadhil Mochammad</name></author><entry><title type="html">What I learned tuning retrieval for a multilingual RAG chatbot</title><link href="https://fadhilmch.github.io/posts/tuning-retrieval-for-multilingual-rag/" rel="alternate" type="text/html" title="What I learned tuning retrieval for a multilingual RAG chatbot" /><published>2026-10-02T00:00:00+02:00</published><updated>2026-10-02T00:00:00+02:00</updated><id>https://fadhilmch.github.io/posts/tuning-retrieval-for-multilingual-rag</id><content type="html" xml:base="https://fadhilmch.github.io/posts/tuning-retrieval-for-multilingual-rag/"><![CDATA[<style>
.rq-fig .fg { fill: var(--muted); opacity: .45; }
.rq-fig .fm { fill: var(--muted); }
.rq-fig .ring { fill: var(--bg); stroke: var(--muted); stroke-width: 2; }
.rq-fig .ta { font-size: 12px; }
.prose .fig.rq-explainer { overflow: visible; }
</style>

<p>The chatbot I’ve been building answers questions about company policies, most of which are specific to one country. That makes it a multilingual problem twice over. The <strong>questions</strong> come in whatever language people like to write in. The <strong>knowledge base</strong> is multilingual too, and each country’s mix is different.</p>

<p>Every answer depends on retrieval. If the right document isn’t among the few the language model gets to read, no prompt will fix the answer. It will either say it doesn’t know or, worse, write something plausible from the wrong source. So before touching prompts I spent a few weeks on the search step: about two dozen configurations, each scored against the same golden set of questions.</p>

<p>This post covers three countries, each with its own knowledge base:</p>

<ul>
  <li><strong>Sweden</strong>: policy pages in Swedish and in English. Questions arrive in English, Swedish and Indonesian.</li>
  <li><strong>Poland</strong>: bilingual pages, Polish and English side by side. Questions arrive in English and Polish.</li>
  <li><strong>United States</strong>: pages in English only, and questions in English only.</li>
</ul>

<figure class="fig rq-fig">
<svg viewBox="0 0 680 280" role="img" aria-labelledby="rq0t rq0d">
  <title id="rq0t">Which question languages meet which knowledge bases</title>
  <desc id="rq0d">Questions in four languages on the left, three country knowledge bases on the right. English questions go to all three: Sweden, with pages in Swedish and English; Poland, with pages in Polish and English; and the United States, with pages in English. Swedish questions go to Sweden and Polish questions to Poland; in both cases the question shares a language with some of the pages. Indonesian questions go to Sweden, where no page is written in Indonesian.</desc>
  <text class="h" x="10" y="16">QUESTION LANGUAGE</text>
  <text class="h" x="670" y="16" text-anchor="end">KNOWLEDGE BASE</text>
  <path class="sa" d="M150 52 C 300 52, 330 58, 450 58" />
  <path class="sa" d="M150 52 C 300 52, 330 140, 450 140" />
  <path class="sa" d="M150 52 C 300 52, 330 222, 450 222" />
  <path class="sa" d="M150 112 C 300 112, 330 58, 450 58" />
  <path class="sa" d="M150 172 C 300 172, 330 140, 450 140" />
  <path class="sb dash" d="M150 232 C 300 232, 330 58, 450 58" />
  <rect class="box" x="10" y="35" width="140" height="34" rx="6" />
  <text class="t" x="24" y="56">English</text>
  <rect class="box" x="10" y="95" width="140" height="34" rx="6" />
  <text class="t" x="24" y="116">Swedish</text>
  <rect class="box" x="10" y="155" width="140" height="34" rx="6" />
  <text class="t" x="24" y="176">Polish</text>
  <rect class="box" x="10" y="215" width="140" height="34" rx="6" />
  <text class="t" x="24" y="236">Indonesian</text>
  <rect class="box" x="450" y="30" width="220" height="56" rx="6" />
  <text class="t" x="466" y="54" style="font-weight:600">Sweden</text>
  <text class="m" x="466" y="72">pages: Swedish, English</text>
  <rect class="box" x="450" y="112" width="220" height="56" rx="6" />
  <text class="t" x="466" y="136" style="font-weight:600">Poland</text>
  <text class="m" x="466" y="154">pages: Polish, English</text>
  <rect class="box" x="450" y="194" width="220" height="56" rx="6" />
  <text class="t" x="466" y="218" style="font-weight:600">United States</text>
  <text class="m" x="466" y="236">pages: English</text>
  <line class="sa" x1="10" x2="40" y1="268" y2="268" /><text class="m" x="48" y="272">shares a language with the pages</text>
  <line class="sb dash" x1="320" x2="350" y1="268" y2="268" /><text class="m" x="358" y="272">no page in this language</text>
</svg>
<figcaption>Which question languages meet which knowledge bases. Every pairing except one shares a language with at least some of the pages. Indonesian is the exception: it stands in for everyone who asks in a language the documents don't use.</figcaption>
</figure>

<p>So there are six combinations of knowledge base and question language, and I’ll call each one a <strong>cell</strong>. Thinking in cells turned out to matter more than any single setting. In short: English questions were fine all along, the overall average hid that questions in Swedish, Polish and Indonesian were quietly losing their documents, and translating each question into the knowledge base’s languages fixed it.</p>

<h2 id="how-i-scored-each-configuration">How I scored each configuration</h2>

<p>Every question in the golden set is annotated with the page or pages a good answer needs, and tagged with its language and the country it’s about. I wrote about building that set in <a href="/posts/golden-datasets-without-the-gold-rush/">Golden datasets without the gold rush</a>.</p>

<p>Two metrics did most of the work:</p>

<ul>
  <li><strong>Recall@10</strong>: the share of the needed pages that appear in the top 10 results. This is the main one, because a page that isn’t retrieved can’t be used.</li>
  <li><strong>MRR</strong> (mean reciprocal rank): how high the first correct page ranks. When recall ties, a higher MRR means the right page is nearer the top, where the model is more likely to use it.</li>
</ul>

<p>I also tracked how many documents each configuration passed on to the answer model, plus latency and cost, which I’ll give as ratios. Each experiment changed one thing from the configuration before it, so every gain or loss has a single cause. The <a href="/posts/evaluating-an-agent-honestly/">previous post</a> explains why I read these numbers as evidence rather than as a pass/fail gate.</p>

<h2 id="step-1-the-search-mode-ladder">Step 1: the search mode ladder</h2>

<p>I started by trying every query mode the search engine offers, from plain keyword matching to hybrid search with a reranker.</p>

<figure class="fig rq-fig">
<svg viewBox="0 0 680 270" role="img" aria-labelledby="rq1t rq1d">
  <title id="rq1t">Recall@10 by search mode</title>
  <desc id="rq1d">Horizontal bars of Recall@10 for six search modes. Hybrid search without a reranker scores 84.0 percent, vector only 84.0, hybrid with a reranker 83.5, keyword with a reranker 74.0, keyword BM25 53.7, and keyword with Lucene query syntax 47.1. The reranker adds 20.3 points to keyword search and changes hybrid recall by half a point, while raising MRR from 0.593 to 0.711.</desc>
  <text class="h" x="10" y="18">BLUE: WITH RERANKER · GREY: WITHOUT</text>
  <line class="grid" x1="230" x2="230" y1="34" y2="242" />
  <text class="m" x="230" y="258" text-anchor="middle">0%</text>
  <line class="grid" x1="330" x2="330" y1="34" y2="242" />
  <text class="m" x="330" y="258" text-anchor="middle">25%</text>
  <line class="grid" x1="430" x2="430" y1="34" y2="242" />
  <text class="m" x="430" y="258" text-anchor="middle">50%</text>
  <line class="grid" x1="530" x2="530" y1="34" y2="242" />
  <text class="m" x="530" y="258" text-anchor="middle">75%</text>
  <line class="grid" x1="630" x2="630" y1="34" y2="242" />
  <text class="m" x="630" y="258" text-anchor="middle">100%</text>
  <text class="t" x="10" y="53">Hybrid (keyword + vector)</text>
  <g tabindex="0"><title>Hybrid (keyword + vector): 84.0% Recall@10</title><rect class="fg" x="230" y="40" width="336.0" height="18" rx="3" /></g>
  <text class="m" x="572.0" y="53">84.0</text>
  <text class="t" x="10" y="87">Vector only</text>
  <g tabindex="0"><title>Vector only: 84.0% Recall@10</title><rect class="fg" x="230" y="74" width="336.0" height="18" rx="3" /></g>
  <text class="m" x="572.0" y="87">84.0</text>
  <text class="t" x="10" y="121">Hybrid + reranker</text>
  <g tabindex="0"><title>Hybrid + reranker: 83.5% Recall@10</title><rect class="fa" x="230" y="108" width="334.0" height="18" rx="3" /></g>
  <text class="m" x="570.0" y="121">83.5</text>
  <text class="t" x="10" y="155">Keyword + reranker</text>
  <g tabindex="0"><title>Keyword + reranker: 74.0% Recall@10</title><rect class="fa" x="230" y="142" width="296.0" height="18" rx="3" /></g>
  <text class="m" x="532.0" y="155">74.0</text>
  <text class="t" x="10" y="189">Keyword (BM25)</text>
  <g tabindex="0"><title>Keyword (BM25): 53.7% Recall@10</title><rect class="fg" x="230" y="176" width="214.8" height="18" rx="3" /></g>
  <text class="m" x="450.8" y="189">53.7</text>
  <text class="t" x="10" y="223">Keyword, Lucene syntax</text>
  <g tabindex="0"><title>Keyword, Lucene syntax: 47.1% Recall@10</title><rect class="fg" x="230" y="210" width="188.4" height="18" rx="3" /></g>
  <text class="m" x="424.4" y="223">47.1</text>
</svg>
<figcaption>Recall@10 for each search mode, all six cells together. Hover a bar for its value.</figcaption>
</figure>

<p>Keyword search alone found the right pages about half the time. Vector search, which matches meaning rather than words, jumped to 84%, and hybrid search, which runs both and fuses the results, landed in the same place.</p>

<p>The <strong>reranker</strong> is a model that reads the question and each candidate page together and scores how well they match. On keyword search it added 20 points of recall. On hybrid search, recall barely moved (84.0% to 83.5%), but MRR jumped from 0.593 to 0.711: the same pages, with the right one much nearer the top.</p>

<p>I kept it, for two reasons. The answer model reads the top results first, so ranking matters. And the reranker’s score is what makes it possible to cut the list with a threshold later. But that flat recall number was hiding something, and it came back in the next step.</p>

<h2 id="step-2-make-it-cheaper-then-look-at-the-cells">Step 2: make it cheaper, then look at the cells</h2>

<p>Ten documents per question is a lot for the answer model to read, so the next changes were about passing on less.</p>

<ul>
  <li>A <strong>threshold</strong> on the reranker score drops weak candidates instead of always passing ten.</li>
  <li>A <strong>country filter</strong> limits the search to documents for the country the question is about, plus those that apply everywhere.</li>
</ul>

<p>Together they cost half a point (83.5% to 83.0%) and cut the average number of documents passed on from 10 to 7.9. Up to this point, every configuration had also appended the country name to the query text (“… in Sweden”). With the filter in place that looked redundant, and it hurt other kinds of questions, so the next change removed it. Recall dropped to 81.0%.</p>

<p>Two and a half points down from the base, for a cleaner query and 6.8 documents. Then I split the results into cells.</p>

<figure class="fig rq-fig">
<svg viewBox="0 0 680 338" role="img" aria-labelledby="rq2t rq2d">
  <title id="rq2t">What the average hid: recall by knowledge base and question language</title>
  <desc id="rq2d">Recall@10 for each knowledge base and question language under three configurations: hybrid search without a reranker, with a reranker, and with the country filter, a reranker threshold and the country name no longer appended. Sweden: English questions 94.5, 93.5, 94.0 percent; Swedish questions 90, 70, 55; Indonesian questions 90, 100, 85. Poland: English questions 79.3, 82.0, 83.3; Polish questions 85, 80, 75. United States: English questions 67.8, 68.6, 65.3.</desc>
  <circle class="ring" cx="15" cy="14" r="5" />
  <text class="m" x="25" y="18">hybrid, no reranker</text>
  <circle class="fm" cx="179.3" cy="14" r="5" />
  <text class="m" x="189.3" y="18">+ reranker</text>
  <circle class="fb" cx="283.3" cy="14" r="5" />
  <text class="m" x="293.3" y="18">+ filter, threshold, no country name</text>
  <line class="grid" x1="230.0" x2="230.0" y1="38" y2="310" />
  <text class="m" x="230.0" y="328" text-anchor="middle">50%</text>
  <line class="grid" x1="312.0" x2="312.0" y1="38" y2="310" />
  <text class="m" x="312.0" y="328" text-anchor="middle">60%</text>
  <line class="grid" x1="394.0" x2="394.0" y1="38" y2="310" />
  <text class="m" x="394.0" y="328" text-anchor="middle">70%</text>
  <line class="grid" x1="476.0" x2="476.0" y1="38" y2="310" />
  <text class="m" x="476.0" y="328" text-anchor="middle">80%</text>
  <line class="grid" x1="558.0" x2="558.0" y1="38" y2="310" />
  <text class="m" x="558.0" y="328" text-anchor="middle">90%</text>
  <line class="grid" x1="640.0" x2="640.0" y1="38" y2="310" />
  <text class="m" x="640.0" y="328" text-anchor="middle">100%</text>
  <text class="h" x="10" y="50">SWEDEN · PAGES IN SWEDISH AND ENGLISH</text>
  <text class="t" x="22" y="78">English questions</text>
  <line class="ln" x1="586.7" x2="594.9" y1="74" y2="74" />
  <g tabindex="0"><title>English questions, hybrid, no reranker: 94.5%</title><circle class="ring" cx="594.9" cy="74" r="6" /></g>
  <g tabindex="0"><title>English questions, + reranker: 93.5%</title><circle class="fm" cx="586.7" cy="74" r="6" /></g>
  <g tabindex="0"><title>English questions, + filter, threshold, no country name: 94.0%</title><circle class="fb" cx="590.8" cy="74" r="6" /></g>
  <text class="m" x="590.8" y="64" text-anchor="middle">94.0</text>
  <text class="t" x="22" y="108">Swedish questions</text>
  <line class="ln" x1="271.0" x2="558.0" y1="104" y2="104" />
  <g tabindex="0"><title>Swedish questions, hybrid, no reranker: 90.0%</title><circle class="ring" cx="558.0" cy="104" r="6" /></g>
  <g tabindex="0"><title>Swedish questions, + reranker: 70.0%</title><circle class="fm" cx="394.0" cy="104" r="6" /></g>
  <g tabindex="0"><title>Swedish questions, + filter, threshold, no country name: 55.0%</title><circle class="fb" cx="271.0" cy="104" r="6" /></g>
  <text class="m" x="271.0" y="94" text-anchor="middle">55.0</text>
  <text class="t" x="22" y="138">Indonesian questions</text>
  <line class="ln" x1="517.0" x2="640.0" y1="134" y2="134" />
  <g tabindex="0"><title>Indonesian questions, hybrid, no reranker: 90.0%</title><circle class="ring" cx="558.0" cy="134" r="6" /></g>
  <g tabindex="0"><title>Indonesian questions, + reranker: 100.0%</title><circle class="fm" cx="640.0" cy="134" r="6" /></g>
  <g tabindex="0"><title>Indonesian questions, + filter, threshold, no country name: 85.0%</title><circle class="fb" cx="517.0" cy="134" r="6" /></g>
  <text class="m" x="517.0" y="124" text-anchor="middle">85.0</text>
  <text class="h" x="10" y="170">POLAND · PAGES IN POLISH AND ENGLISH</text>
  <text class="t" x="22" y="198">English questions</text>
  <line class="ln" x1="470.3" x2="503.1" y1="194" y2="194" />
  <g tabindex="0"><title>English questions, hybrid, no reranker: 79.3%</title><circle class="ring" cx="470.3" cy="194" r="6" /></g>
  <g tabindex="0"><title>English questions, + reranker: 82.0%</title><circle class="fm" cx="492.4" cy="194" r="6" /></g>
  <g tabindex="0"><title>English questions, + filter, threshold, no country name: 83.3%</title><circle class="fb" cx="503.1" cy="194" r="6" /></g>
  <text class="m" x="503.1" y="184" text-anchor="middle">83.3</text>
  <text class="t" x="22" y="228">Polish questions</text>
  <line class="ln" x1="435.0" x2="517.0" y1="224" y2="224" />
  <g tabindex="0"><title>Polish questions, hybrid, no reranker: 85.0%</title><circle class="ring" cx="517.0" cy="224" r="6" /></g>
  <g tabindex="0"><title>Polish questions, + reranker: 80.0%</title><circle class="fm" cx="476.0" cy="224" r="6" /></g>
  <g tabindex="0"><title>Polish questions, + filter, threshold, no country name: 75.0%</title><circle class="fb" cx="435.0" cy="224" r="6" /></g>
  <text class="m" x="435.0" y="214" text-anchor="middle">75.0</text>
  <text class="h" x="10" y="260">UNITED STATES · PAGES IN ENGLISH</text>
  <text class="t" x="22" y="288">English questions</text>
  <line class="ln" x1="355.5" x2="382.5" y1="284" y2="284" />
  <g tabindex="0"><title>English questions, hybrid, no reranker: 67.8%</title><circle class="ring" cx="376.0" cy="284" r="6" /></g>
  <g tabindex="0"><title>English questions, + reranker: 68.6%</title><circle class="fm" cx="382.5" cy="284" r="6" /></g>
  <g tabindex="0"><title>English questions, + filter, threshold, no country name: 65.3%</title><circle class="fb" cx="355.5" cy="284" r="6" /></g>
  <text class="m" x="355.5" y="274" text-anchor="middle">65.3</text>
</svg>
<figcaption>Recall@10 for each cell, grouped by knowledge base. Overall recall fell 3 points across these steps. Swedish questions lost 35.</figcaption>
</figure>

<p><strong>English questions held up in every knowledge base.</strong> Every country has English pages, so an English question always shares a language with something it can find. Sweden’s English questions stayed around 94%, Poland’s rose 4 points, and the United States’ dipped by 2.5. Nothing dramatic.</p>

<p><strong>The other three cells told a different story:</strong></p>

<ul>
  <li><strong>Swedish questions against Sweden’s mixed knowledge base</strong> fell hardest: 90% with plain hybrid search, 70% with the reranker, 55% at the end. The MRR numbers add a twist. Without the reranker, these questions found their pages but ranked them poorly, with an MRR of just 0.17. The reranker pulled the pages it kept up the list (MRR 0.45), but pushed others out of the top 10 entirely.</li>
  <li><strong>Polish questions against Poland’s bilingual pages</strong> slid from 85% to 75%. A smaller loss than Swedish, even though both are questions in the local language.</li>
  <li><strong>Indonesian questions against Sweden</strong>, which share no language with any page, actually did well with the reranker (100%, and MRR up from 0.37 to 0.78). They only fell, to 85%, when the country name came out of the query.</li>
</ul>

<p>I don’t have a complete explanation for the reranker step. The cell that lost most to it, Swedish questions, is one where the question does share a language with some of the pages, while the cell with no shared language gained. So it isn’t simply “non-English questions do worse”. My guess is that the reranker is weaker on Swedish text in particular. The last step is easier to read: removing the appended country name cost all three non-English cells 5 to 15 points, while the English cells as a whole didn’t change. An English word like “Sweden” seems to have been giving those questions an anchor in the English pages.</p>

<p>These cells are small, so one question moves the number by several points. Even so, a 35-point drop for Swedish questions is not something an overall average should be allowed to hide.</p>

<h2 id="step-3-translate-into-the-knowledge-bases-languages">Step 3: translate into the knowledge base’s languages</h2>

<p>The fix was to stop searching with the question as written, and search in the knowledge base’s languages instead. The walkthrough below goes through it step by step. Steps 1 and 2 recap the single search path, 3 to 5 cover translation and merging, and 6 covers a third path I’ll come back to.</p>

<figure class="fig rq-explainer">
<iframe id="rq-explainer" src="/assets/posts/multilingual-rag/retrieval-explainer.html" title="Animated walkthrough: two translated queries and a max-score merge" loading="lazy" style="display:block;width:100%;height:640px;border:0"></iframe>
<figcaption>Six steps. Use the numbered buttons or the arrows, or drag the timeline. The example question and documents are invented.</figcaption>
</figure>
<script>
window.addEventListener('message', function (event) {
  if (event.origin !== window.location.origin) return;
  var data = event.data;
  if (data && data.explainer === 'multilingual-rag' && data.height) {
    document.getElementById('rq-explainer').style.height = data.height + 'px';
  }
});
</script>

<p>Every question now becomes two queries, one per language the knowledge base is written in:</p>

<ol>
  <li>An <strong>English</strong> translation, because every country has English pages.</li>
  <li>A translation into the <strong>country’s own language</strong>: Swedish for Sweden, Polish for Poland. For the United States, whose pages are all English, both queries end up in English.</li>
</ol>

<p>The language the person wrote in is ignored. An Indonesian question about Sweden gets English and Swedish queries, not an Indonesian one, because no pages are written in Indonesian and searching for Indonesian text would find nothing. Picking languages from the knowledge base’s side, not the user’s, is the part of this design that matters most.</p>

<p>Both queries run in parallel through the same hybrid search and reranker. The two result lists are merged by keeping each page’s higher score, and the same threshold applies to the merged list.</p>

<figure class="fig rq-fig">
<svg viewBox="0 0 680 338" role="img" aria-labelledby="rq4t rq4d">
  <title id="rq4t">Recall by knowledge base and question language, one query vs two translated queries</title>
  <desc id="rq4d">Recall@10 for each knowledge base and question language, before and after translating each question into English and the country language. Sweden: English questions 94.0 to 93.5 percent; Swedish questions 55 to 100; Indonesian questions 85 to 100. Poland: English questions 83.3 both times; Polish questions 75 to 85. United States: 65.3 both times.</desc>
  <circle class="fb" cx="15" cy="14" r="5" />
  <text class="m" x="25" y="18">one query, no country name</text>
  <circle class="fa" cx="226.20000000000002" cy="14" r="5" />
  <text class="m" x="236.20000000000002" y="18">translated: English + local language</text>
  <line class="grid" x1="230.0" x2="230.0" y1="38" y2="310" />
  <text class="m" x="230.0" y="328" text-anchor="middle">50%</text>
  <line class="grid" x1="312.0" x2="312.0" y1="38" y2="310" />
  <text class="m" x="312.0" y="328" text-anchor="middle">60%</text>
  <line class="grid" x1="394.0" x2="394.0" y1="38" y2="310" />
  <text class="m" x="394.0" y="328" text-anchor="middle">70%</text>
  <line class="grid" x1="476.0" x2="476.0" y1="38" y2="310" />
  <text class="m" x="476.0" y="328" text-anchor="middle">80%</text>
  <line class="grid" x1="558.0" x2="558.0" y1="38" y2="310" />
  <text class="m" x="558.0" y="328" text-anchor="middle">90%</text>
  <line class="grid" x1="640.0" x2="640.0" y1="38" y2="310" />
  <text class="m" x="640.0" y="328" text-anchor="middle">100%</text>
  <text class="h" x="10" y="50">SWEDEN · PAGES IN SWEDISH AND ENGLISH</text>
  <text class="t" x="22" y="78">English questions</text>
  <line class="ln" x1="586.7" x2="590.8" y1="74" y2="74" />
  <g tabindex="0"><title>English questions, one query, no country name: 94.0%</title><circle class="fb" cx="590.8" cy="74" r="6" /></g>
  <g tabindex="0"><title>English questions, translated: English + local language: 93.5%</title><circle class="fa" cx="586.7" cy="74" r="6" /></g>
  <text class="m" x="586.7" y="64" text-anchor="middle">93.5</text>
  <text class="t" x="22" y="108">Swedish questions</text>
  <line class="ln" x1="271.0" x2="640.0" y1="104" y2="104" />
  <g tabindex="0"><title>Swedish questions, one query, no country name: 55.0%</title><circle class="fb" cx="271.0" cy="104" r="6" /></g>
  <g tabindex="0"><title>Swedish questions, translated: English + local language: 100.0%</title><circle class="fa" cx="640.0" cy="104" r="6" /></g>
  <text class="m" x="640.0" y="94" text-anchor="middle">100.0</text>
  <text class="t" x="22" y="138">Indonesian questions</text>
  <line class="ln" x1="517.0" x2="640.0" y1="134" y2="134" />
  <g tabindex="0"><title>Indonesian questions, one query, no country name: 85.0%</title><circle class="fb" cx="517.0" cy="134" r="6" /></g>
  <g tabindex="0"><title>Indonesian questions, translated: English + local language: 100.0%</title><circle class="fa" cx="640.0" cy="134" r="6" /></g>
  <text class="m" x="640.0" y="124" text-anchor="middle">100.0</text>
  <text class="h" x="10" y="170">POLAND · PAGES IN POLISH AND ENGLISH</text>
  <text class="t" x="22" y="198">English questions</text>
  <line class="ln" x1="503.1" x2="503.1" y1="194" y2="194" />
  <g tabindex="0"><title>English questions, one query, no country name: 83.3%</title><circle class="fb" cx="503.1" cy="194" r="6" /></g>
  <g tabindex="0"><title>English questions, translated: English + local language: 83.3%</title><circle class="fa" cx="503.1" cy="194" r="6" /></g>
  <text class="m" x="503.1" y="184" text-anchor="middle">83.3</text>
  <text class="t" x="22" y="228">Polish questions</text>
  <line class="ln" x1="435.0" x2="517.0" y1="224" y2="224" />
  <g tabindex="0"><title>Polish questions, one query, no country name: 75.0%</title><circle class="fb" cx="435.0" cy="224" r="6" /></g>
  <g tabindex="0"><title>Polish questions, translated: English + local language: 85.0%</title><circle class="fa" cx="517.0" cy="224" r="6" /></g>
  <text class="m" x="517.0" y="214" text-anchor="middle">85.0</text>
  <text class="h" x="10" y="260">UNITED STATES · PAGES IN ENGLISH</text>
  <text class="t" x="22" y="288">English questions</text>
  <line class="ln" x1="355.5" x2="355.5" y1="284" y2="284" />
  <g tabindex="0"><title>English questions, one query, no country name: 65.3%</title><circle class="fb" cx="355.5" cy="284" r="6" /></g>
  <g tabindex="0"><title>English questions, translated: English + local language: 65.3%</title><circle class="fa" cx="355.5" cy="284" r="6" /></g>
  <text class="m" x="355.5" y="274" text-anchor="middle">65.3</text>
</svg>
<figcaption>Recall@10 for each cell, before and after translating each question into the knowledge base's languages.</figcaption>
</figure>

<p>Overall recall went from 81.0% to 85.5%. The gain landed exactly where the problem was:</p>

<ul>
  <li><strong>Swedish questions</strong> went from 55% to 100%, and their MRR from 0.39 to 0.67. Both problems from step 2, missing pages and poor ranking, improved at once.</li>
  <li><strong>Indonesian questions</strong> went from 85% to 100%.</li>
  <li><strong>Polish questions</strong> went from 75% to 85%, back to where they started before the reranker.</li>
  <li><strong>English questions</strong> stayed within a point of where they were, in all three knowledge bases. For an English question, the English translation is close to the original, and in the United States there’s nothing else to translate into.</li>
</ul>

<p>Fifteen questions that used to miss now found their page, and fourteen of them were asked in Swedish, Indonesian or Polish. One English question got worse. The model reads 6.4 documents on average, slightly fewer than before. Compared with the version that still had the country name in the query, recall is 2.5 points higher: a proper translation does what the appended word was doing by accident.</p>

<p>What I can’t tell you yet is how the work splits between the two queries. I didn’t run each one on its own, so I don’t know how much of the gain comes from the English query and how much from the local one. That’s the next experiment.</p>

<p>The cost is one translation call per question, which is the only language-model call in retrieval. Retrieval latency a little more than doubled, but it’s still a small part of the time a full answer takes, and cost per question rose by about a quarter.</p>

<h2 id="merging-keep-the-best-score-dont-fuse-ranks">Merging: keep the best score, don’t fuse ranks</h2>

<p>The usual way to merge ranked lists is <strong>reciprocal rank fusion</strong> (RRF). Each document gets 1/(60 + rank) for every list it appears in, and the totals are summed. It ignores the scores completely, which is its strength when the lists come from systems whose scores aren’t comparable.</p>

<p>Here the scores are comparable, because both paths are scored by the same reranker. Taking the maximum keeps that information: a page one path is confident about stays near the top even if the other path missed it. RRF instead rewards pages that turn up everywhere, so a page that was mediocre on both paths can overtake one that a single path was sure of. Step 5 of the walkthrough shows this happening.</p>

<p>When I tested both, recall was the same (84.0% against 83.7%), but max score ranked the right page higher: MRR 0.706 against 0.637. There’s a practical reason too. A threshold only makes sense on a score with a fixed meaning, and RRF totals don’t have one.</p>

<p>The condition matters, though. If your paths are different retrievers with different score scales, such as raw keyword scores next to vector similarities, the max is meaningless and RRF or score normalisation is the right choice.</p>

<h2 id="hyde-neutral-alone-marginal-together">HyDE: neutral alone, marginal together</h2>

<p><strong>HyDE</strong> (Hypothetical Document Embeddings) asks a language model to write a plausible answer, then searches with that passage instead of the question. The idea is that an answer-shaped passage looks more like the documents than a question does.</p>

<p>Used on its own instead of the question, it changed nothing overall.</p>

<figure class="fig rq-fig">
<svg viewBox="0 0 680 350" role="img" aria-labelledby="rq3t rq3d">
  <title id="rq3t">HyDE used alone: change in Recall@10 by knowledge base and question language</title>
  <desc id="rq3d">Change in Recall@10 when the question is replaced by a HyDE passage, against the same search without HyDE. All questions: no change. Sweden: English questions minus 0.5 points, Swedish questions plus 5.0, Indonesian questions minus 5.0. Poland: English questions plus 0.7, Polish questions no change. United States: no change.</desc>
  <text class="h" x="10" y="18">BLUE HELPS · ORANGE HURTS · POINTS OF RECALL@10</text>
  <line class="grid" x1="270" x2="270" y1="34" y2="320" />
  <text class="m" x="270" y="336" text-anchor="middle">-10</text>
  <line class="grid" x1="360" x2="360" y1="34" y2="320" />
  <text class="m" x="360" y="336" text-anchor="middle">-5</text>
  <line class="grid" x1="450" x2="450" y1="34" y2="320" />
  <text class="m" x="450" y="336" text-anchor="middle">0</text>
  <line class="grid" x1="540" x2="540" y1="34" y2="320" />
  <text class="m" x="540" y="336" text-anchor="middle">+5</text>
  <line class="grid" x1="630" x2="630" y1="34" y2="320" />
  <text class="m" x="630" y="336" text-anchor="middle">+10</text>
  <text class="t" x="10" y="53">All questions</text>
  <g tabindex="0"><title>All questions: +0.0 points</title><rect class="fm" x="450.0" y="40" width="1.5" height="18" rx="3" /></g>
  <text class="m" x="458" y="53">no change</text>
  <text class="h" x="10" y="76">SWEDEN · PAGES IN SWEDISH AND ENGLISH</text>
  <text class="t" x="22" y="107">English questions</text>
  <g tabindex="0"><title>English questions: -0.5 points</title><rect class="fb" x="441.0" y="94" width="9.0" height="18" rx="3" /></g>
  <text class="m" x="435.0" y="107" text-anchor="end">-0.5</text>
  <text class="t" x="22" y="135">Swedish questions</text>
  <g tabindex="0"><title>Swedish questions: +5.0 points</title><rect class="fa" x="450.0" y="122" width="90.0" height="18" rx="3" /></g>
  <text class="m" x="546.0" y="135" text-anchor="start">+5.0</text>
  <text class="t" x="22" y="163">Indonesian questions</text>
  <g tabindex="0"><title>Indonesian questions: -5.0 points</title><rect class="fb" x="360.0" y="150" width="90.0" height="18" rx="3" /></g>
  <text class="m" x="354.0" y="163" text-anchor="end">-5.0</text>
  <text class="h" x="10" y="188">POLAND · PAGES IN POLISH AND ENGLISH</text>
  <text class="t" x="22" y="219">English questions</text>
  <g tabindex="0"><title>English questions: +0.7 points</title><rect class="fa" x="450.0" y="206" width="12.6" height="18" rx="3" /></g>
  <text class="m" x="468.6" y="219" text-anchor="start">+0.7</text>
  <text class="t" x="22" y="247">Polish questions</text>
  <g tabindex="0"><title>Polish questions: +0.0 points</title><rect class="fm" x="450.0" y="234" width="1.5" height="18" rx="3" /></g>
  <text class="m" x="458" y="247">no change</text>
  <text class="h" x="10" y="272">UNITED STATES · PAGES IN ENGLISH</text>
  <text class="t" x="22" y="303">English questions</text>
  <g tabindex="0"><title>English questions: +0.0 points</title><rect class="fm" x="450.0" y="290" width="1.5" height="18" rx="3" /></g>
  <text class="m" x="458" y="303">no change</text>
  <line class="ln" x1="450" x2="450" y1="34" y2="320" />
</svg>
<figcaption>Change in Recall@10 for each cell when the question is replaced by a HyDE passage, against the same search without HyDE.</figcaption>
</figure>

<p>Overall recall stayed at 83.0%: eight questions gained, seven lost, and retrieval was four times slower. Underneath, Swedish questions gained 5 points and Indonesian questions lost 5. That mix is what made it worth one more try. A technique that finds <em>different</em> pages can be useful as an extra path even if it doesn’t beat the original, so I added it as a <strong>third parallel path</strong> with the same max-score merge. A path that only adds candidates can’t push good results out, because the merge keeps whichever score is highest.</p>

<p>Next to the two translated queries, it added 0.6 points: 86.1% against 85.5%, three questions gained (two of them in the United States) and one lost. That cost 2.4 times the retrieval latency and 4.6 times the language-model tokens. For a chatbot, where people are waiting for an answer, that isn’t worth it, so the chat setup uses two paths.</p>

<p>The lesson is to judge an extra path against your best setup, not a weak one. Once the translations were in place, most of what HyDE could have caught was already being found.</p>

<h2 id="every-step-every-cell">Every step, every cell</h2>

<p>The matrix below puts the whole story in one place. Rows are knowledge bases, columns are question languages, and each cell shows how that combination did. Step through the configurations, or press play, and watch which cells move. Switch to MRR to see the ranking side.</p>

<style>
.rq-lab { margin: 32px 0; padding: 20px; border: 1px solid var(--line); border-radius: 8px; background: var(--panel); box-sizing: border-box; }
@media (min-width: 900px) { .prose .rq-lab { width: min(680px, calc(100vw - 72px)); } }
.prose .rq-lab h3 { margin: 0 0 6px; }
.prose .rq-lab p { margin: 0 0 12px; font-size: 14px; color: var(--muted); }
.rq-lab .rq-js { display: none; }
.rq-lab.is-live .rq-js { display: flex; }
.rq-bar { flex-wrap: wrap; gap: 6px; align-items: center; margin: 0 0 10px; }
.rq-lab button { font: 12px 'Geist Mono', monospace; padding: 6px 10px; min-height: 36px; border: 1px solid var(--line); border-radius: 6px; background: var(--bg); color: var(--fg); cursor: pointer; }
.rq-lab button[aria-pressed="true"] { border-color: var(--fig-a); background: color-mix(in oklch, var(--fig-a) 16%, var(--bg)); }
.rq-lab button:focus-visible { outline: 2px solid var(--fig-a); outline-offset: 2px; }
.rq-lab .rq-play { min-width: 72px; }
.rq-bar .rq-label { font: 11px 'Geist Mono', monospace; color: var(--muted); letter-spacing: .04em; margin-right: 4px; }
.prose .rq-lab .rq-note { min-height: 3.2em; margin: 4px 0 14px; color: var(--fg); font-size: 14px; }
.rq-wrap { overflow-x: auto; }
table.rq-grid { width: 100%; min-width: 520px; border-collapse: separate; border-spacing: 4px; font: 12px 'Geist Mono', monospace; margin: 0; }
.rq-grid th { text-align: left; font-weight: 500; color: var(--muted); padding: 4px 6px; vertical-align: bottom; border: 0; background: none; }
.rq-grid th[scope="row"] { color: var(--fg); vertical-align: middle; width: 26%; }
.rq-grid th[scope="row"] small { display: block; color: var(--muted); font-weight: 400; font-size: 11px; }
.rq-grid td { height: 58px; padding: 6px 8px; border: 0; border-radius: 6px; vertical-align: top; background: color-mix(in oklch, var(--fig-a) var(--p, 8%), var(--panel)); transition: background-color .45s ease; }
.rq-grid td.na { background: repeating-linear-gradient(135deg, transparent 0 6px, var(--line) 6px 7px); color: var(--muted); vertical-align: middle; }
.rq-grid td.nomatch { box-shadow: inset 0 0 0 1.5px var(--fig-b); }
.rq-val { display: block; font-size: 18px; font-weight: 600; color: var(--fg); }
.rq-delta { font-size: 11px; color: var(--fg); opacity: .8; }
.rq-delta.up::before { content: '▲ '; color: var(--fig-a); }
.rq-delta.down::before { content: '▼ '; color: var(--fig-b); }
.rq-sum { display: flex; flex-wrap: wrap; gap: 8px 22px; margin-top: 12px; font: 12px 'Geist Mono', monospace; color: var(--muted); }
.rq-sum b { color: var(--fg); font-size: 14px; font-weight: 600; }
.rq-legend { margin-top: 10px; font: 11px/1.6 'Geist Mono', monospace; color: var(--muted); }
.rq-legend i { display: inline-block; width: 12px; height: 12px; border-radius: 3px; vertical-align: -2px; margin-right: 4px; }
@media (prefers-reduced-motion: reduce) { .rq-grid td { transition: none; } }
@media (max-width: 420px) { .rq-lab { padding: 14px; } }
</style>

<section class="rq-lab" id="rq-matrix" aria-labelledby="rq-matrix-title">
  <h3 id="rq-matrix-title">Six cells, six configurations</h3>
  <p>Rows are knowledge bases, columns are the language the question was asked in. Darker cells score higher.</p>
  <div class="rq-bar rq-js" role="group" aria-label="Configuration">
    <span class="rq-label">STEP</span>
  </div>
  <div class="rq-bar rq-js" role="group" aria-label="Metric and playback">
    <span class="rq-label">SHOW</span>
    <button type="button" data-metric="recall" aria-pressed="true">Recall@10</button>
    <button type="button" data-metric="mrr" aria-pressed="false">MRR</button>
    <button type="button" class="rq-play">Play</button>
  </div>
  <p class="rq-note" aria-live="polite">Configuration used for chat: English and country-language translations, merged by max score.</p>
  <div class="rq-wrap">
    <table class="rq-grid">
      <thead>
        <tr><th scope="col">Knowledge base</th><th scope="col">English</th><th scope="col">Swedish</th><th scope="col">Polish</th><th scope="col">Indonesian</th></tr>
      </thead>
      <tbody>
        <tr>
          <th scope="row">Sweden<small>Swedish + English pages</small></th>
          <td data-cell="SE-en"><span class="rq-val">93.5%</span><span class="rq-delta"></span></td>
          <td data-cell="SE-sw"><span class="rq-val">100.0%</span><span class="rq-delta"></span></td>
          <td class="na">not asked</td>
          <td data-cell="SE-in" class="nomatch"><span class="rq-val">100.0%</span><span class="rq-delta"></span></td>
        </tr>
        <tr>
          <th scope="row">Poland<small>Polish + English pages</small></th>
          <td data-cell="PL-en"><span class="rq-val">83.3%</span><span class="rq-delta"></span></td>
          <td class="na">not asked</td>
          <td data-cell="PL-po"><span class="rq-val">85.0%</span><span class="rq-delta"></span></td>
          <td class="na">not asked</td>
        </tr>
        <tr>
          <th scope="row">United States<small>English pages</small></th>
          <td data-cell="US-en"><span class="rq-val">65.3%</span><span class="rq-delta"></span></td>
          <td class="na">not asked</td>
          <td class="na">not asked</td>
          <td class="na">not asked</td>
        </tr>
      </tbody>
    </table>
  </div>
  <div class="rq-sum" aria-live="polite">
    <span>all questions <b data-sum="recall">85.5%</b></span>
    <span>MRR <b data-sum="mrr">0.705</b></span>
    <span>documents kept <b data-sum="docs">6.4</b></span>
  </div>
  <div class="rq-legend"><i style="box-shadow: inset 0 0 0 1.5px var(--fig-b)"></i>orange outline: no page is written in the question's language &nbsp; <i style="background: repeating-linear-gradient(135deg, transparent 0 4px, var(--line) 4px 5px)"></i>hatched: no questions of this kind &nbsp; ▲ ▼ change from the previous step</div>
</section>

<script>
(() => {
  'use strict';
  const root = document.getElementById('rq-matrix');
  if (!root) return;
  // Recall@10 (%) and MRR per cell, recomputed per knowledge base and question language.
  const STEPS = [
    { label: '1 Hybrid', note: 'Keyword and vector search fused, no reranker. Swedish questions find their pages but rank them low: MRR 0.17.',
      all: [84.0, 0.593, 10.0], cells: { 'SE-en': [94.5, 0.691], 'SE-sw': [90.0, 0.170], 'SE-in': [90.0, 0.372], 'PL-en': [79.3, 0.654], 'PL-po': [85.0, 0.576], 'US-en': [67.8, 0.575] } },
    { label: '2 + reranker', note: 'The reranker lifts ranking almost everywhere, and costs Swedish questions 20 points of recall.',
      all: [83.5, 0.711, 10.0], cells: { 'SE-en': [93.5, 0.828], 'SE-sw': [70.0, 0.447], 'SE-in': [100.0, 0.782], 'PL-en': [82.0, 0.756], 'PL-po': [80.0, 0.710], 'US-en': [68.6, 0.522] } },
    { label: '3 + threshold, filter', note: 'A score threshold and a country filter: 7.9 documents instead of 10, for half a point of recall.',
      all: [83.0, 0.715, 7.9], cells: { 'SE-en': [94.5, 0.843], 'SE-sw': [65.0, 0.437], 'SE-in': [100.0, 0.782], 'PL-en': [82.0, 0.762], 'PL-po': [80.0, 0.710], 'US-en': [66.1, 0.512] } },
    { label: '4 − country name', note: 'The country name is no longer appended to the query. Every non-English cell drops. English cells hold.',
      all: [81.0, 0.700, 6.8], cells: { 'SE-en': [94.0, 0.816], 'SE-sw': [55.0, 0.392], 'SE-in': [85.0, 0.738], 'PL-en': [83.3, 0.761], 'PL-po': [75.0, 0.667], 'US-en': [65.3, 0.529] } },
    { label: '5 + translation', note: 'Configuration used for chat: English and country-language translations, merged by max score.',
      all: [85.5, 0.705, 6.4], cells: { 'SE-en': [93.5, 0.784], 'SE-sw': [100.0, 0.665], 'SE-in': [100.0, 0.701], 'PL-en': [83.3, 0.762], 'PL-po': [85.0, 0.654], 'US-en': [65.3, 0.532] } },
    { label: '6 + HyDE', note: 'A HyDE passage as a third path: 0.6 points more overall, for 2.4 times the retrieval latency.',
      all: [86.1, 0.706, 7.4], cells: { 'SE-en': [94.0, 0.771], 'SE-sw': [100.0, 0.664], 'SE-in': [100.0, 0.717], 'PL-en': [82.7, 0.749], 'PL-po': [85.0, 0.671], 'US-en': [67.8, 0.561] } }
  ];
  const stepBar = root.querySelector('[aria-label="Configuration"]');
  const metricButtons = Array.from(root.querySelectorAll('[data-metric]'));
  const playButton = root.querySelector('.rq-play');
  const note = root.querySelector('.rq-note');
  const cells = Array.from(root.querySelectorAll('td[data-cell]'));
  const sums = { recall: root.querySelector('[data-sum="recall"]'), mrr: root.querySelector('[data-sum="mrr"]'), docs: root.querySelector('[data-sum="docs"]') };
  const reduced = window.matchMedia && window.matchMedia('(prefers-reduced-motion: reduce)').matches;
  let step = 4, metric = 'recall', timer = 0;

  const stepButtons = STEPS.map((s, i) => {
    const b = document.createElement('button');
    b.type = 'button';
    b.textContent = s.label;
    b.addEventListener('click', () => { stop(); show(i); });
    stepBar.appendChild(b);
    return b;
  });

  const value = (pair) => metric === 'recall' ? pair[0] : pair[1];
  const format = (v) => metric === 'recall' ? v.toFixed(1) + '%' : v.toFixed(3);
  // Shade by value: recall 50..100% or MRR 0.1..0.9 maps to 6..56% of the accent colour.
  const shade = (v) => {
    const t = metric === 'recall' ? (v - 50) / 50 : (v - 0.1) / 0.8;
    return (6 + 50 * Math.max(0, Math.min(1, t))).toFixed(1) + '%';
  };

  function show(i) {
    step = i;
    const cur = STEPS[i], prev = i > 0 ? STEPS[i - 1] : null;
    stepButtons.forEach((b, k) => b.setAttribute('aria-pressed', String(k === i)));
    note.textContent = cur.note;
    cells.forEach(td => {
      const key = td.dataset.cell;
      const v = value(cur.cells[key]);
      td.style.setProperty('--p', shade(v));
      td.querySelector('.rq-val').textContent = format(v);
      td.title = key + ': ' + format(v);
      const d = td.querySelector('.rq-delta');
      d.className = 'rq-delta';
      if (!prev) { d.textContent = ''; return; }
      const diff = v - value(prev.cells[key]);
      const small = metric === 'recall' ? Math.abs(diff) < 0.05 : Math.abs(diff) < 0.0005;
      if (small) { d.textContent = 'no change'; return; }
      d.classList.add(diff > 0 ? 'up' : 'down');
      d.textContent = metric === 'recall' ? Math.abs(diff).toFixed(1) + ' pts' : Math.abs(diff).toFixed(3);
    });
    sums.recall.textContent = cur.all[0].toFixed(1) + '%';
    sums.mrr.textContent = cur.all[1].toFixed(3);
    sums.docs.textContent = cur.all[2].toFixed(1);
  }

  function stop() { clearInterval(timer); timer = 0; playButton.textContent = 'Play'; }
  playButton.addEventListener('click', () => {
    if (timer) { stop(); return; }
    show(0);
    playButton.textContent = 'Pause';
    timer = setInterval(() => {
      if (step >= STEPS.length - 1) { stop(); return; }
      show(step + 1);
    }, reduced ? 3000 : 2200);
  });
  metricButtons.forEach(b => b.addEventListener('click', () => {
    metric = b.dataset.metric;
    metricButtons.forEach(x => x.setAttribute('aria-pressed', String(x === b)));
    show(step);
  }));

  root.classList.add('is-live');
  show(step);
})();
</script>

<p>Compared with the hybrid-plus-reranker base, the configuration I use for chat finds the right pages more often (85.5% against 83.5%), ranks them about as well (MRR 0.705 against 0.711), and passes on 6 documents instead of 10. The overall gain looks modest because most questions are in English, and English questions had nothing to gain. Swedish and Indonesian questions went from the worst cells to perfect ones, and Polish questions won back what they had lost.</p>

<p>I also checked top-k. Keeping only 5 results cost over 3 points, and keeping 20 added under a point for twice the reading, so 10 stayed.</p>

<h2 id="whats-still-open">What’s still open</h2>

<ul>
  <li><strong>The United States sits at 65% to 69% in every configuration.</strong> It’s the simplest cell, English questions against English pages, and nothing on the search side moved it. When no retrieval change moves a number, I’d look at the documents next: whether the answer is missing, buried in a long page, or phrased very differently from how people ask.</li>
  <li><strong>Measuring each query on its own.</strong> An English-only run and a local-language-only run would show how the work splits between the two translations, cell by cell.</li>
</ul>

<h2 id="what-id-tell-myself-at-the-start">What I’d tell myself at the start</h2>

<ol>
  <li><strong>Map your questions’ languages against your knowledge base’s languages before tuning anything.</strong> The cells behave differently, and the same setting can help one and hurt another.</li>
  <li><strong>Split every result by those cells before trusting it.</strong> My biggest problem showed up as a three-point dip in the average.</li>
  <li><strong>Judge the reranker by MRR, and check it per language.</strong> Its overall recall barely moved, while it cost Swedish questions 20 points.</li>
  <li><strong>Search in the knowledge base’s languages, not the user’s.</strong> English plus the country’s own language, whatever the question was written in. Don’t rely on a stray English word in the query to do that job.</li>
  <li><strong>Merge on scores when they share a scale, and keep a threshold.</strong> Use RRF when they don’t.</li>
  <li><strong>Try new techniques as extra paths, and measure the gain against your best setup.</strong></li>
</ol>

<h2 id="references">References</h2>

<ol>
  <li>Robertson, S. and Zaragoza, H. <em>The Probabilistic Relevance Framework: BM25 and Beyond</em>. Foundations and Trends in Information Retrieval, 2009. <a href="https://doi.org/10.1561/1500000019">https://doi.org/10.1561/1500000019</a></li>
  <li>Karpukhin, V. et al. <em>Dense Passage Retrieval for Open-Domain Question Answering</em>. EMNLP, 2020. <a href="https://arxiv.org/abs/2004.04906">https://arxiv.org/abs/2004.04906</a></li>
  <li>Nogueira, R. and Cho, K. <em>Passage Re-ranking with BERT</em>. arXiv, 2019. <a href="https://arxiv.org/abs/1901.04085">https://arxiv.org/abs/1901.04085</a>: the cross-encoder reranking idea.</li>
  <li>Cormack, G. V., Clarke, C. L. A. and Büttcher, S. <em>Reciprocal Rank Fusion Outperforms Condorcet and Individual Rank Learning Methods</em>. SIGIR, 2009. <a href="https://doi.org/10.1145/1571941.1572114">https://doi.org/10.1145/1571941.1572114</a></li>
  <li>Gao, L., Ma, X., Lin, J. and Callan, J. <em>Precise Zero-Shot Dense Retrieval without Relevance Labels</em>. ACL, 2023. <a href="https://arxiv.org/abs/2212.10496">https://arxiv.org/abs/2212.10496</a>: HyDE.</li>
  <li>Rackauckas, Z. <em>RAG-Fusion: a New Take on Retrieval-Augmented Generation</em>. arXiv, 2024. <a href="https://arxiv.org/abs/2402.03367">https://arxiv.org/abs/2402.03367</a>: multiple generated queries merged with RRF.</li>
  <li>Thakur, N. et al. <em>BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models</em>. NeurIPS Datasets and Benchmarks, 2021. <a href="https://arxiv.org/abs/2104.08663">https://arxiv.org/abs/2104.08663</a></li>
  <li>Zhang, X. et al. <em>MIRACL: A Multilingual Retrieval Dataset Covering 18 Diverse Languages</em>. TACL, 2023. <a href="https://arxiv.org/abs/2210.09984">https://arxiv.org/abs/2210.09984</a>: a benchmark for retrieval when questions and documents span many languages.</li>
  <li>Lewis, P. et al. <em>Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks</em>. NeurIPS, 2020. <a href="https://arxiv.org/abs/2005.11401">https://arxiv.org/abs/2005.11401</a></li>
</ol>]]></content><author><name>Fadhil Mochammad</name></author><category term="agents" /><summary type="html"><![CDATA[]]></summary></entry><entry><title type="html">Deconstructing PlanOut: a coin toss you can repeat</title><link href="https://fadhilmch.github.io/posts/deconstructing-planout/" rel="alternate" type="text/html" title="Deconstructing PlanOut: a coin toss you can repeat" /><published>2026-09-30T00:00:00+02:00</published><updated>2026-09-30T00:00:00+02:00</updated><id>https://fadhilmch.github.io/posts/deconstructing-planout</id><content type="html" xml:base="https://fadhilmch.github.io/posts/deconstructing-planout/"><![CDATA[<p>Andi opens a checkout page. We want to test two buttons: A is blue, B is pink. He gets blue today. Tomorrow, he opens the page on another server. He should still get blue.</p>

<p>That sounds like a small requirement. It rules out the most obvious implementation: toss a fresh coin on every request. You could save every user’s choice in a database, but then every server needs access to that record. PlanOut offers another way: make the coin toss repeatable from the inputs.</p>

<p>This is the part of experimentation systems I want to take apart here. One shopper, one button, and the actual numbers from the archived Python implementation. No production traffic or company internals: the example below uses invented IDs.</p>

<style>
.prose .fig.planout-fig { max-width: 540px; }
.fig.planout-fig svg { min-width: 340px; max-width: 540px; }
.planout-fig .t { font-size: 14px; }
.planout-fig .m, .planout-fig .ta, .planout-fig .tb { font-size: 12px; }
.planout-fig .h { font-size: 12px; }
</style>

<h2 id="what-planout-was-trying-to-separate">What PlanOut was trying to separate</h2>

<p>The product owns the checkout page. The experiment owns the rule for choosing the button. Those should not become one tangled piece of code.</p>

<p>PlanOut was developed at Facebook as a language and framework for defining online experiments. Its 2014 paper describes how an experiment maps a unit, such as a user ID, to parameter values. A parameter might be a button colour, a piece of text, or a discount. Product code asks for that value and uses it.</p>

<p>In other words, the experiment says “blue”. It does not draw the button, count purchases, or tell us whether blue is better.</p>

<figure class="fig planout-fig">
<svg viewBox="0 0 420 324" role="img" aria-labelledby="po-flow-t po-flow-d">
<title id="po-flow-t">Flow</title>
<desc id="po-flow-d">Andi enters the checkout. The experiment chooses blue. The product renders the button. The outcome pipeline counts purchases separately.</desc>
<rect class="box" x="12" y="12" width="396" height="60" rx="6" /><text class="t" x="24" y="36">Andi opens checkout</text><text class="m" x="24" y="55">unit: user-42</text><path class="ln" d="M210,72 v20" /><rect class="box" x="12" y="92" width="396" height="60" rx="6" /><text class="t" x="24" y="116">Experiment definition</text><text class="m" x="24" y="135">choose button = blue</text><path class="ln" d="M210,152 v20" /><rect class="box" x="12" y="172" width="396" height="60" rx="6" /><text class="t" x="24" y="196">Product code</text><text class="m" x="24" y="215">use blue when rendering checkout</text><path class="ln dash" d="M210,232 v20" /><rect class="box" x="12" y="252" width="396" height="60" rx="6" /><text class="t" x="24" y="276">Outcome pipeline</text><text class="m" x="24" y="295">record purchases; analyse later</text>
</svg>
<figcaption>PlanOut returns parameter values. Rendering and measuring the result live outside that decision.</figcaption>
</figure>

<p>Here is a small definition using PlanOut’s Python API:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">from</span> <span class="n">planout.experiment</span> <span class="kn">import</span> <span class="n">SimpleExperiment</span>
<span class="kn">from</span> <span class="n">planout.ops.random</span> <span class="kn">import</span> <span class="n">UniformChoice</span>

<span class="k">class</span> <span class="nc">CheckoutButton</span><span class="p">(</span><span class="n">SimpleExperiment</span><span class="p">):</span>
    <span class="k">def</span> <span class="nf">setup</span><span class="p">(</span><span class="n">self</span><span class="p">):</span>
        <span class="n">self</span><span class="p">.</span><span class="n">salt</span> <span class="o">=</span> <span class="sh">"</span><span class="s">checkout-v1</span><span class="sh">"</span>

    <span class="k">def</span> <span class="nf">assign</span><span class="p">(</span><span class="n">self</span><span class="p">,</span> <span class="n">params</span><span class="p">,</span> <span class="n">userid</span><span class="p">):</span>
        <span class="n">params</span><span class="p">.</span><span class="n">button</span> <span class="o">=</span> <span class="nc">UniformChoice</span><span class="p">(</span>
            <span class="n">choices</span><span class="o">=</span><span class="p">[</span><span class="sh">"</span><span class="s">blue</span><span class="sh">"</span><span class="p">,</span> <span class="sh">"</span><span class="s">pink</span><span class="sh">"</span><span class="p">],</span> <span class="n">unit</span><span class="o">=</span><span class="n">userid</span>
        <span class="p">)</span>

<span class="n">experiment</span> <span class="o">=</span> <span class="nc">CheckoutButton</span><span class="p">(</span><span class="n">userid</span><span class="o">=</span><span class="sh">"</span><span class="s">user-42</span><span class="sh">"</span><span class="p">)</span>
<span class="n">colour</span> <span class="o">=</span> <span class="n">experiment</span><span class="p">.</span><span class="nf">get</span><span class="p">(</span><span class="sh">"</span><span class="s">button</span><span class="sh">"</span><span class="p">)</span>  <span class="c1"># "blue"
</span></code></pre></div></div>

<p>Andi’s invented ID is <code class="language-plaintext highlighter-rouge">user-42</code>. It is the <em>unit</em>: the thing we randomise. If we used a session ID instead, we’d be randomising sessions, and Andi could get a different button on his next visit. Choosing the unit is a design decision, not a detail to leave to the SDK.</p>

<h2 id="a-coin-toss-without-a-coin">A coin toss without a coin</h2>

<p>A hash function turns an input string into a fixed-length output. The same input always gives the same output. Small changes to the input usually produce very different outputs.</p>

<p>PlanOut uses SHA-1 as a mixing function for assignment. This is not a password scheme or a way to hide someone’s group. We want repeatable, well-spread values, not secrecy.</p>

<p>For this example, PlanOut joins three pieces with dots:</p>

<ol>
  <li>The experiment salt: <code class="language-plaintext highlighter-rouge">checkout-v1</code>.</li>
  <li>The parameter salt: <code class="language-plaintext highlighter-rouge">button</code>, supplied from the assigned variable name.</li>
  <li>The unit: <code class="language-plaintext highlighter-rouge">user-42</code>.</li>
</ol>

<p>The complete key is <code class="language-plaintext highlighter-rouge">checkout-v1.button.user-42</code>. The default separator is a dot; explicit salt overrides and multi-part units have their own paths in the source.</p>

<figure class="fig planout-fig">
<svg viewBox="0 0 420 324" role="img" aria-labelledby="po-hash-t po-hash-d">
<title id="po-hash-t">Hash</title>
<desc id="po-hash-d">The exact key checkout-v1.button.user-42 is hashed. Its first 15 hex digits become integer 235347851596851754. Remainder modulo two is zero, selecting blue.</desc>
<g><rect class="box" x="12" y="12" width="396" height="60" rx="6" /><text class="t" x="24" y="36">checkout-v1.button.user-42</text><text class="m" x="24" y="55">salt + parameter + unit</text></g><g><path class="ln" d="M210,72 v20" /></g><g><rect class="box" x="12" y="92" width="396" height="60" rx="6" /><text class="t" x="24" y="116">SHA-1 → keep first 15 hex digits</text><text class="m" x="24" y="135">3441f9fc514ea2a</text></g><g><path class="ln" d="M210,152 v20" /></g><g><rect class="box" x="12" y="172" width="396" height="60" rx="6" /><text class="t" x="24" y="196">Convert hex to integer</text><text class="m" x="24" y="215">235347851596851754</text></g><g><path class="ln" d="M210,232 v20" /></g><g><rect class="box" x="12" y="252" width="396" height="60" rx="6" /><text class="t" x="24" y="276">h % 2 = 0 → choices[0] → blue</text><text class="m" x="24" y="295">same key, same answer</text></g>
</svg>
<figcaption>The computed path for Andi. Nothing is freshly drawn on the second request.</figcaption>
</figure>

<style>
.po-playground { margin: 32px 0; padding: 20px; border: 1px solid var(--line); border-radius: 8px; background: var(--panel); }
.po-playground[hidden], .po-playground [hidden] { display: none; }
.prose .po-playground h3 { margin: 0 0 8px; }
.prose .po-playground p { margin: 0 0 16px; font-size: 14px; color: var(--muted); }
.po-playground label { display: block; margin-bottom: 6px; font-size: 14px; }
.po-playground input { box-sizing: border-box; width: 100%; min-width: 0; padding: 10px 12px; border: 1px solid var(--line); border-radius: 4px; background: var(--bg); color: var(--fg); font: 14px 'Geist Mono', monospace; }
.po-playground input:focus-visible { outline: 2px solid var(--l0); outline-offset: 3px; }
.prose .po-playground .po-status { margin: 12px 0 0; }
.po-playground dl { margin: 20px 0 0; }
.po-playground .po-row { border-top: 1px solid var(--line); padding: 10px 0; }
.po-playground dt { color: var(--muted); font-size: 12px; }
.po-playground dd { margin: 4px 0 0; overflow-wrap: anywhere; font: 13px/1.6 'Geist Mono', monospace; }
.po-playground dd code { padding: 0; border: 0; background: none; font: inherit; }
.po-playground .po-choice { display: inline-block; margin-top: 8px; padding: 4px 10px; border: 1px solid currentColor; border-radius: 4px; font-weight: 600; }
.po-playground .po-choice[data-choice="blue"] { color: var(--l0); }
.po-playground .po-choice[data-choice="pink"] { color: var(--l3); }
</style>

<section class="po-playground" id="po-playground" aria-labelledby="po-playground-title" hidden="">
<h3 id="po-playground-title">Try the coin toss</h3>
<p>Change the user ID and follow the actual calculation. The salt stays <code>checkout-v1</code>, the parameter stays <code>button</code>. Nothing is sent to a server.</p>
<label for="po-user-id">User ID</label>
<input id="po-user-id" type="text" value="user-42" placeholder="e.g. user-42" autocomplete="off" spellcheck="false" aria-describedby="po-status" />
<p class="po-status" id="po-status" role="status" aria-live="polite" aria-atomic="true">Computing the assignment...</p>
<dl id="po-results" hidden="">
<div class="po-row"><dt>1. Full key: salt.parameter.unit</dt><dd><code id="po-key"></code></dd></div>
<div class="po-row"><dt>2. SHA-1 of the UTF-8 key</dt><dd><code id="po-hash"></code></dd></div>
<div class="po-row"><dt>3. Keep the first 15 hex digits (60 bits)</dt><dd><code id="po-prefix"></code></dd></div>
<div class="po-row"><dt>4. Convert hex to an integer</dt><dd><code id="po-integer"></code></dd></div>
<div class="po-row"><dt>5. Remainder selects a choice</dt><dd><code id="po-modulo"></code><br /><span class="po-choice" id="po-choice"></span></dd></div>
</dl>
</section>
<noscript><p>The example above works without JavaScript. Enable JavaScript to try other user IDs.</p></noscript>
<script>
(() => {
  'use strict';
  const widget = document.getElementById('po-playground');
  const input = document.getElementById('po-user-id');
  const status = document.getElementById('po-status');
  const results = document.getElementById('po-results');
  // SHA-1 is PlanOut's assignment mixer here, not a security primitive.
  if (!window.crypto || !window.crypto.subtle || typeof TextEncoder === 'undefined' || typeof BigInt === 'undefined') {
    const note = document.createElement('p');
    note.textContent = 'The interactive example needs a browser with Web Crypto and BigInt on HTTPS. The static calculation above still applies.';
    widget.replaceWith(note);
    return;
  }
  widget.hidden = false;
  let revision = 0;
  async function update() {
    const current = ++revision;
    const unit = input.value; // Preserve the exact ID, including spaces and Unicode.
    results.hidden = true;
    if (unit === '') {
      status.textContent = 'Enter a user ID to compute its assignment.';
      return;
    }
    status.textContent = 'Computing the assignment...';
    try {
      const key = 'checkout-v1.button.' + unit;
      const bytes = new TextEncoder().encode(key);
      const digest = await window.crypto.subtle.digest('SHA-1', bytes);
      if (current !== revision) return; // An older digest must not overwrite a newer ID.
      const hash = Array.from(new Uint8Array(digest), byte => byte.toString(16).padStart(2, '0')).join('');
      const prefix = hash.slice(0, 15);
      // Number would lose precision for a 60-bit hash. Keep this exact with BigInt.
      const integer = BigInt('0x' + prefix);
      const remainder = integer % BigInt(2);
      const choice = ['blue', 'pink'][Number(remainder)];
      document.getElementById('po-key').textContent = key;
      document.getElementById('po-hash').textContent = hash;
      document.getElementById('po-prefix').textContent = prefix;
      document.getElementById('po-integer').textContent = integer.toString();
      document.getElementById('po-modulo').textContent = 'h % 2 = ' + remainder + ' → choices[' + remainder + ']';
      const badge = document.getElementById('po-choice');
      badge.textContent = choice;
      badge.dataset.choice = choice;
      results.hidden = false;
      status.textContent = 'Assigned to ' + choice + '. Same ID, same answer.';
    } catch (error) {
      if (current !== revision) return;
      status.textContent = 'This browser could not compute SHA-1. The static example above is still available.';
    }
  }
  input.addEventListener('input', update);
  update();
})();
</script>

<p>SHA-1 produces 40 hexadecimal digits. The Python implementation keeps the first 15. Each hex digit holds four bits, so that prefix gives a 60-bit integer. For Andi:</p>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>key:          checkout-v1.button.user-42
SHA-1:        3441f9fc514ea2ac8f04b1c44a64d90aa350f4fb
first 15:     3441f9fc514ea2a
integer h:    235347851596851754
h % 2:        0
choices[0]:   blue
</code></pre></div></div>

<p><code class="language-plaintext highlighter-rouge">%</code> means remainder. Dividing an even number by 2 leaves remainder 0; an odd number leaves 1. With two choices, that is the whole <code class="language-plaintext highlighter-rouge">UniformChoice</code> decision: <code class="language-plaintext highlighter-rouge">choices[h % 2]</code>.</p>

<p>A thousand servers can repeat it. None needs a stored row saying “Andi got blue”, as long as they agree on the input and implementation.</p>

<h2 id="the-bucket-picture-with-one-important-distinction">The bucket picture, with one important distinction</h2>

<p>It is tempting to explain every assignment as “put the hash on a line from zero to one, then split the line in half”. That is a useful picture for weighted choices. It is not the exact code path for <code class="language-plaintext highlighter-rouge">UniformChoice</code>.</p>

<p>The archived Python implementation has two different operations:</p>

<ul>
  <li><code class="language-plaintext highlighter-rouge">UniformChoice</code>: use the integer’s remainder to select a list position.</li>
  <li><code class="language-plaintext highlighter-rouge">WeightedChoice</code>: divide the integer by a scale, then compare it with cumulative weights.</li>
</ul>

<p>For the second operation, the scale is <code class="language-plaintext highlighter-rouge">2**60 - 1</code>, or <code class="language-plaintext highlighter-rouge">1,152,921,504,606,846,975</code>. Andi’s normalized value is about <code class="language-plaintext highlighter-rouge">0.2041317216</code>. With weights <code class="language-plaintext highlighter-rouge">[0.5, 0.5]</code>, it lands in blue’s half of the line.</p>

<figure class="fig planout-fig">
<svg viewBox="0 0 420 296" role="img" aria-labelledby="po-bucket-t po-bucket-d">
<title id="po-bucket-t">Bucket</title>
<desc id="po-bucket-d">UniformChoice selects blue using remainder zero. WeightedChoice instead places normalized value 0.2041 in the first half of a line from zero to one.</desc>
<text class="h" x="12" y="22">UNIFORMCHOICE: LIST INDEX</text><rect class="box" x="12" y="38" width="190" height="60" rx="6" /><text class="t" x="24" y="62">0 → blue</text><text class="m" x="24" y="81">even hash</text><rect class="box" x="218" y="38" width="190" height="60" rx="6" /><text class="t" x="230" y="62">1 → pink</text><text class="m" x="230" y="81">odd hash</text><circle class="fa" cx="107" cy="118" r="6" /><text class="ta" x="107" y="145" text-anchor="middle">Andi: remainder 0</text><text class="h" x="12" y="188">WEIGHTEDCHOICE: CUMULATIVE WEIGHTS</text><rect class="fa" x="12" y="205" width="198" height="32" opacity=".5" /><rect class="fb" x="210" y="205" width="198" height="32" opacity=".5" /><text class="t" x="111" y="226" text-anchor="middle">blue</text><text class="t" x="309" y="226" text-anchor="middle">pink</text><text class="m" x="12" y="258">0</text><text class="m" x="210" y="258" text-anchor="middle">0.5</text><text class="m" x="408" y="258" text-anchor="end">1</text><path class="sa" d="M93,200 v44" /><text class="ta" x="93" y="280" text-anchor="middle">Andi: 0.2041</text>
</svg>
<figcaption>Two different operators. The example uses the same blue/pink order and equal weights, but matching probabilities do not guarantee matching users.</figcaption>
</figure>

<p>Both operations happen to give Andi blue here. That does not mean they assign every user identically. Replacing one with the other can move users even if both advertise a 50/50 split.</p>

<p>A small source-level detail matters if you want compatibility: PlanOut converts that scale to a floating-point value, and uses <code class="language-plaintext highlighter-rouge">&lt;=</code> at the weighted boundaries. The mathematical normalization includes 1 at the maximum hash. Replacing the denominator with <code class="language-plaintext highlighter-rouge">2**60</code>, changing the comparison, or switching to another port without checking its hash behaviour is not a harmless cleanup.</p>

<h2 id="a-salt-is-the-name-of-the-shuffle">A salt is the name of the shuffle</h2>

<p>Think of the salt as the label on a shuffled deck. With the same deck label and the same user ID, draw the same card again. Change the label and you get a new shuffle.</p>

<p>Here are computed results for 10,000 invented users, <code class="language-plaintext highlighter-rouge">user-0</code> through <code class="language-plaintext highlighter-rouge">user-9999</code>, using the same key construction and <code class="language-plaintext highlighter-rouge">UniformChoice</code> modulo rule:</p>

<table>
  <thead>
    <tr>
      <th>Change</th>
      <th style="text-align: right">Users whose button changed</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Repeat <code class="language-plaintext highlighter-rouge">checkout-v1</code>, same choices</td>
      <td style="text-align: right">0 / 10,000</td>
    </tr>
    <tr>
      <td>Change salt to <code class="language-plaintext highlighter-rouge">checkout-v2</code></td>
      <td style="text-align: right">4,984 / 10,000</td>
    </tr>
    <tr>
      <td>Keep salt, reverse <code class="language-plaintext highlighter-rouge">[blue, pink]</code></td>
      <td style="text-align: right">10,000 / 10,000</td>
    </tr>
  </tbody>
</table>

<p>The original split was 5,004 blue and 4,996 pink. These are reproducible results for this toy population, not a promise that every experiment produces those exact counts.</p>

<figure class="fig planout-fig">
<svg viewBox="0 0 420 306" role="img" aria-labelledby="po-salt-t po-salt-d">
<title id="po-salt-t">Salt</title>
<desc id="po-salt-d">Among ten thousand users, repeating the definition changes zero assignments, changing the experiment salt changes 4984, and reversing the choices changes all ten thousand.</desc>
<text class="h" x="12" y="20">USERS CHANGED / 10,000</text><text class="t" x="12" y="52">Same salt, same choices</text><rect class="box" x="12" y="64" width="396" height="20" rx="4" /><text class="m" x="408" y="104" text-anchor="end">0 / 10,000</text><text class="t" x="12" y="142">New salt</text><rect class="box" x="12" y="154" width="396" height="20" rx="4" /><rect class="fa" x="12" y="154" width="197.3664" height="20" rx="4" /><text class="m" x="408" y="194" text-anchor="end">4,984 / 10,000</text><text class="t" x="12" y="232">Reversed choices</text><rect class="box" x="12" y="244" width="396" height="20" rx="4" /><rect class="fb" x="12" y="244" width="396.0" height="20" rx="4" /><text class="m" x="408" y="284" text-anchor="end">10,000 / 10,000</text>
</svg>
<figcaption>Computed on invented user-0 through user-9999. Bars share the same scale; the input change, not new coin tosses, moves users.</figcaption>
</figure>

<style>
.po-lab { margin:32px 0; padding:20px; border:1px solid var(--line); border-radius:8px; background:var(--panel); }
.po-lab[hidden] { display:none; }
.prose .po-lab h3 { margin:0 0 8px; }
.prose .po-lab p { margin:0 0 12px; font-size:14px; color:var(--muted); }
.po-lab button { padding:10px 14px; border:1px solid var(--line); border-radius:4px; background:var(--bg); color:var(--fg); font:inherit; cursor:pointer; min-height:44px; }
.po-lab button:focus-visible { outline:2px solid var(--l0); outline-offset:3px; }
.po-lab button:disabled { opacity:.6; cursor:wait; }
.po-lab code { overflow-wrap:anywhere; }
.po-lab .po-tally { display:flex; justify-content:space-between; gap:12px; margin:16px 0 8px; font:13px/1.5 'Geist Mono',monospace; }
.po-blue { color:var(--l0); }
.po-pink { color:var(--l3); }
.po-dot-field { position:relative; height:204px; border:1px solid var(--line); border-radius:4px; background:linear-gradient(to right,transparent calc(50% - 1px),var(--line) 50%,transparent calc(50% + 1px)); overflow:hidden; }
.po-dot { position:absolute; left:0; top:0; width:6px; height:6px; border-radius:50%; transition:transform .55s ease,background-color .55s ease; }
.po-dot[data-choice="0"] { background:var(--l0); }
.po-dot[data-choice="1"] { background:var(--l3); }
.po-split-track { height:24px; display:flex; border:1px solid var(--line); border-radius:4px; overflow:hidden; background:var(--bg); }
.po-split-track span { display:block; }
.po-split-track .po-blue-fill { background:var(--l0); }
.po-split-track .po-pink-fill { background:var(--l3); }
.po-lab progress { width:100%; height:8px; accent-color:var(--l0); }
.prose .po-lab .po-lab-status { margin-top:12px; margin-bottom:0; }
@media(prefers-reduced-motion:reduce) { .po-dot { transition:none; } }
@media(max-width:420px) { .po-lab { padding:14px; } }
</style>

<section class="po-lab" id="po-salt-shifter" aria-labelledby="po-salt-title" hidden="">
<h3 id="po-salt-title">Change the shuffle, not the users</h3>
<p>These are the same 200 IDs, <code>user-0</code> through <code>user-199</code>. Each dot uses SHA-1 and the first 15 hex digits, then <code>h % 2</code>. Change only the salt and see which users cross over.</p>
<p>Salt: <code id="po-current-salt">checkout-v1</code> · parameter: <code>button</code></p>
<button type="button" id="po-change-salt">Change salt</button>
<div class="po-tally"><span class="po-blue" id="po-salt-blue">Blue: 0</span><span class="po-pink" id="po-salt-pink">Pink: 0</span></div>
<div class="po-dot-field" id="po-dot-field" aria-hidden="true"></div>
<p class="po-lab-status" id="po-salt-status" role="status" aria-live="polite">Computing 200 assignments...</p>
</section>
<noscript><p>The static salt comparison above shows the same principle. Enable JavaScript to reshuffle 200 users yourself.</p></noscript>

<p>For Andi, the new salt produces prefix <code class="language-plaintext highlighter-rouge">5c87192d9dc8829</code>, integer <code class="language-plaintext highlighter-rouge">416707841066043433</code>, and remainder 1. He moves to pink. Reversing the original choices also gives him pink, but for a different reason: his original index is still 0; index 0 now means pink.</p>

<p>That last row is the trap. Keeping the hash stable is not enough. Keep the meaning of its output stable too.</p>

<p>A new experiment can use a new salt deliberately. A live experiment should not casually rename its salt, change ID formatting, reorder choices, edit weights, or switch hashing algorithms. Any of those can change the experience for people already enrolled. Stateless repeatability is not the same thing as a saved sticky assignment that survives definition changes.</p>

<p>The parameter name acts as a salt too. <code class="language-plaintext highlighter-rouge">button</code> and <code class="language-plaintext highlighter-rouge">hint</code> get different hash keys for the same user, unless you intentionally share an explicit salt. This keeps two parameters from being perfectly tied to each other by accident.</p>

<h2 id="namespaces-pick-the-experiment-before-the-button">Namespaces: pick the experiment before the button</h2>

<p>Now suppose two teams both want to test the checkout page. One tests the button; the other tests the layout. We do not want Andi in both at the same time.</p>

<p>A namespace is a shared allocation space. In the Python implementation, a user hashes into an integer segment. Experiments reserve disjoint sets of segment IDs. The user’s segment decides which experiment, if any, runs. Only then does that experiment choose its parameter values.</p>

<figure class="fig planout-fig">
<svg viewBox="0 0 420 326" role="img" aria-labelledby="po-namespace-t po-namespace-d">
<title id="po-namespace-t">Namespace</title>
<desc id="po-namespace-d">Illustrative noncontiguous slots: button experiment owns 0, 3, 7; layout owns 1, 5. Other slots give defaults. A user first routes to one experiment, then gets its parameters.</desc>
<text class="h" x="12" y="20">ONE NAMESPACE · 10 ILLUSTRATIVE SLOTS</text><rect class="fa" x="12" y="36" width="72" height="48" rx="5" opacity=".65" /><text class="t" x="48" y="64" text-anchor="middle">0</text><rect class="fb" x="92" y="36" width="72" height="48" rx="5" opacity=".65" /><text class="t" x="128" y="64" text-anchor="middle">1</text><rect class="box" x="172" y="36" width="72" height="48" rx="5" opacity=".65" /><text class="t" x="208" y="64" text-anchor="middle">2</text><rect class="fa" x="252" y="36" width="72" height="48" rx="5" opacity=".65" /><text class="t" x="288" y="64" text-anchor="middle">3</text><rect class="box" x="332" y="36" width="72" height="48" rx="5" opacity=".65" /><text class="t" x="368" y="64" text-anchor="middle">4</text><rect class="fb" x="12" y="100" width="72" height="48" rx="5" opacity=".65" /><text class="t" x="48" y="128" text-anchor="middle">5</text><rect class="box" x="92" y="100" width="72" height="48" rx="5" opacity=".65" /><text class="t" x="128" y="128" text-anchor="middle">6</text><rect class="fa" x="172" y="100" width="72" height="48" rx="5" opacity=".65" /><text class="t" x="208" y="128" text-anchor="middle">7</text><rect class="box" x="252" y="100" width="72" height="48" rx="5" opacity=".65" /><text class="t" x="288" y="128" text-anchor="middle">8</text><rect class="box" x="332" y="100" width="72" height="48" rx="5" opacity=".65" /><text class="t" x="368" y="128" text-anchor="middle">9</text><text class="ta" x="12" y="186">Button: slots 0, 3, 7</text><text class="tb" x="12" y="208">Layout: slots 1, 5</text><text class="m" x="12" y="230">Other slots: default experience</text><rect class="box" x="12" y="250" width="396" height="64" rx="6" /><text class="t" x="24" y="274">User → segment → one experiment</text><text class="m" x="24" y="293">then: experiment → button value</text>
</svg>
<figcaption>Two decisions: namespace routing, then treatment assignment. Separate namespaces can still overlap in people.</figcaption>
</figure>

<p>The numbered tiles above are an illustration, not Andi’s computed upstream namespace segment. The real <code class="language-plaintext highlighter-rouge">add_experiment()</code> samples from the available segments, so an experiment’s allocation need not be a neat contiguous range.</p>

<p>Mutual exclusion applies inside that namespace: a segment cannot belong to both checkout experiments at once. A separate namespace for the homepage can also enrol Andi. Namespaces do not automatically prevent interactions between experiments on different surfaces.</p>

<p>The namespace name, unit, segment count and allocation history are part of the routing contract. Preserving only the button’s salt will not preserve routing if those change.</p>

<h2 id="assignment-is-not-exposure">Assignment is not exposure</h2>

<p>We have computed a button colour. Has Andi seen the button? Not necessarily. The page might fail before rendering it. The application might ask for the colour on a route where checkout is never shown.</p>

<p>PlanOut gives you a logging hook. In the default Python experiment behaviour, the first eligible <code class="language-plaintext highlighter-rouge">get()</code> on an instance triggers exposure logging before returning the parameter. The instance remembers that it logged, so repeated <code class="language-plaintext highlighter-rouge">get()</code> calls on that same instance do not each emit a new automatic exposure. A new instance is a new situation; there is no global per-user deduplication promise.</p>

<figure class="fig planout-fig">
<svg viewBox="0 0 420 324" role="img" aria-labelledby="po-logging-t po-logging-d">
<title id="po-logging-t">Logging</title>
<desc id="po-logging-d">Computing assignment is followed by first eligible get, which triggers a log hook. Rendering comes later and can fail. A purchase is another event.</desc>
<rect class="box" x="12" y="12" width="396" height="60" rx="6" /><text class="t" x="24" y="36">Compute assignment</text><text class="m" x="24" y="55">blue is the intended value</text><path class="ln" d="M210,72 v20" /><rect class="box" x="12" y="92" width="396" height="60" rx="6" /><text class="t" x="24" y="116">First eligible get("button")</text><text class="m" x="24" y="135">automatic log hook on this instance</text><path class="ln dash" d="M210,152 v20" /><rect class="box" x="12" y="172" width="396" height="60" rx="6" /><text class="t" x="24" y="196">Render button</text><text class="m" x="24" y="215">may fail; get() is not viewability</text><path class="ln dash" d="M210,232 v20" /><rect class="box" x="12" y="252" width="396" height="60" rx="6" /><text class="t" x="24" y="276">Purchase event</text><text class="m" x="24" y="295">outcome, not assignment</text>
</svg>
<figcaption>A parameter-access log is useful evidence, but it is not proof that a person saw the pixels.</figcaption>
</figure>

<p>That hook records parameter access, not human viewability. Automatic logging can be disabled, and an experiment marked out of experiment does not log an exposure. The base experiment delegates the log to your implementation; <code class="language-plaintext highlighter-rouge">SimpleExperiment</code> writes JSON to a file. Getting a production event pipeline right is still your job.</p>

<p>A useful record connects the unit, experiment, parameter values and time to the definition that produced them. Upstream also records a checksum when available. In a production system I would make the definition version, override, surface and event ID explicit, then decide where the application should emit the event that analysis treats as exposure.</p>

<p>A hash can reconstruct an intended choice if the original inputs and definition survive. It cannot reconstruct a missing render event or a fallback that nobody recorded.</p>

<h2 id="what-the-hash-cannot-tell-you">What the hash cannot tell you</h2>

<section class="po-lab" id="po-split-simulator" aria-labelledby="po-split-title" hidden="">
<h3 id="po-split-title">Run 10,000 repeatable coin tosses</h3>
<p>Assign <code>user-0</code> through <code>user-9999</code> with <code>checkout-v1.button</code> and <code>h % 2</code>. A 50/50 rule does not promise exactly 5,000 in each group. Run it again: the same inputs give the same counts.</p>
<button type="button" id="po-run-split">Run 10,000 users</button>
<div class="po-tally"><span class="po-blue" id="po-split-blue">Blue: 0</span><span class="po-pink" id="po-split-pink">Pink: 0</span></div>
<div class="po-split-track" aria-hidden="true"><span class="po-blue-fill" id="po-blue-fill" style="width:0%"></span><span class="po-pink-fill" id="po-pink-fill" style="width:0%"></span></div>
<label for="po-split-progress">Users assigned</label>
<progress id="po-split-progress" max="10000" value="0">0 / 10000</progress>
<p class="po-lab-status" id="po-split-status" role="status" aria-live="polite">Ready. All computation stays in your browser.</p>
</section>
<noscript><p>Enable JavaScript to run the 10,000-user assignment example. The SRM explanation below works without it.</p></noscript>
<script>
(() => {
  'use strict';
  const saltWidget = document.getElementById('po-salt-shifter');
  const splitWidget = document.getElementById('po-split-simulator');
  if (!window.crypto || !window.crypto.subtle || typeof TextEncoder === 'undefined' || typeof BigInt === 'undefined') {
    [saltWidget, splitWidget].forEach(widget => {
      const note = document.createElement('p');
      note.textContent = 'This interactive example needs Web Crypto and BigInt on HTTPS. The static examples still apply.';
      widget.replaceWith(note);
    });
    return;
  }
  saltWidget.hidden = splitWidget.hidden = false;
  const encoder = new TextEncoder();
  // PlanOut's SHA-1 mixer, not a cryptographic security feature. Never round the 60-bit integer.
  async function bucket(salt, unit) {
    const digest = await crypto.subtle.digest('SHA-1', encoder.encode(salt + '.button.' + unit));
    const hex = Array.from(new Uint8Array(digest), byte => byte.toString(16).padStart(2, '0')).join('');
    return Number(BigInt('0x' + hex.slice(0, 15)) % BigInt(2));
  }
  const field = document.getElementById('po-dot-field');
  const saltButton = document.getElementById('po-change-salt');
  const saltStatus = document.getElementById('po-salt-status');
  const dots = Array.from({length:200}, (_, i) => {
    const dot = document.createElement('span');
    dot.className = 'po-dot';
    dot.title = 'user-' + i;
    field.appendChild(dot);
    return dot;
  });
  let saltVersion = 1;
  let assignments = [];
  function drawDots() {
    if (!assignments.length) return;
    const half = field.clientWidth / 2;
    const counts = [0, 0];
    assignments.forEach((choice, i) => {
      const index = counts[choice]++;
      const x = choice * half + 10 + (index % 8) * ((half - 22) / 8);
      const y = 12 + Math.floor(index / 8) * 12;
      dots[i].dataset.choice = String(choice);
      dots[i].style.transform = 'translate(' + x + 'px,' + y + 'px)';
    });
    field.style.height = (Math.ceil(Math.max(...counts) / 8) * 12 + 24) + 'px';
    document.getElementById('po-salt-blue').textContent = 'Blue: ' + counts[0];
    document.getElementById('po-salt-pink').textContent = 'Pink: ' + counts[1];
  }
  async function reshuffle() {
    saltButton.disabled = true;
    saltStatus.textContent = 'Computing 200 assignments...';
    const salt = 'checkout-v' + saltVersion;
    try {
      const next = await Promise.all(dots.map((_, i) => bucket(salt, 'user-' + i)));
      const changed = next.filter((choice, i) => assignments.length && choice !== assignments[i]).length;
      const initial = assignments.length === 0;
      assignments = next;
      document.getElementById('po-current-salt').textContent = salt;
      drawDots();
      saltStatus.textContent = initial ? '200 users assigned. Change salt to reshuffle them.' : changed + ' of 200 users changed sides. Same users; new salt.';
    } catch (error) {
      saltStatus.textContent = 'SHA-1 could not be computed. The static salt comparison is still available.';
    } finally { saltButton.disabled = false; }
  }
  saltButton.addEventListener('click', () => { saltVersion++; reshuffle(); });
  if (typeof ResizeObserver !== 'undefined') new ResizeObserver(drawDots).observe(field);
  else window.addEventListener('resize', drawDots);
  reshuffle();
  const runButton = document.getElementById('po-run-split');
  const splitStatus = document.getElementById('po-split-status');
  runButton.addEventListener('click', async () => {
    runButton.disabled = true;
    const counts = [0, 0];
    document.getElementById('po-split-progress').value = 0;
    document.getElementById('po-blue-fill').style.width = '0%';
    document.getElementById('po-pink-fill').style.width = '0%';
    document.getElementById('po-split-blue').textContent = 'Blue: 0';
    document.getElementById('po-split-pink').textContent = 'Pink: 0';
    splitStatus.textContent = 'Computing 10,000 real SHA-1 assignments...';
    try {
      // Bounded batches keep the page responsive and let actual progress paint.
      for (let start = 0; start < 10000; start += 250) {
        const batch = await Promise.all(Array.from({length:250}, (_, offset) => bucket('checkout-v1', 'user-' + (start + offset))));
        batch.forEach(choice => counts[choice]++);
        const completed = start + 250;
        document.getElementById('po-split-blue').textContent = 'Blue: ' + counts[0].toLocaleString();
        document.getElementById('po-split-pink').textContent = 'Pink: ' + counts[1].toLocaleString();
        document.getElementById('po-blue-fill').style.width = (counts[0] / 10000 * 100) + '%';
        document.getElementById('po-pink-fill').style.width = (counts[1] / 10000 * 100) + '%';
        document.getElementById('po-split-progress').value = completed;
        await new Promise(resolve => requestAnimationFrame(resolve));
      }
      splitStatus.textContent = '10,000 assigned: ' + (counts[0] / 100).toFixed(2) + '% blue, ' + (counts[1] / 100).toFixed(2) + '% pink. Repeat to get exactly the same counts. This is assignment, not exposure data or an SRM test.';
    } catch (error) {
      splitStatus.textContent = 'The computation stopped before finishing. Try again; the static explanation remains available.';
    } finally { runButton.disabled = false; }
  });
})();
</script>

<p>Suppose the button experiment expects an even split, but the logged population contains 700 blue and 300 pink users. That is a sample ratio mismatch, or SRM: the observed group counts are unexpectedly far from the planned allocation.</p>

<p>It is a warning to investigate the experiment, not evidence that pink loses. Maybe assignment changed. Maybe one group has missing events. Maybe the eligibility filter is different on the two paths. Check the population you count, the planned allocation, and the logging before reading an effect estimate. A fixed 50/50 check also does not fit a bandit that intentionally moves traffic.</p>

<p>PlanOut does not perform that investigation or estimate lift for you. Repeatable assignment is one part of trustworthy experimentation. Exposure data, outcome definitions and statistical analysis are separate parts.</p>

<h2 id="archived-does-not-mean-the-idea-failed">Archived does not mean the idea failed</h2>

<p>The public repository was archived on June 1, 2021 and is read-only. That is verifiable. An exact public explanation from the maintainers for <em>why</em> it was archived is not established in the sources linked below.</p>

<p>So I would not turn the archive badge into a story about Facebook abandoning experimentation, maintenance costs, or a single successor replacing PlanOut. Its own documentation described the interpreter as one way to define experiments at Facebook, running on top of QuickExperiment. The public library was never the whole internal platform.</p>

<p>What has changed for someone choosing a tool today is the available scope. A language gives you definitions and operators. Modern experimentation products can also provide configuration delivery, allocation controls, exposure handling, metric definitions and analysis. Those are reasons to evaluate a broader platform, not proof of the archive’s motive.</p>

<h2 id="what-i-would-use-instead-today">What I would use instead today</h2>

<p>There is no official replacement established by the PlanOut sources. The replacement depends on which job you need done.</p>

<ul>
  <li>For owning the assignment implementation, a small SDK can work, but stable identity, versioned definitions and compatibility tests become your responsibility. GrowthBook’s <a href="https://docs.growthbook.io/lib/build-your-own">build-your-own guide</a> is a useful current example of an SDK contract.</li>
  <li>For an experimentation platform, <a href="https://docs.growthbook.io/app/sticky-bucketing">GrowthBook</a> and <a href="https://docs.statsig.com/experiments-plus">Statsig</a> are examples worth evaluating. Compare their allocation, logging, warehouse and analysis paths against your requirements. They are alternatives, not drop-in replicas of PlanOut’s byte-level behaviour.</li>
  <li>For adaptive allocation, the policy is another layer. A bandit updates traffic based on rewards; a deterministic hash does not learn which button sells more. I covered that distinction in <a href="/posts/thompson-sampling-explained-with-coupons/">Thompson sampling, explained with coupons</a>.</li>
</ul>

<p>If I were moving an existing experiment, I would first make a test set of IDs, salts, definitions, overrides and expected outputs. Run old and new systems against it. Explain every disagreement before moving live traffic. A new SDK producing “roughly half in each group” is not enough if it puts different people in those groups.</p>

<p>The lesson I keep from PlanOut is smaller than a whole platform: separate the experiment definition from product code, make assignment repeatable, and make the boundary between a decision and an observation explicit. Andi getting blue twice is easy to demonstrate. Knowing whether blue helped him is a different system.</p>

<h2 id="references">References</h2>

<ul>
  <li>Bakshy, Eckles and Bernstein, <a href="https://hci.stanford.edu/publications/2014/planout/planout-www2014.pdf">Designing and Deploying Online Field Experiments</a>, WWW 2014. The design motivation and system boundaries.</li>
  <li><a href="https://github.com/facebookarchive/planout">PlanOut repository</a> and <a href="https://github.com/facebookarchive/planout/releases">release page</a>. Public archival status and date.</li>
  <li>Python source: <a href="https://github.com/facebookarchive/planout/blob/master/python/planout/ops/random.py">random operators</a>, <a href="https://github.com/facebookarchive/planout/blob/master/python/planout/assignment.py">assignment</a>, <a href="https://github.com/facebookarchive/planout/blob/master/python/planout/namespace.py">namespaces</a>, and <a href="https://github.com/facebookarchive/planout/blob/master/python/planout/experiment.py">experiment logging</a>. The mechanics here refer to this implementation, not every language port.</li>
  <li>Fabijan et al., <a href="https://www.microsoft.com/en-us/research/publication/diagnosing-sample-ratio-mismatch-in-online-controlled-experiments-a-taxonomy-and-rules-of-thumb-for-practitioners/">Diagnosing Sample Ratio Mismatch</a>, KDD 2019. Why the count mismatch is a diagnostic warning, not a treatment-effect result.</li>
  <li><a href="https://github.com/facebook/planout/blob/master/python/docs/08-about-planout.md">About PlanOut</a>. Its relationship to QuickExperiment.</li>
  <li><a href="https://docs.growthbook.io/lib/build-your-own">GrowthBook SDK guide</a>, <a href="https://docs.growthbook.io/app/sticky-bucketing">sticky bucketing</a>, and <a href="https://docs.statsig.com/experiments-plus">Statsig Experiments</a>. Examples of current alternatives, not evidence of an official succession.</li>
</ul>]]></content><author><name>Fadhil Mochammad</name></author><category term="exp" /><category term="systems" /><summary type="html"><![CDATA[Andi opens a checkout page. We want to test two buttons: A is blue, B is pink. He gets blue today. Tomorrow, he opens the page on another server. He should still get blue.]]></summary></entry><entry><title type="html">Evaluating an agent honestly</title><link href="https://fadhilmch.github.io/posts/evaluating-an-agent-honestly/" rel="alternate" type="text/html" title="Evaluating an agent honestly" /><published>2026-09-10T00:00:00+02:00</published><updated>2026-09-10T00:00:00+02:00</updated><id>https://fadhilmch.github.io/posts/evaluating-an-agent-honestly</id><content type="html" xml:base="https://fadhilmch.github.io/posts/evaluating-an-agent-honestly/"><![CDATA[<p>A single score cannot tell you whether an agent is useful. I work on a production assistant that answers questions by searching a body of documents and writing a reply from what it finds, a pattern usually called retrieval-augmented generation, or RAG. It’s agentic, meaning a model decides which steps to take, so there are more places for it to go wrong than in a fixed pipeline. This post is about how I evaluate it, and about one decision that surprises people: the evaluation produces a report for engineers, not an automatic pass or fail on deployment.</p>

<h2 id="why-it-looks-better-stops-working">Why “it looks better” stops working</h2>

<p>In the first weeks of building an agent you can iterate by feel. You try a handful of questions, read the answers, change a prompt, try again. That’s a perfectly good way to start.</p>

<p>It breaks down once the system has some reach. The assistant I work on serves several markets, takes questions in any language and mixes global and region-specific knowledge. A change that improves one language can quietly degrade another. Nobody can read enough answers by hand to notice. And because the output is fluent, a worse answer doesn’t look worse.</p>

<p>So I treat evaluation as engineering evidence. Compare a change with a baseline, look at the failures one by one, and only then decide what to change next. This is what people mean by <strong>evaluation-driven development</strong>: the eval isn’t a check at the end, it’s what steers the design.</p>

<h2 id="four-metrics-four-different-failures">Four metrics, four different failures</h2>

<p>I measure four things: answer relevance, retrieval recall, groundedness and latency. I want to be specific about why those four, because each one exists to catch something the others miss.</p>

<ul>
  <li><strong>Answer relevance:</strong> does the answer address what the user asked?</li>
  <li><strong>Retrieval recall:</strong> did the search find the documents that contain the answer? This one needs a reference, the source a good answer should come from, which is why my golden examples record one.</li>
  <li><strong>Groundedness:</strong> is every claim in the answer supported by what was retrieved? An answer can be relevant and still invent a detail.</li>
  <li><strong>Latency:</strong> how long did the user wait? A correct answer that arrives too late is, for most purposes, not a good one.</li>
</ul>

<figure class="fig">
<svg viewBox="0 0 680 290" role="img" aria-labelledby="ev1t ev1d">
  <title id="ev1t">Which metric catches which failure</title>
  <desc id="ev1d">A matrix of six failure cases against four metrics. Ignoring the question and off-topic answers are caught by relevance. A missing document is caught by retrieval recall. Unsupported claims are caught by groundedness. A slow answer is caught by latency. A fluent, grounded answer built on a stale source is caught by none of the four.</desc>
  <text class="h" x="10" y="18">FAILURE CASE</text>
  <text class="h" x="395" y="18" text-anchor="middle">RELEVANCE</text>
  <text class="h" x="470" y="18" text-anchor="middle">RECALL</text>
  <text class="h" x="550" y="18" text-anchor="middle">GROUNDED</text>
  <text class="h" x="630" y="18" text-anchor="middle">LATENCY</text>
  <line class="grid" x1="10" x2="670" y1="28" y2="28" />
  <text class="t" x="10" y="52">Answer ignores the question</text>
  <line class="grid" x1="10" x2="670" y1="66" y2="66" />
  <g tabindex="0"><title>Answer ignores the question: caught by relevance</title><circle class="fa" cx="395" cy="48" r="6" /></g>
  <text class="t" x="10" y="90">Right document never retrieved</text>
  <line class="grid" x1="10" x2="670" y1="104" y2="104" />
  <g tabindex="0"><title>Right document never retrieved: caught by retrieval recall</title><circle class="fa" cx="470" cy="86" r="6" /></g>
  <text class="t" x="10" y="128">Answer claims more than sources say</text>
  <line class="grid" x1="10" x2="670" y1="142" y2="142" />
  <g tabindex="0"><title>Answer claims more than sources say: caught by groundedness</title><circle class="fa" cx="550" cy="124" r="6" /></g>
  <text class="t" x="10" y="166">Faithful to sources, but off-topic</text>
  <line class="grid" x1="10" x2="670" y1="180" y2="180" />
  <g tabindex="0"><title>Faithful to sources, but off-topic: caught by relevance</title><circle class="fa" cx="395" cy="162" r="6" /></g>
  <text class="t" x="10" y="204">Good answer, arrives far too late</text>
  <line class="grid" x1="10" x2="670" y1="218" y2="218" />
  <g tabindex="0"><title>Good answer, arrives far too late: caught by latency</title><circle class="fa" cx="630" cy="200" r="6" /></g>
  <text class="t" x="10" y="242">Fluent and grounded, but source is stale</text>
  <line class="grid" x1="10" x2="670" y1="256" y2="256" />
  <g tabindex="0"><title>Fluent and grounded, but source is stale: none of the four metrics flags it</title><text class="tb" x="512" y="242" text-anchor="middle">none of the four</text></g>
</svg>
<figcaption>Each metric covers a different failure, and each can look healthy while another one is failing. The last row is the reminder that a passing dashboard is not the same as a good answer.</figcaption>
</figure>

<p>Two rows in that matrix deserve a comment. “Faithful to sources, but off-topic” shows why groundedness alone is not enough: an answer can quote its sources perfectly and still not answer the question. And the last row shows the limit of all of them. If the source document is out of date, the assistant can be relevant, well retrieved and faithfully grounded in something that is no longer true. No automated metric flags that. A person reading the failures does, or a content review does.</p>

<p>This is why I don’t collapse the four into one number. An average lets a good latency score hide a poor recall score. Reading them side by side tells you which part of the system to look at.</p>

<h2 id="two-kinds-of-evaluation-two-jobs">Two kinds of evaluation, two jobs</h2>

<p>There are two evaluations in the loop, and they answer different questions.</p>

<p><strong>Offline evaluation</strong> runs on demand against a fixed set of questions. In my case it runs whenever there’s an architecture change or a significant code change, and it’s also what I use to check that a new market or language is ready. Because the test set is fixed, two runs are comparable. It answers: is this change better than what we have?</p>

<p><strong>Online evaluation</strong> runs against live traffic. Its job is to detect regressions and bugs in production, the things nobody thought to put in the test set. It answers: what did we miss?</p>

<figure class="fig">
<svg viewBox="0 0 680 216" role="img" aria-labelledby="ev2t ev2d">
  <title id="ev2t">Offline evaluation and online monitoring in one loop</title>
  <desc id="ev2d">A change goes through an on-demand offline evaluation run, which produces a report against a baseline. An engineer reads the report and decides whether to release. After release, online evaluation watches live traffic. A regression or new failure found there is turned into a new test case, which feeds the next change.</desc>
  <defs><marker id="ev2a" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="arrow" d="M0,0 L10,5 L0,10 z" /></marker></defs>
  <text class="h" x="10" y="16">BEFORE RELEASE · OFFLINE</text>
  <text class="h" x="10" y="208">AFTER RELEASE · ONLINE</text>
  <rect class="box" x="10" y="26" width="132" height="54" rx="6" />
  <text class="t" x="22" y="48">Change</text><text class="m" x="22" y="65">code or design</text>
  <rect class="box" x="186" y="26" width="132" height="54" rx="6" style="stroke:var(--fig-a)" />
  <text class="t" x="198" y="48">Offline run</text><text class="m" x="198" y="65">on demand</text>
  <rect class="box" x="362" y="26" width="132" height="54" rx="6" style="stroke:var(--fig-a)" />
  <text class="t" x="374" y="48">Report</text><text class="m" x="374" y="65">vs baseline</text>
  <rect class="box" x="538" y="26" width="132" height="54" rx="6" />
  <text class="t" x="550" y="48">Engineer</text><text class="m" x="550" y="65">reads, decides</text>
  <path class="ln" d="M142,53 L184,53" marker-end="url(#ev2a)" />
  <path class="ln" d="M318,53 L360,53" marker-end="url(#ev2a)" />
  <path class="ln" d="M494,53 L536,53" marker-end="url(#ev2a)" />
  <rect class="box" x="538" y="120" width="132" height="54" rx="6" />
  <text class="t" x="550" y="142">Release</text><text class="m" x="550" y="159">ship it</text>
  <rect class="box" x="362" y="120" width="132" height="54" rx="6" style="stroke:var(--fig-a)" />
  <text class="t" x="374" y="142">Online eval</text><text class="m" x="374" y="159">live traffic</text>
  <rect class="box" x="186" y="120" width="132" height="54" rx="6" style="stroke:var(--fig-b)" />
  <text class="tb" x="198" y="142">Regression</text><text class="m" x="198" y="159">or new failure</text>
  <rect class="box" x="10" y="120" width="132" height="54" rx="6" />
  <text class="t" x="22" y="142">Test set</text><text class="m" x="22" y="159">add the case</text>
  <path class="ln" d="M604,80 L604,118" marker-end="url(#ev2a)" />
  <path class="ln" d="M538,147 L496,147" marker-end="url(#ev2a)" />
  <path class="ln" d="M362,147 L320,147" marker-end="url(#ev2a)" />
  <path class="ln" d="M186,147 L144,147" marker-end="url(#ev2a)" />
  <path class="ln" d="M76,120 L76,82" marker-end="url(#ev2a)" />
</svg>
<figcaption>The two kinds of evaluation do different jobs. Offline runs answer "is this change better?" before release. Online evaluation answers "what did we miss?" after it, and feeds the offline set.</figcaption>
</figure>

<p>The loop closes when a failure found online becomes a new offline case. Over time the test set stops being what you thought users would ask and becomes a record of what actually broke. Neither kind is a substitute for the other, and neither is a substitute for understanding the task: someone has to read the failures.</p>

<p>Readiness reviews for a new market combine the offline results with something else: the history of issues that came through support tickets. The eval tells you how the system does on the test questions, and the ticket history tells you what people have actually run into. In practice, the findings from that combination have mostly led to improvements in the source content, not in the code.</p>

<h2 id="why-the-report-is-not-a-gate">Why the report is not a gate</h2>

<p>The default instinct, and it’s a reasonable one, is to wire the eval into CI and block a release if the score falls below a threshold. It’s what we do with unit tests. I chose not to, and the pipeline I built produces a comprehensive report that compares the current run with a baseline or with agreed criteria, and leaves the decision to an engineer.</p>

<p>My reasoning comes down to three points.</p>

<p><strong>The scores are noisy.</strong> Metrics like relevance and groundedness are commonly scored by a language model acting as a judge, and the agent itself is not deterministic. Run the same system twice and you’ll get slightly different numbers. A hard threshold turns that noise into a coin flip:</p>

<figure class="fig">
<svg viewBox="0 0 680 262" role="img" aria-labelledby="ev3t ev3d">
  <title id="ev3t">Twelve runs of an unchanged system against a fixed threshold</title>
  <desc id="ev3d">Simulated. Twelve repeated evaluation runs of the same system score between 0.78 and 0.82. A pass mark at 0.80 sits in the middle of that spread, so 5 of the twelve runs fail with no change to the system.</desc>
  <line class="grid" x1="60" x2="650" y1="200" y2="200" />
  <line class="grid" x1="60" x2="650" y1="146.7" y2="146.7" />
  <line class="grid" x1="60" x2="650" y1="93.3" y2="93.3" />
  <line class="grid" x1="60" x2="650" y1="40" y2="40" />
  <text class="m" x="52" y="204" text-anchor="end">0.74</text>
  <text class="m" x="52" y="150" text-anchor="end">0.78</text>
  <text class="m" x="52" y="97" text-anchor="end">0.82</text>
  <text class="m" x="52" y="44" text-anchor="end">0.86</text>
  <line class="sb dash" x1="60" x2="650" y1="120.0" y2="120.0" />
  <text class="tb" x="650" y="136.0" text-anchor="end">pass mark 0.80</text>
  <g tabindex="0"><title>Run 1: score 0.795, fails the 0.80 pass mark</title><circle class="fb" cx="93.8" cy="126.7" r="6" /></g>
  <text class="m" x="93.8" y="218" text-anchor="middle">1</text>
  <g tabindex="0"><title>Run 2: score 0.810, passes the 0.80 pass mark</title><circle class="fa" cx="141.2" cy="106.7" r="6" /></g>
  <text class="m" x="141.2" y="218" text-anchor="middle">2</text>
  <g tabindex="0"><title>Run 3: score 0.795, fails the 0.80 pass mark</title><circle class="fb" cx="188.8" cy="126.7" r="6" /></g>
  <text class="m" x="188.8" y="218" text-anchor="middle">3</text>
  <g tabindex="0"><title>Run 4: score 0.794, fails the 0.80 pass mark</title><circle class="fb" cx="236.2" cy="128.0" r="6" /></g>
  <text class="m" x="236.2" y="218" text-anchor="middle">4</text>
  <g tabindex="0"><title>Run 5: score 0.781, fails the 0.80 pass mark</title><circle class="fb" cx="283.8" cy="145.3" r="6" /></g>
  <text class="m" x="283.8" y="218" text-anchor="middle">5</text>
  <g tabindex="0"><title>Run 6: score 0.796, fails the 0.80 pass mark</title><circle class="fb" cx="331.2" cy="125.3" r="6" /></g>
  <text class="m" x="331.2" y="218" text-anchor="middle">6</text>
  <g tabindex="0"><title>Run 7: score 0.822, passes the 0.80 pass mark</title><circle class="fa" cx="378.8" cy="90.7" r="6" /></g>
  <text class="m" x="378.8" y="218" text-anchor="middle">7</text>
  <g tabindex="0"><title>Run 8: score 0.808, passes the 0.80 pass mark</title><circle class="fa" cx="426.2" cy="109.3" r="6" /></g>
  <text class="m" x="426.2" y="218" text-anchor="middle">8</text>
  <g tabindex="0"><title>Run 9: score 0.821, passes the 0.80 pass mark</title><circle class="fa" cx="473.8" cy="92.0" r="6" /></g>
  <text class="m" x="473.8" y="218" text-anchor="middle">9</text>
  <g tabindex="0"><title>Run 10: score 0.805, passes the 0.80 pass mark</title><circle class="fa" cx="521.2" cy="113.3" r="6" /></g>
  <text class="m" x="521.2" y="218" text-anchor="middle">10</text>
  <g tabindex="0"><title>Run 11: score 0.808, passes the 0.80 pass mark</title><circle class="fa" cx="568.8" cy="109.3" r="6" /></g>
  <text class="m" x="568.8" y="218" text-anchor="middle">11</text>
  <g tabindex="0"><title>Run 12: score 0.804, passes the 0.80 pass mark</title><circle class="fa" cx="616.2" cy="114.7" r="6" /></g>
  <text class="m" x="616.2" y="218" text-anchor="middle">12</text>
  <text class="m" x="355" y="238" text-anchor="middle">run number, same code and same test set each time</text>
  <text class="m" transform="translate(14 120) rotate(-90)" text-anchor="middle">answer score</text>
  <text class="h" x="60" y="20">BLUE PASSES · ORANGE FAILS</text>
</svg>
<figcaption>Simulated data. Nothing changed between runs, yet 5 of 12 fail. A hard threshold turns run-to-run noise into a red cross.</figcaption>
</figure>

<p><strong>A single threshold hides the trade-offs.</strong> A change that lifts recall while costing a little relevance on a narrow slice might be exactly the change you want. A gate sees one number and says no. A report shows the per-metric differences, the slices that moved, and the actual examples that got worse, so a person can judge whether the trade is worth it.</p>

<p><strong>A red cross ends the conversation.</strong> When a gate fails, the natural response is to make it pass: tune until the number moves, or lower the bar. When a report arrives, the natural response is to read it. I’d like the team to argue about specific failing examples, not about the threshold.</p>

<h2 id="where-a-gate-does-belong">Where a gate does belong</h2>

<p>I’m not against gates in general. A hard gate suits checks that are cheap, deterministic and clearly binary. A latency budget is a good example: either the response time is inside the budget or it isn’t, and there’s little to interpret. The same goes for a safety filter that must never be bypassed, or a schema check on tool output. Those I’d happily block on.</p>

<p>The metrics I’d keep as reports are the ones that need reading: relevance, and the finer judgements about groundedness. The rule of thumb I use is that if a person would want to look at the case before deciding, it shouldn’t be an automatic block.</p>

<h2 id="what-changes-on-a-team">What changes on a team</h2>

<p>Making the evaluation a report changes how people use it. It becomes something you read before a review, alongside the diff, and not a gate that you try to get past. Reports that people actually want to read need to be legible: baseline next to current, differences by metric and slice, and the worst examples one click away. If nobody reads the report, you’ve built a gate with extra steps, so it’s worth spending effort on the reading experience.</p>

<p>The practical lesson is to make evaluation part of the development loop, keep the reports readable by the people making the next decision, and keep a human in the seat where the decision is a judgement call.</p>

<h2 id="references">References</h2>

<ol>
  <li>Lewis, P. et al. <em>Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks</em>. NeurIPS, 2020. <a href="https://arxiv.org/abs/2005.11401">https://arxiv.org/abs/2005.11401</a></li>
  <li>Xia, B. et al. <em>Evaluation-Driven Development and Operations of LLM Agents: A Process Model and Reference Architecture</em>. arXiv, 2024. <a href="https://arxiv.org/abs/2411.13768">https://arxiv.org/abs/2411.13768</a></li>
  <li>Es, S. et al. <em>Ragas: Automated Evaluation of Retrieval Augmented Generation</em>. arXiv, 2023. <a href="https://arxiv.org/abs/2309.15217">https://arxiv.org/abs/2309.15217</a> — relevance, context and faithfulness metrics.</li>
  <li>Saad-Falcon, J. et al. <em>ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems</em>. NAACL, 2024. <a href="https://arxiv.org/abs/2311.09476">https://arxiv.org/abs/2311.09476</a></li>
  <li>Barnett, S. et al. <em>Seven Failure Points When Engineering a Retrieval Augmented Generation System</em>. arXiv, 2024. <a href="https://arxiv.org/abs/2401.05856">https://arxiv.org/abs/2401.05856</a></li>
  <li>Zheng, L. et al. <em>Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena</em>. NeurIPS Datasets and Benchmarks, 2023. <a href="https://arxiv.org/abs/2306.05685">https://arxiv.org/abs/2306.05685</a></li>
  <li>Miller, E. <em>Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations</em>. arXiv, 2024. <a href="https://arxiv.org/abs/2411.00640">https://arxiv.org/abs/2411.00640</a> — quantifying run-to-run noise.</li>
  <li>Jones, C., Wilkes, J. and Murphy, N. <em>Service Level Objectives</em>. In <em>Site Reliability Engineering</em>, Google, O’Reilly, 2016. <a href="https://sre.google/sre-book/service-level-objectives/">https://sre.google/sre-book/service-level-objectives/</a> — latency budgets as SLOs.</li>
</ol>]]></content><author><name>Fadhil Mochammad</name></author><category term="agents" /><summary type="html"><![CDATA[A single score cannot tell you whether an agent is useful. I work on a production assistant that answers questions by searching a body of documents and writing a reply from what it finds, a pattern usually called retrieval-augmented generation, or RAG. It’s agentic, meaning a model decides which steps to take, so there are more places for it to go wrong than in a fixed pipeline. This post is about how I evaluate it, and about one decision that surprises people: the evaluation produces a report for engineers, not an automatic pass or fail on deployment.]]></summary></entry><entry><title type="html">An agent that triages its own feedback</title><link href="https://fadhilmch.github.io/posts/an-agent-that-triages-its-own-feedback/" rel="alternate" type="text/html" title="An agent that triages its own feedback" /><published>2026-08-02T00:00:00+02:00</published><updated>2026-08-02T00:00:00+02:00</updated><id>https://fadhilmch.github.io/posts/an-agent-that-triages-its-own-feedback</id><content type="html" xml:base="https://fadhilmch.github.io/posts/an-agent-that-triages-its-own-feedback/"><![CDATA[<p>Once an AI assistant is in production, users start telling you when it gets things wrong. Someone clicks a thumbs-down, or leaves a comment like “this isn’t right for my situation”. That feedback is valuable, and it is also close to useless as it arrives. It says that something went wrong. It doesn’t say what, or whose problem it is.</p>

<p>Working out the what and the who is a small investigation every time. You find the conversation, read the question, look at what the assistant retrieved and what it said, check whether the right information exists anywhere, and decide whether to fix a document or fix the software. Then you find the person who owns that thing and explain it to them. Done by hand, on every flagged item, it eats the time of the people who should be improving the assistant.</p>

<p>I built and run a workflow that does the first pass of that investigation. This post describes how it’s put together and, more to the point, where I chose not to let the agent decide anything.</p>

<h2 id="the-shape-of-the-loop">The shape of the loop</h2>

<p>The workflow runs weekly, or on demand when someone wants an answer sooner. It reads the production traces for flagged conversations, and for each one it works through the same steps.</p>

<figure class="fig">
<svg viewBox="0 0 680 216" role="img" aria-labelledby="tr1t tr1d">
  <title id="tr1t">The feedback triage loop</title>
  <desc id="tr1d">Flagged feedback is replayed, then the supporting evidence is checked, then the likely cause is classified, then the case is routed to an owner. The output is a report and a follow-up ticket. If the evidence is missing or conflicting after the check, the case stops and goes to a person.</desc>
  <defs><marker id="tr1a" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="arrow" d="M0,0 L10,5 L0,10 z" /></marker></defs>
  <rect class="box" x="9" y="30" width="110" height="54" rx="6" />
  <text class="t" x="21" y="52">Feedback</text><text class="m" x="21" y="69">flagged trace</text>
  <rect class="box" x="147" y="30" width="110" height="54" rx="6" style="stroke:var(--fig-a)" />
  <text class="t" x="159" y="52">Replay</text><text class="m" x="159" y="69">ask it again</text>
  <rect class="box" x="285" y="30" width="110" height="54" rx="6" style="stroke:var(--fig-a)" />
  <text class="t" x="297" y="52">Evidence</text><text class="m" x="297" y="69">sources found</text>
  <rect class="box" x="423" y="30" width="110" height="54" rx="6" style="stroke:var(--fig-a)" />
  <text class="t" x="435" y="52">Classify</text><text class="m" x="435" y="69">likely cause</text>
  <rect class="box" x="561" y="30" width="110" height="54" rx="6" />
  <text class="t" x="573" y="52">Route</text><text class="m" x="573" y="69">to an owner</text>
  <path class="ln" d="M119,57 L145,57" marker-end="url(#tr1a)" />
  <path class="ln" d="M257,57 L283,57" marker-end="url(#tr1a)" />
  <path class="ln" d="M395,57 L421,57" marker-end="url(#tr1a)" />
  <path class="ln" d="M533,57 L559,57" marker-end="url(#tr1a)" />
  <rect class="box" x="561" y="140" width="110" height="54" rx="6" />
  <text class="t" x="573" y="162">Report</text><text class="m" x="573" y="179">and a ticket</text>
  <path class="ln" d="M616,84 L616,138" marker-end="url(#tr1a)" />
  <rect class="box" x="230" y="140" width="270" height="54" rx="6" style="stroke:var(--fig-b)" />
  <text class="tb" x="242" y="162">Missing or conflicting evidence</text><text class="m" x="242" y="179">stop here: a person reviews</text>
  <path class="sb dash" d="M340,84 L340,138" marker-end="url(#tr1a)" />
  <text class="m" x="9" y="112">weekly, or on demand</text>
  <text class="m" x="9" y="210">blue outline: agent steps</text>
</svg>
<figcaption>Three steps are agent work. The fourth is a lookup. Anywhere the evidence doesn't line up, the loop stops and hands the case to a person.</figcaption>
</figure>

<p><strong>Replay.</strong> The trace records what the user asked and what the assistant did. The workflow asks the same question again, in the language it was originally asked. This matters more than it sounds. Some complaints describe a problem that has since been fixed, some can’t be reproduced, and some reproduce every time. A flag you can’t reproduce is a different kind of case from one you can.</p>

<p><strong>Check the evidence.</strong> This is the step I’d insist on in any version of this system. The workflow looks at what the assistant retrieved and asks whether the information needed for a good answer was actually there. Was the relevant document in the knowledge base at all? Was it retrieved? Does the answer follow from what was retrieved? Everything after this step depends on the answer, so it comes before any conclusion, not after.</p>

<p><strong>Classify.</strong> From the evidence, the workflow picks a likely cause from a fixed list. The most important split on that list is between two families of problem:</p>

<ul>
  <li><strong>A content gap.</strong> The knowledge the assistant needed is missing, out of date, unclear or contradicts another document. Fixing the assistant’s code won’t help. The indexed content needs to change.</li>
  <li><strong>A configuration or architecture issue.</strong> The information exists, but the application didn’t use it well: retrieval missed it, a setting excluded it, or the flow around the model handled the question badly. The fix is in the software.</li>
</ul>

<p><strong>Route.</strong> Each cause on the list maps to an owner, and that mapping is a plain lookup table. More on that below.</p>

<p>The workflow handles feedback, traces and conversations in whatever language they arrived in, since the assistant is used in many. The output is an investigation report per case, and it can open a follow-up ticket for the owner.</p>

<h2 id="why-the-split-matters">Why the split matters</h2>

<p>It’s easy to treat all bad feedback as “the AI got it wrong” and send it to the engineers. That’s a waste, and in my experience it’s also the wrong diagnosis for a lot of cases. If the answer was wrong because a policy document was ambiguous, the engineer can do nothing useful with the ticket. The person who maintains that document can.</p>

<p>Separating content gaps from application faults sends work to the people who can act on it. It also gives you a quiet second benefit. Counting causes over a few weeks tells you where the assistant’s weak spots really are. If most of the problems turn out to be content gaps, the most useful improvement isn’t a cleverer prompt but better source material. Which way it falls will differ from system to system, but you can’t know until the causes are counted.</p>

<p>A made-up example shows the difference. Suppose a user asks how a policy applies in their region and marks the answer as wrong. The replay reproduces the same answer. The evidence check finds that the regional document exists and was retrieved, but that the answer quotes the global one. That points at how the application chose between sources, so the cause is configuration, and the ticket goes to the engineers. Change one fact: the regional document doesn’t exist, so the assistant fell back to the global one, which was the best it had. Now the cause is a content gap, and the ticket goes to whoever maintains that content. The complaint was identical both times, and the right owner was different.</p>

<h2 id="where-the-agent-stops">Where the agent stops</h2>

<p>An agent that always produces an answer is dangerous in a triage role, because the answer is often a confident guess. So the workflow has a rule: <strong>missing or conflicting evidence goes to a human.</strong></p>

<p>Missing evidence means the replay couldn’t be run or the trace was incomplete. Conflicting evidence means the signals disagree, for example the right document was retrieved but the answer contradicts it in a way that could point at either the content or the model. In both cases the workflow doesn’t pick a side. It says what it found, says what it couldn’t establish, and hands the case over with the material already gathered.</p>

<p>I think this is the difference between a triage tool people trust and one they route around. The first time it confidently blames the wrong team, the reports stop being read. A tool that says “I can’t tell, here is what I’ve got” loses very little, and it keeps its credibility for the cases where it does commit.</p>

<h2 id="judgement-by-the-model-decisions-by-code">Judgement by the model, decisions by code</h2>

<p>The last design choice is about who does what. The model is good at the fuzzy parts: reading a trace, comparing an answer with retrieved passages, deciding which cause on the list fits. It’s a poor choice for deciding who gets paged, because you want that to be predictable and easy to change.</p>

<p>So the cause-to-owner step is deterministic. Something like this:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">OWNER_BY_CAUSE</span> <span class="o">=</span> <span class="p">{</span>
    <span class="sh">"</span><span class="s">content_gap</span><span class="sh">"</span><span class="p">:</span> <span class="sh">"</span><span class="s">content-owners</span><span class="sh">"</span><span class="p">,</span>
    <span class="sh">"</span><span class="s">configuration</span><span class="sh">"</span><span class="p">:</span> <span class="sh">"</span><span class="s">application-team</span><span class="sh">"</span><span class="p">,</span>
    <span class="sh">"</span><span class="s">architecture</span><span class="sh">"</span><span class="p">:</span> <span class="sh">"</span><span class="s">engineering</span><span class="sh">"</span><span class="p">,</span>
<span class="p">}</span>

<span class="k">def</span> <span class="nf">route</span><span class="p">(</span><span class="n">cause</span><span class="p">):</span>
    <span class="c1"># An unknown or missing cause never guesses; it goes to a person.
</span>    <span class="k">return</span> <span class="n">OWNER_BY_CAUSE</span><span class="p">.</span><span class="nf">get</span><span class="p">(</span><span class="n">cause</span><span class="p">,</span> <span class="sh">"</span><span class="s">human-review</span><span class="sh">"</span><span class="p">)</span>
</code></pre></div></div>

<p>This is a sketch, not the real table, but the principle is what I’m after. If a team reorganises, you edit a mapping. If someone asks why a case went where it did, the answer is a line of code and not “the model felt like it”. And because the classifier has to choose from a fixed list, its output is something the rest of the system can check.</p>

<figure class="fig">
<svg viewBox="0 0 680 168" role="img" aria-labelledby="tr2t tr2d">
  <title id="tr2t">Which steps are model judgement and which are plain code</title>
  <desc id="tr2d">Replay, evidence check and cause classification are done by the agent using judgement. Owner lookup, report writing and ticket creation are deterministic code. A cause outside the fixed list routes to human review.</desc>
  <defs><marker id="tr2a" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="arrow" d="M0,0 L10,5 L0,10 z" /></marker></defs>
  <text class="h" x="10" y="16">MODEL JUDGEMENT</text>
  <text class="h" x="670" y="16" text-anchor="end">PLAIN CODE</text>
  <rect class="box" x="10" y="30" width="120" height="50" rx="6" style="stroke:var(--fig-a)" />
  <text class="t" x="22" y="51">Replay</text><text class="m" x="22" y="68">re-ask</text>
  <rect class="box" x="150" y="30" width="120" height="50" rx="6" style="stroke:var(--fig-a)" />
  <text class="t" x="162" y="51">Evidence</text><text class="m" x="162" y="68">compare</text>
  <rect class="box" x="290" y="30" width="120" height="50" rx="6" style="stroke:var(--fig-a)" />
  <text class="t" x="302" y="51">Cause</text><text class="m" x="302" y="68">from a list</text>
  <rect class="box" x="430" y="30" width="110" height="50" rx="6" />
  <text class="t" x="442" y="51">Owner</text><text class="m" x="442" y="68">table lookup</text>
  <rect class="box" x="560" y="30" width="110" height="50" rx="6" />
  <text class="t" x="572" y="51">Ticket</text><text class="m" x="572" y="68">report</text>
  <path class="ln" d="M130,55 L148,55" marker-end="url(#tr2a)" />
  <path class="ln" d="M270,55 L288,55" marker-end="url(#tr2a)" />
  <path class="ln" d="M410,55 L428,55" marker-end="url(#tr2a)" />
  <path class="ln" d="M540,55 L558,55" marker-end="url(#tr2a)" />
  <line class="dash ln" x1="420" x2="420" y1="24" y2="92" />
  <rect class="box" x="290" y="110" width="250" height="46" rx="6" style="stroke:var(--fig-b)" />
  <text class="tb" x="302" y="130">Not on the list, or unsure</text><text class="m" x="302" y="147">routes to a person</text>
  <path class="sb dash" d="M350,80 L350,108" marker-end="url(#tr2a)" />
</svg>
<figcaption>The dashed line is where judgement ends. Everything to its right can be read, tested and changed like any other code.</figcaption>
</figure>

<h2 id="what-i-havent-measured">What I haven’t measured</h2>

<p>The workflow has helped the team deal with recurring issues and content gaps. I haven’t formally measured how often its root-cause call and its owner assignment are right, and I’d rather say so than quote a number. The comparison I’d make at this stage isn’t accuracy but effort: how much of the manual investigation is already done when a case reaches a person. Because it produces a report with the replay and the evidence attached, a reviewer starts from a worked lead instead of a blank page. My own estimate is that it removes most of the digging per item, based on comparing the manual and the automated routes. It’s an estimate, not a controlled measurement.</p>

<p>If you’re building something like this, the parts I’d carry over are these: check the evidence before concluding anything, make the cause list small and explicit, keep routing deterministic, and give the agent a clear way to say it doesn’t know. The goal is to shorten the investigation, not to automate certainty.</p>

<h2 id="references">References</h2>

<ol>
  <li>Lewis, P. et al. <em>Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks</em>. NeurIPS, 2020. <a href="https://arxiv.org/abs/2005.11401">https://arxiv.org/abs/2005.11401</a></li>
  <li>Barnett, S. et al. <em>Seven Failure Points When Engineering a Retrieval Augmented Generation System</em>. arXiv, 2024. <a href="https://arxiv.org/abs/2401.05856">https://arxiv.org/abs/2401.05856</a> — a taxonomy of RAG failures.</li>
  <li>Anthropic. <em>Building Effective AI Agents</em>. Anthropic Engineering, 2024. <a href="https://www.anthropic.com/engineering/building-effective-agents">https://www.anthropic.com/engineering/building-effective-agents</a> — routing and workflows versus agents.</li>
  <li>Es, S. et al. <em>Ragas: Automated Evaluation of Retrieval Augmented Generation</em>. arXiv, 2023. <a href="https://arxiv.org/abs/2309.15217">https://arxiv.org/abs/2309.15217</a></li>
  <li>Zheng, L. et al. <em>Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena</em>. NeurIPS Datasets and Benchmarks, 2023. <a href="https://arxiv.org/abs/2306.05685">https://arxiv.org/abs/2306.05685</a></li>
  <li>Geifman, Y. and El-Yaniv, R. <em>Selective Classification for Deep Neural Networks</em>. arXiv, 2017. <a href="https://arxiv.org/abs/1705.08500">https://arxiv.org/abs/1705.08500</a> — letting a model abstain when unsure.</li>
</ol>]]></content><author><name>Fadhil Mochammad</name></author><category term="agents" /><category term="mlops" /><summary type="html"><![CDATA[Once an AI assistant is in production, users start telling you when it gets things wrong. Someone clicks a thumbs-down, or leaves a comment like “this isn’t right for my situation”. That feedback is valuable, and it is also close to useless as it arrives. It says that something went wrong. It doesn’t say what, or whose problem it is.]]></summary></entry><entry><title type="html">Golden datasets without the gold rush</title><link href="https://fadhilmch.github.io/posts/golden-datasets-without-the-gold-rush/" rel="alternate" type="text/html" title="Golden datasets without the gold rush" /><published>2026-06-20T00:00:00+02:00</published><updated>2026-06-20T00:00:00+02:00</updated><id>https://fadhilmch.github.io/posts/golden-datasets-without-the-gold-rush</id><content type="html" xml:base="https://fadhilmch.github.io/posts/golden-datasets-without-the-gold-rush/"><![CDATA[<p>A <strong>golden dataset</strong> is the set of questions, with known-good answers, that you run a system against to see whether a change made it better or worse. For a retrieval-augmented generation (RAG) assistant, meaning one that searches a body of documents and writes an answer from what it finds, the golden set is what turns “it feels better” into evidence.</p>

<p>The usual advice for building one is to sit down and write fifty question and answer pairs by hand. That advice is fine for a prototype. It stops working when the assistant serves several markets, several languages and a knowledge base that keeps changing, because nobody has the time to write, and re-write, enough examples to cover all of that.</p>

<p>I work on a production assistant of this kind, and I’m building a pipeline that generates the candidates instead. This post describes the design: simulated personas propose evaluation records, and a person approves each one before it counts. I’ll keep it about the mechanism, not the system it runs in, and I won’t give sizes or rates because those aren’t mine to share.</p>

<h2 id="what-a-record-contains">What a record contains</h2>

<p>An evaluation record is more than a question and an answer. In the pipeline I’m describing, each one has:</p>

<ul>
  <li><strong>A query</strong>, written the way a user would type it.</li>
  <li><strong>An expected answer</strong>, the reference the assistant’s answer is compared with.</li>
  <li><strong>A source document URL</strong>, the place in the knowledge base where the answer comes from.</li>
  <li><strong>Metadata</strong>: the language of the query, the geography it applies to, and the intent behind it.</li>
</ul>

<p>The source URL matters most. It <strong>anchors</strong> the record to something checkable. An expected answer with no source is an opinion. An expected answer that points at a specific document can be verified by opening the document, and it gives you a second thing to measure: whether the assistant’s retrieval found that document at all. I’ll come back to that in the last section.</p>

<p>The metadata is what lets you slice results later. An average score across everything hides a lot. If the same assistant does well in one language and badly in another, you want to see that, which means every record needs to say which language it belongs to.</p>

<h2 id="where-personas-come-in">Where personas come in</h2>

<p>The obvious way to generate questions with a language model is to hand it a document and ask for questions about it. The result is predictable. The questions echo the document’s own vocabulary, they’re all polite and well-formed, and they all have a tidy answer in the text. Real users don’t ask like that.</p>

<p>A <strong>persona</strong> is a short description of a kind of user: who they are, where they are, what they’re trying to get done, how they tend to phrase things. The generator is asked to write queries as that person would, about a document sampled from the knowledge index. Different personas produce different questions from the same source:</p>

<ul>
  <li>someone who knows the topic well and asks a narrow, specific question;</li>
  <li>someone new who doesn’t know the right term and describes the situation instead;</li>
  <li>someone whose question only applies in their region;</li>
  <li>someone who writes in another language, or mixes two.</li>
</ul>

<figure class="fig">
<svg viewBox="0 0 680 236" role="img" aria-labelledby="gd1t gd1d">
  <title id="gd1t">Generating a golden dataset with a human approval gate</title>
  <desc id="gd1d">Source documents from the knowledge index and a simulated persona feed a generator that proposes records containing a query, an expected answer, a source URL and metadata. A human reviewer checks each proposal. Approved records join the golden set. Rejected or edited ones are dropped or fixed.</desc>
  <defs><marker id="gd1a" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="arrow" d="M0,0 L10,5 L0,10 z" /></marker></defs>
  <text class="h" x="10" y="16">GENERATE LIBERALLY</text>
  <text class="h" x="520" y="16">APPROVE CONSERVATIVELY</text>
  <rect class="box" x="10" y="30" width="150" height="54" rx="6" />
  <text class="t" x="22" y="52">Index</text><text class="m" x="22" y="69">source documents</text>
  <rect class="box" x="180" y="30" width="150" height="54" rx="6" />
  <text class="t" x="192" y="52">Persona</text><text class="m" x="192" y="69">asks like a user</text>
  <rect class="box" x="350" y="30" width="150" height="54" rx="6" />
  <text class="t" x="362" y="52">Proposed record</text><text class="m" x="362" y="69">query, answer, URL</text>
  <rect class="box" x="520" y="30" width="150" height="54" rx="6" />
  <text class="t" x="532" y="52">Human review</text><text class="m" x="532" y="69">every record</text>
  <path class="ln" d="M160,57 L178,57" marker-end="url(#gd1a)" />
  <path class="ln" d="M330,57 L348,57" marker-end="url(#gd1a)" />
  <path class="ln" d="M500,57 L518,57" marker-end="url(#gd1a)" />
  <rect class="box" x="520" y="130" width="150" height="54" rx="6" style="stroke:var(--fig-a)" />
  <text class="t" x="532" y="152">Golden set</text><text class="m" x="532" y="169">approved only</text>
  <rect class="box" x="290" y="130" width="150" height="54" rx="6" />
  <text class="t" x="302" y="152">Edit or drop</text><text class="m" x="302" y="169">back to the pool</text>
  <path class="sa" d="M620,84 L620,128" marker-end="url(#gd1a)" />
  <text class="ta" x="628" y="110">approve</text>
  <path class="sb" d="M560,84 C560,110 480,110 440,140" marker-end="url(#gd1a)" />
  <text class="tb" x="470" y="104">reject</text>
  <text class="m" x="10" y="214">Metadata (language, geography, intent) travels with every record.</text>
</svg>
<figcaption>Generation is cheap and wide. The gate is narrow and manual, and nothing reaches the golden set without passing through it.</figcaption>
</figure>

<p>Personas also give you a coverage lever. If a language or an intent is thin in the set, you ask for more of it by changing the persona mix, instead of asking a colleague to write more by hand. The pipeline is configurable, and the persona mix is the natural knob for that: you change the request, not the code.</p>

<h2 id="generate-liberally-approve-conservatively">Generate liberally, approve conservatively</h2>

<p>The design decision I care most about is where the human sits. Generation is cheap, so I want it to over-produce and be varied. Approval is the expensive part, so it needs to be strict. Nothing a model generates becomes golden on its own, whatever score it gets from another model.</p>

<p>The reason is simple. A golden set is the ruler you measure everything else with. If the ruler is wrong, every comparison you make with it is wrong in a way that’s hard to see, because the numbers still look like numbers. A model checking a model’s homework tends to agree with itself.</p>

<p>What does the reviewer actually check? Three things, together, for each record:</p>

<ol>
  <li><strong>The query.</strong> Would a real person plausibly ask this? Is it in the language and for the geography the metadata says?</li>
  <li><strong>The expected answer.</strong> Is it correct, and complete enough to compare against? Does it say more than the source does?</li>
  <li><strong>The source.</strong> Open the document. Does it support the answer? Is it current, and is it the right document for this user, or one that has been superseded?</li>
</ol>

<p>Checking all three at once is faster than it sounds, because the record puts the three side by side. The reviewer is reading, comparing and deciding, not researching from scratch. That’s the whole trade: the model does the tedious drafting and the person keeps the judgement.</p>

<h2 id="what-goes-wrong-in-generated-data">What goes wrong in generated data</h2>

<p>Generated evaluation data fails in a handful of predictable ways, and it helps to know which check catches which.</p>

<figure class="fig">
<svg viewBox="0 0 680 236" role="img" aria-labelledby="gd2t gd2d">
  <title id="gd2t">Failure modes in generated records and the review check that catches each</title>
  <desc id="gd2d">A matrix of five failure modes against four checks: query, answer, source and set. Nobody would ask this is caught by the query check. Answer says more than the source is caught by answer and source. Stale source is caught by source. Wrong language or intent tag is caught by query. Near-duplicate records are caught by looking at the set as a whole.</desc>
  <text class="h" x="10" y="18">FAILURE IN A GENERATED RECORD</text>
  <text class="h" x="410" y="18" text-anchor="middle">QUERY</text>
  <text class="h" x="475" y="18" text-anchor="middle">ANSWER</text>
  <text class="h" x="545" y="18" text-anchor="middle">SOURCE</text>
  <text class="h" x="620" y="18" text-anchor="middle">SET</text>
  <line class="grid" x1="10" x2="670" y1="28" y2="28" />
  <text class="t" x="10" y="52">Nobody would ask this</text>
  <text class="t" x="10" y="90">Answer says more than the source</text>
  <text class="t" x="10" y="128">Source is stale or superseded</text>
  <text class="t" x="10" y="166">Wrong language or intent tag</text>
  <text class="t" x="10" y="204">Near-duplicate of another record</text>
  <line class="grid" x1="10" x2="670" y1="66" y2="66" />
  <line class="grid" x1="10" x2="670" y1="104" y2="104" />
  <line class="grid" x1="10" x2="670" y1="142" y2="142" />
  <line class="grid" x1="10" x2="670" y1="180" y2="180" />
  <g tabindex="0"><title>Nobody would ask this: caught by reading the query</title><circle class="fa" cx="410" cy="48" r="6" /></g>
  <g tabindex="0"><title>Answer says more than the source: caught by comparing the answer, and by opening the source</title><circle class="fa" cx="475" cy="86" r="6" /><circle class="fa" cx="545" cy="86" r="6" /></g>
  <g tabindex="0"><title>Stale source: caught by opening the source</title><circle class="fa" cx="545" cy="124" r="6" /></g>
  <g tabindex="0"><title>Wrong language or intent tag: caught by reading the query against its metadata</title><circle class="fa" cx="410" cy="162" r="6" /></g>
  <g tabindex="0"><title>Near-duplicate: only visible when looking at the set as a whole</title><circle class="fa" cx="620" cy="200" r="6" /></g>
</svg>
<figcaption>A dot means that check catches the failure. No single check covers everything, and the last row shows why coverage has to be looked at across the set, not record by record.</figcaption>
</figure>

<p>The last row is the one people skip. Each record can be individually fine and the set can still be lopsided: twenty near-identical questions about the same document, and nothing about the one that users actually struggle with. Reviewing records one by one won’t show you that. You need to look at the distribution by document, language, geography and intent, and then adjust the personas or the sampling.</p>

<h2 id="a-bigger-set-is-not-a-better-set">A bigger set is not a better set</h2>

<p>It’s tempting to report the size of the golden set as an achievement. I’d avoid that. A set padded with near-duplicates is worse than a smaller one that covers the real spread of what people ask, because it gives you a confident number about a narrow slice. The measures I’d care about are coverage, how much reviewers agree with each other, and how many records they had to edit. A headline count tells you none of that.</p>

<p>The same goes for keeping the set alive. The knowledge base changes, so a record that was correct last quarter can point at a document that has since been rewritten. Because each record carries its source URL, you can find those records mechanically and send them back for review, instead of discovering the problem when a score drops for no visible reason.</p>

<h2 id="what-the-anchor-buys-you-later">What the anchor buys you later</h2>

<p>Because each record names its source, the evaluation can ask two separate questions. Did the assistant give a good answer? And did it retrieve the document that answer should come from? When the first is bad and the second is good, the fault is in how the answer was written. When both are bad, look at retrieval. That split is much harder to make when the golden set is only questions and answers, and it’s the main reason I insist on the URL.</p>

<p>If you’re building something similar, the parts I’d keep are these: anchor every record to a source, use personas to widen what gets asked, keep a person as the only route into the golden set, and review the set as a whole as well as record by record. The rest is tooling.</p>

<h2 id="references">References</h2>

<ol>
  <li>Lewis, P. et al. <em>Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks</em>. NeurIPS, 2020. <a href="https://arxiv.org/abs/2005.11401">https://arxiv.org/abs/2005.11401</a></li>
  <li>Es, S. et al. <em>Ragas: Automated Evaluation of Retrieval Augmented Generation</em>. arXiv, 2023. <a href="https://arxiv.org/abs/2309.15217">https://arxiv.org/abs/2309.15217</a></li>
  <li>Saad-Falcon, J. et al. <em>ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems</em>. NAACL, 2024. <a href="https://arxiv.org/abs/2311.09476">https://arxiv.org/abs/2311.09476</a> — synthetic evaluation data plus human annotations.</li>
  <li>Ge, T. et al. <em>Scaling Synthetic Data Creation with 1,000,000,000 Personas</em>. arXiv, 2024. <a href="https://arxiv.org/abs/2406.20094">https://arxiv.org/abs/2406.20094</a></li>
  <li>Zheng, L. et al. <em>Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena</em>. NeurIPS Datasets and Benchmarks, 2023. <a href="https://arxiv.org/abs/2306.05685">https://arxiv.org/abs/2306.05685</a></li>
  <li>Panickssery, A., Bowman, S. R. and Feng, S. <em>LLM Evaluators Recognize and Favor Their Own Generations</em>. arXiv, 2024. <a href="https://arxiv.org/abs/2404.13076">https://arxiv.org/abs/2404.13076</a> — why a model shouldn’t grade its own output.</li>
</ol>]]></content><author><name>Fadhil Mochammad</name></author><category term="agents" /><summary type="html"><![CDATA[A golden dataset is the set of questions, with known-good answers, that you run a system against to see whether a change made it better or worse. For a retrieval-augmented generation (RAG) assistant, meaning one that searches a body of documents and writes an answer from what it finds, the golden set is what turns “it feels better” into evidence.]]></summary></entry><entry><title type="html">Bandits in production: Thompson sampling, explained with coupons</title><link href="https://fadhilmch.github.io/posts/thompson-sampling-explained-with-coupons/" rel="alternate" type="text/html" title="Bandits in production: Thompson sampling, explained with coupons" /><published>2026-04-11T00:00:00+02:00</published><updated>2026-04-11T00:00:00+02:00</updated><id>https://fadhilmch.github.io/posts/thompson-sampling-explained-with-coupons</id><content type="html" xml:base="https://fadhilmch.github.io/posts/thompson-sampling-explained-with-coupons/"><![CDATA[<p>Say you have three discount coupons and want to know which one gets the most people to buy. The textbook answer is an A/B test: split traffic evenly, wait until the result is significant, then ship the winner. It works, but for the whole test a third of your users see each losing coupon, even after the data has made it fairly obvious they’re losing.</p>

<p>A <strong>multi-armed bandit</strong> handles that by moving traffic toward what’s working while the test is still running. On the experimentation platform I worked on, we ran bandits for exactly this kind of problem, coupon optimisation and adaptive traffic allocation, using <strong>Thompson sampling</strong>. This post explains the algorithm with coupons, then shows how we fitted it into a production system.</p>

<h2 id="explore-or-exploit">Explore or exploit</h2>

<p>Every bandit balances two needs. <strong>Exploiting</strong> means showing the coupon that looks best so far, to earn conversions now. <strong>Exploring</strong> means still showing the others sometimes, because “looks best” is based on limited data and might be wrong. Pure exploitation locks in an early lucky guess; pure exploration is just an A/B test. Thompson sampling gets the balance from how uncertain it is about each coupon.</p>

<h2 id="thompson-sampling-in-one-picture">Thompson sampling in one picture</h2>

<p>For each coupon, keep two numbers: how many people saw it and how many converted. Those two numbers define a <strong>Beta distribution</strong>, the bandit’s belief about that coupon’s true conversion rate, here the share of people who buy after seeing it. Little data gives a wide curve; lots of data gives a narrow one.</p>

<figure class="fig">
<svg viewBox="0 0 680 252" role="img" aria-labelledby="t1t t1d">
  <title id="t1t">Posterior beliefs about three coupons</title>
  <desc id="t1d">Beta posterior curves for three coupons. A, 12 conversions out of 200, is narrow around 6%. B, 20 out of 200, is narrow around 10%. C, 3 out of 40, is wide around 7.5%. One random draw from each curve is marked on the axis.</desc>
  <line class="grid" x1="60" x2="650" y1="210" y2="210" />
  <text class="m" x="60.0" y="228" text-anchor="middle">0%</text>
  <text class="m" x="207.5" y="228" text-anchor="middle">5%</text>
  <text class="m" x="355.0" y="228" text-anchor="middle">10%</text>
  <text class="m" x="502.50000000000006" y="228" text-anchor="middle">15%</text>
  <text class="m" x="650.0" y="228" text-anchor="middle">20%</text>
  <text class="m" x="650" y="244" text-anchor="end">conversion rate</text>
  <g tabindex="0"><title>Coupon A: 12 / 200 conversions: belief about its rate</title><path d="M63.0,210.0 L65.9,210.0 L68.8,210.0 L71.8,210.0 L74.8,210.0 L77.7,210.0 L80.7,210.0 L83.6,210.0 L86.5,210.0 L89.5,210.0 L92.4,210.0 L95.4,210.0 L98.3,210.0 L101.3,210.0 L104.2,209.9 L107.2,209.9 L110.2,209.8 L113.1,209.7 L116.0,209.5 L119.0,209.2 L121.9,208.9 L124.9,208.4 L127.8,207.7 L130.8,206.8 L133.8,205.7 L136.7,204.3 L139.6,202.7 L142.6,200.6 L145.6,198.3 L148.5,195.5 L151.4,192.3 L154.4,188.6 L157.4,184.5 L160.3,180.0 L163.2,175.0 L166.2,169.7 L169.1,163.9 L172.1,157.8 L175.0,151.3 L178.0,144.6 L180.9,137.7 L183.9,130.7 L186.8,123.6 L189.8,116.4 L192.8,109.3 L195.7,102.4 L198.7,95.6 L201.6,89.1 L204.6,83.0 L207.5,77.2 L210.4,71.8 L213.4,66.9 L216.3,62.5 L219.3,58.7 L222.2,55.4 L225.2,52.7 L228.1,50.6 L231.1,49.2 L234.0,48.3 L237.0,48.0 L239.9,48.3 L242.9,49.1 L245.8,50.5 L248.8,52.4 L251.8,54.7 L254.7,57.5 L257.6,60.6 L260.6,64.1 L263.6,68.0 L266.5,72.1 L269.4,76.4 L272.4,81.0 L275.3,85.7 L278.3,90.5 L281.2,95.4 L284.2,100.4 L287.1,105.4 L290.1,110.4 L293.0,115.4 L296.0,120.3 L298.9,125.1 L301.9,129.9 L304.9,134.5 L307.8,139.0 L310.8,143.3 L313.7,147.6 L316.6,151.6 L319.6,155.5 L322.5,159.2 L325.5,162.8 L328.4,166.1 L331.4,169.3 L334.3,172.4 L337.3,175.2 L340.2,177.9 L343.2,180.4 L346.1,182.8 L349.1,185.0 L352.1,187.1 L355.0,189.0 L357.9,190.8 L360.9,192.5 L363.8,194.1 L366.8,195.5 L369.7,196.8 L372.7,198.0 L375.6,199.1 L378.6,200.2 L381.5,201.1 L384.5,202.0 L387.4,202.7 L390.4,203.5 L393.3,204.1 L396.3,204.7 L399.2,205.2 L402.2,205.7 L405.1,206.2 L408.1,206.6 L411.1,206.9 L414.0,207.3 L416.9,207.6 L419.9,207.8 L422.9,208.1 L425.8,208.3 L428.8,208.5 L431.7,208.6 L434.6,208.8 L437.6,208.9 L440.6,209.1 L443.5,209.2 L446.4,209.3 L449.4,209.4 L452.4,209.4 L455.3,209.5 L458.2,209.6 L461.2,209.6 L464.2,209.7 L467.1,209.7 L470.1,209.7 L473.0,209.8 L475.9,209.8 L478.9,209.8 L481.8,209.8 L484.8,209.9 L487.7,209.9 L490.7,209.9 L493.6,209.9 L496.6,209.9 L499.5,209.9 L502.5,209.9 L505.4,209.9 L508.4,210.0 L511.3,210.0 L514.3,210.0 L517.2,210.0 L520.2,210.0 L523.1,210.0 L526.1,210.0 L529.0,210.0 L532.0,210.0 L535.0,210.0 L537.9,210.0 L540.8,210.0 L543.8,210.0 L546.8,210.0 L549.7,210.0 L552.6,210.0 L555.6,210.0 L558.5,210.0 L561.5,210.0 L564.5,210.0 L567.4,210.0 L570.3,210.0 L573.3,210.0 L576.2,210.0 L579.2,210.0 L582.1,210.0 L585.1,210.0 L588.0,210.0 L591.0,210.0 L593.9,210.0 L596.9,210.0 L599.8,210.0 L602.8,210.0 L605.8,210.0 L608.7,210.0 L611.6,210.0 L614.6,210.0 L617.5,210.0 L620.5,210.0 L623.4,210.0 L626.4,210.0 L629.4,210.0 L632.3,210.0 L635.2,210.0 L638.2,210.0 L641.1,210.0 L644.1,210.0 L647.0,210.0 L650.0,210.0" style="fill:none;stroke:var(--muted);stroke-width:2" /></g>
  <g tabindex="0"><title>Coupon B: 20 / 200 conversions: belief about its rate</title><path d="M63.0,210.0 L65.9,210.0 L68.8,210.0 L71.8,210.0 L74.8,210.0 L77.7,210.0 L80.7,210.0 L83.6,210.0 L86.5,210.0 L89.5,210.0 L92.4,210.0 L95.4,210.0 L98.3,210.0 L101.3,210.0 L104.2,210.0 L107.2,210.0 L110.2,210.0 L113.1,210.0 L116.0,210.0 L119.0,210.0 L121.9,210.0 L124.9,210.0 L127.8,210.0 L130.8,210.0 L133.8,210.0 L136.7,210.0 L139.6,210.0 L142.6,210.0 L145.6,210.0 L148.5,210.0 L151.4,210.0 L154.4,210.0 L157.4,210.0 L160.3,210.0 L163.2,210.0 L166.2,210.0 L169.1,209.9 L172.1,209.9 L175.0,209.9 L178.0,209.8 L180.9,209.8 L183.9,209.7 L186.8,209.6 L189.8,209.5 L192.8,209.4 L195.7,209.2 L198.7,208.9 L201.6,208.7 L204.6,208.3 L207.5,207.9 L210.4,207.5 L213.4,206.9 L216.3,206.2 L219.3,205.5 L222.2,204.6 L225.2,203.6 L228.1,202.5 L231.1,201.2 L234.0,199.8 L237.0,198.2 L239.9,196.4 L242.9,194.5 L245.8,192.4 L248.8,190.1 L251.8,187.6 L254.7,184.9 L257.6,182.1 L260.6,179.1 L263.6,175.8 L266.5,172.5 L269.4,168.9 L272.4,165.2 L275.3,161.4 L278.3,157.5 L281.2,153.5 L284.2,149.4 L287.1,145.2 L290.1,141.0 L293.0,136.7 L296.0,132.5 L298.9,128.3 L301.9,124.2 L304.9,120.1 L307.8,116.2 L310.8,112.3 L313.7,108.6 L316.6,105.1 L319.6,101.8 L322.5,98.7 L325.5,95.7 L328.4,93.1 L331.4,90.7 L334.3,88.5 L337.3,86.6 L340.2,85.0 L343.2,83.7 L346.1,82.7 L349.1,82.0 L352.1,81.5 L355.0,81.4 L357.9,81.5 L360.9,82.0 L363.8,82.7 L366.8,83.6 L369.7,84.8 L372.7,86.3 L375.6,88.0 L378.6,89.8 L381.5,91.9 L384.5,94.2 L387.4,96.7 L390.4,99.3 L393.3,102.0 L396.3,104.9 L399.2,107.8 L402.2,110.9 L405.1,114.0 L408.1,117.2 L411.1,120.4 L414.0,123.7 L416.9,127.0 L419.9,130.2 L422.9,133.5 L425.8,136.8 L428.8,140.0 L431.7,143.2 L434.6,146.3 L437.6,149.4 L440.6,152.4 L443.5,155.3 L446.4,158.2 L449.4,161.0 L452.4,163.7 L455.3,166.3 L458.2,168.8 L461.2,171.2 L464.2,173.5 L467.1,175.8 L470.1,177.9 L473.0,179.9 L475.9,181.9 L478.9,183.8 L481.8,185.5 L484.8,187.2 L487.7,188.8 L490.7,190.3 L493.6,191.7 L496.6,193.0 L499.5,194.3 L502.5,195.5 L505.4,196.6 L508.4,197.6 L511.3,198.6 L514.3,199.5 L517.2,200.3 L520.2,201.1 L523.1,201.8 L526.1,202.5 L529.0,203.1 L532.0,203.7 L535.0,204.3 L537.9,204.8 L540.8,205.2 L543.8,205.6 L546.8,206.0 L549.7,206.4 L552.6,206.7 L555.6,207.0 L558.5,207.3 L561.5,207.6 L564.5,207.8 L567.4,208.0 L570.3,208.2 L573.3,208.4 L576.2,208.5 L579.2,208.7 L582.1,208.8 L585.1,208.9 L588.0,209.0 L591.0,209.1 L593.9,209.2 L596.9,209.3 L599.8,209.4 L602.8,209.4 L605.8,209.5 L608.7,209.6 L611.6,209.6 L614.6,209.6 L617.5,209.7 L620.5,209.7 L623.4,209.8 L626.4,209.8 L629.4,209.8 L632.3,209.8 L635.2,209.8 L638.2,209.9 L641.1,209.9 L644.1,209.9 L647.0,209.9 L650.0,209.9" style="fill:none;stroke:var(--fig-a);stroke-width:2" /></g>
  <g tabindex="0"><title>Coupon C: 3 / 40 conversions: belief about its rate</title><path d="M63.0,210.0 L65.9,210.0 L68.8,209.9 L71.8,209.8 L74.8,209.7 L77.7,209.5 L80.7,209.3 L83.6,208.9 L86.5,208.6 L89.5,208.1 L92.4,207.6 L95.4,206.9 L98.3,206.3 L101.3,205.5 L104.2,204.7 L107.2,203.8 L110.2,202.8 L113.1,201.8 L116.0,200.7 L119.0,199.5 L121.9,198.3 L124.9,197.1 L127.8,195.8 L130.8,194.4 L133.8,193.1 L136.7,191.6 L139.6,190.2 L142.6,188.8 L145.6,187.3 L148.5,185.8 L151.4,184.3 L154.4,182.8 L157.4,181.3 L160.3,179.8 L163.2,178.2 L166.2,176.7 L169.1,175.3 L172.1,173.8 L175.0,172.3 L178.0,170.9 L180.9,169.5 L183.9,168.1 L186.8,166.7 L189.8,165.4 L192.8,164.1 L195.7,162.8 L198.7,161.6 L201.6,160.4 L204.6,159.3 L207.5,158.1 L210.4,157.1 L213.4,156.0 L216.3,155.1 L219.3,154.1 L222.2,153.2 L225.2,152.4 L228.1,151.6 L231.1,150.8 L234.0,150.1 L237.0,149.4 L239.9,148.8 L242.9,148.2 L245.8,147.7 L248.8,147.2 L251.8,146.8 L254.7,146.4 L257.6,146.0 L260.6,145.7 L263.6,145.5 L266.5,145.2 L269.4,145.1 L272.4,144.9 L275.3,144.8 L278.3,144.8 L281.2,144.8 L284.2,144.8 L287.1,144.8 L290.1,144.9 L293.0,145.0 L296.0,145.2 L298.9,145.4 L301.9,145.6 L304.9,145.9 L307.8,146.2 L310.8,146.5 L313.7,146.8 L316.6,147.2 L319.6,147.6 L322.5,148.0 L325.5,148.4 L328.4,148.9 L331.4,149.4 L334.3,149.9 L337.3,150.4 L340.2,150.9 L343.2,151.5 L346.1,152.1 L349.1,152.7 L352.1,153.3 L355.0,153.9 L357.9,154.5 L360.9,155.2 L363.8,155.8 L366.8,156.5 L369.7,157.1 L372.7,157.8 L375.6,158.5 L378.6,159.2 L381.5,159.9 L384.5,160.6 L387.4,161.3 L390.4,162.0 L393.3,162.7 L396.3,163.5 L399.2,164.2 L402.2,164.9 L405.1,165.6 L408.1,166.3 L411.1,167.1 L414.0,167.8 L416.9,168.5 L419.9,169.2 L422.9,169.9 L425.8,170.6 L428.8,171.3 L431.7,172.1 L434.6,172.8 L437.6,173.5 L440.6,174.1 L443.5,174.8 L446.4,175.5 L449.4,176.2 L452.4,176.9 L455.3,177.5 L458.2,178.2 L461.2,178.8 L464.2,179.5 L467.1,180.1 L470.1,180.7 L473.0,181.4 L475.9,182.0 L478.9,182.6 L481.8,183.2 L484.8,183.8 L487.7,184.4 L490.7,184.9 L493.6,185.5 L496.6,186.1 L499.5,186.6 L502.5,187.1 L505.4,187.7 L508.4,188.2 L511.3,188.7 L514.3,189.2 L517.2,189.7 L520.2,190.2 L523.1,190.7 L526.1,191.2 L529.0,191.6 L532.0,192.1 L535.0,192.5 L537.9,193.0 L540.8,193.4 L543.8,193.8 L546.8,194.3 L549.7,194.7 L552.6,195.1 L555.6,195.5 L558.5,195.8 L561.5,196.2 L564.5,196.6 L567.4,196.9 L570.3,197.3 L573.3,197.6 L576.2,198.0 L579.2,198.3 L582.1,198.6 L585.1,198.9 L588.0,199.3 L591.0,199.6 L593.9,199.8 L596.9,200.1 L599.8,200.4 L602.8,200.7 L605.8,201.0 L608.7,201.2 L611.6,201.5 L614.6,201.7 L617.5,202.0 L620.5,202.2 L623.4,202.4 L626.4,202.7 L629.4,202.9 L632.3,203.1 L635.2,203.3 L638.2,203.5 L641.1,203.7 L644.1,203.9 L647.0,204.1 L650.0,204.3" style="fill:none;stroke:var(--fig-b);stroke-width:2" /></g>
  <text x="213.4" y="62.0" text-anchor="end" style="fill:var(--muted);font-size:11px">A · 12/200 · narrow, low</text>
  <text x="390.4" y="95.4" text-anchor="start" style="fill:var(--fig-a);font-size:11px">B · 20/200 · narrow, high</text>
  <text x="458.2" y="170.2" text-anchor="start" style="fill:var(--fig-b);font-size:11px">C · 3/40 · wide: unsure</text>
  <circle cx="215.3" cy="210" r="5" style="fill:var(--muted);stroke:var(--bg);stroke-width:2" />
  <circle cx="246.9" cy="210" r="5" style="fill:var(--fig-a);stroke:var(--bg);stroke-width:2" />
  <circle cx="470.7" cy="210" r="5" style="fill:var(--fig-b);stroke:var(--bg);stroke-width:2" />
  <text class="m" x="60" y="16">dots on the axis: one random draw per coupon · this round C drew highest, so C is shown</text>
</svg>
<figcaption>What the bandit believes after some traffic. Each curve is a Beta distribution over a coupon's true conversion rate. C has little data, so its curve is wide and it still wins some draws. Probability of being best: B 58%, C 38%, A under 5%.</figcaption>
</figure>

<p>The whole algorithm is: for each user, draw one random value from each coupon’s curve, and show the coupon with the highest draw.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">import</span> <span class="n">random</span>

<span class="k">def</span> <span class="nf">choose</span><span class="p">(</span><span class="n">stats</span><span class="p">):</span>
    <span class="c1"># stats: {"A": (conversions, views), ...}
</span>    <span class="n">draws</span> <span class="o">=</span> <span class="p">{</span>
        <span class="n">coupon</span><span class="p">:</span> <span class="n">random</span><span class="p">.</span><span class="nf">betavariate</span><span class="p">(</span><span class="mi">1</span> <span class="o">+</span> <span class="n">conv</span><span class="p">,</span> <span class="mi">1</span> <span class="o">+</span> <span class="n">views</span> <span class="o">-</span> <span class="n">conv</span><span class="p">)</span>
        <span class="k">for</span> <span class="n">coupon</span><span class="p">,</span> <span class="p">(</span><span class="n">conv</span><span class="p">,</span> <span class="n">views</span><span class="p">)</span> <span class="ow">in</span> <span class="n">stats</span><span class="p">.</span><span class="nf">items</span><span class="p">()</span>
    <span class="p">}</span>
    <span class="k">return</span> <span class="nf">max</span><span class="p">(</span><span class="n">draws</span><span class="p">,</span> <span class="n">key</span><span class="o">=</span><span class="n">draws</span><span class="p">.</span><span class="n">get</span><span class="p">)</span>
</code></pre></div></div>

<p>That’s enough to get the balance right. A coupon with a high, narrow curve wins most draws, so it gets most of the traffic. A coupon with a wide curve sometimes draws high, so it still gets explored until the data settles what it’s worth. A coupon with a low, narrow curve almost never wins. Nobody has to tune an exploration rate; uncertainty does the work.</p>

<h2 id="sampling-per-request-or-in-batches">Sampling per request, or in batches</h2>

<p>The textbook version draws per user, updates the counts after every conversion, and draws again. Ours ran differently: the bandit was a <strong>batch job that edits a traffic split</strong>. Batching like this has real advantages over per-request sampling:</p>

<ul>
  <li><strong>Purchases arrive late.</strong> Someone who sees a coupon might buy hours later. Updating per request would mostly update on “no purchase yet”.</li>
  <li><strong>Users see consistent variants.</strong> Drawing per request could show the same person a different coupon on every page.</li>
  <li><strong>The request path stays simple.</strong> Assignment is a hot, latency-sensitive service that already knows how to assign users by a traffic split.</li>
</ul>

<p>Every few hours, the job recomputed each coupon’s posterior from logged exposures and conversions, estimated the probability that each coupon is the best, and wrote those probabilities back as the experiment’s new traffic split.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">def</span> <span class="nf">split_from_posteriors</span><span class="p">(</span><span class="n">stats</span><span class="p">,</span> <span class="n">draws</span><span class="o">=</span><span class="mi">10_000</span><span class="p">):</span>
    <span class="n">wins</span> <span class="o">=</span> <span class="p">{</span><span class="n">coupon</span><span class="p">:</span> <span class="mi">0</span> <span class="k">for</span> <span class="n">coupon</span> <span class="ow">in</span> <span class="n">stats</span><span class="p">}</span>
    <span class="k">for</span> <span class="n">_</span> <span class="ow">in</span> <span class="nf">range</span><span class="p">(</span><span class="n">draws</span><span class="p">):</span>
        <span class="n">wins</span><span class="p">[</span><span class="nf">choose</span><span class="p">(</span><span class="n">stats</span><span class="p">)]</span> <span class="o">+=</span> <span class="mi">1</span>
    <span class="k">return</span> <span class="p">{</span><span class="n">coupon</span><span class="p">:</span> <span class="n">n</span> <span class="o">/</span> <span class="n">draws</span> <span class="k">for</span> <span class="n">coupon</span><span class="p">,</span> <span class="n">n</span> <span class="ow">in</span> <span class="n">wins</span><span class="p">.</span><span class="nf">items</span><span class="p">()}</span>
</code></pre></div></div>

<p>Allocating traffic in proportion to “probability of being best” is exactly what Thompson sampling does on average, just applied in batches. The new split then travels to the assignment service like any other config change, which is why the <a href="/posts/from-polling-to-push/">config delivery work</a> mattered: bandits were among the most frequent writers.</p>

<figure class="fig">
<svg viewBox="0 0 680 226" role="img" aria-labelledby="t4t t4d">
  <title id="t4t">The bandit loop</title>
  <desc id="t4d">The assignment service assigns users using the current split and logs exposures. Conversions join them in an event store. Every few hours a bandit updater recomputes posteriors and the probability each variant is best, and writes a new split to the config store, which flows back to the assignment service.</desc>
  <defs><marker id="t4" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="arrow" d="M0,0 L10,5 L0,10 z" /></marker></defs>
  <rect class="box" x="10" y="20" width="200" height="60" rx="6" /><text class="t" x="22" y="42">Assignment service</text><text class="m" x="22" y="60">hash user → split</text>
  <rect class="box" x="470" y="20" width="200" height="60" rx="6" /><text class="t" x="482" y="42">Event store</text><text class="m" x="482" y="60">exposures, conversions</text>
  <rect class="box" x="470" y="130" width="200" height="86" rx="6" style="stroke:var(--fig-a)" /><text class="t" x="482" y="152">Bandit updater</text><text class="m" x="482" y="170">every few hours:</text><text class="m" x="482" y="186">update posteriors,</text><text class="m" x="482" y="202">P(best) → new split</text>
  <rect class="box" x="10" y="140" width="200" height="60" rx="6" /><text class="t" x="22" y="162">Config store</text><text class="m" x="22" y="180">split: B 90 · C 5 · A 5</text>
  <path class="ln" d="M210,50 L468,50" marker-end="url(#t4)" /><text class="m" x="340" y="42" text-anchor="middle">who saw what, who converted</text>
  <path class="ln" d="M570,80 L570,128" marker-end="url(#t4)" />
  <path class="sa" d="M470,172 L212,172" marker-end="url(#t4)" /><text class="ta" x="340" y="164" text-anchor="middle">write new split</text>
  <path class="ln" d="M110,140 L110,82" marker-end="url(#t4)" /><text class="m" x="118" y="116">config path</text>
</svg>
<figcaption>The loop in production. Nothing in the request path samples anything: the bandit is a batch job that edits the traffic split, and the split travels to the assignment service like any other config change.</figcaption>
</figure>

<h2 id="what-it-looks-like-over-time">What it looks like over time</h2>

<p>Here is a simulated run, with the true conversion rates set to 6%, 10% and 7.5%, and a few hundred users between updates:</p>

<figure class="fig">
<svg viewBox="0 0 680 240" role="img" aria-labelledby="t2t t2d">
  <title id="t2t">Traffic allocation over time</title>
  <desc id="t2d">Stacked bars of traffic share per update. All three coupons start at a third. After a few updates the bandit hesitates between B and C, then settles on about 90% to B, keeping 5% each for A and C.</desc>
  <g tabindex="0"><title>Update 0: coupon B gets 33% of traffic</title><rect x="61.0" y="144.3" width="40.1" height="54.7" rx="2" style="fill:var(--fig-a);opacity:.75" /></g>
  <g tabindex="0"><title>Update 0: coupon C gets 33% of traffic</title><rect x="61.0" y="87.7" width="40.1" height="54.7" rx="2" style="fill:var(--fig-b);opacity:.75" /></g>
  <g tabindex="0"><title>Update 0: coupon A gets 33% of traffic</title><rect x="61.0" y="31.0" width="40.1" height="54.7" rx="2" style="fill:var(--muted);opacity:.75" /></g>
  <g tabindex="0"><title>Update 1: coupon B gets 47% of traffic</title><rect x="103.1" y="121.6" width="40.1" height="77.4" rx="2" style="fill:var(--fig-a);opacity:.75" /></g>
  <g tabindex="0"><title>Update 1: coupon C gets 48% of traffic</title><rect x="103.1" y="39.3" width="40.1" height="80.3" rx="2" style="fill:var(--fig-b);opacity:.75" /></g>
  <g tabindex="0"><title>Update 1: coupon A gets 5% of traffic</title><rect x="103.1" y="31.0" width="40.1" height="6.3" rx="2" style="fill:var(--muted);opacity:.75" /></g>
  <g tabindex="0"><title>Update 2: coupon B gets 90% of traffic</title><rect x="145.3" y="47.5" width="40.1" height="151.5" rx="2" style="fill:var(--fig-a);opacity:.75" /></g>
  <g tabindex="0"><title>Update 2: coupon C gets 5% of traffic</title><rect x="145.3" y="39.2" width="40.1" height="6.2" rx="2" style="fill:var(--fig-b);opacity:.75" /></g>
  <g tabindex="0"><title>Update 2: coupon A gets 5% of traffic</title><rect x="145.3" y="31.0" width="40.1" height="6.2" rx="2" style="fill:var(--muted);opacity:.75" /></g>
  <g tabindex="0"><title>Update 3: coupon B gets 58% of traffic</title><rect x="187.4" y="101.6" width="40.1" height="97.4" rx="2" style="fill:var(--fig-a);opacity:.75" /></g>
  <g tabindex="0"><title>Update 3: coupon C gets 36% of traffic</title><rect x="187.4" y="40.9" width="40.1" height="58.7" rx="2" style="fill:var(--fig-b);opacity:.75" /></g>
  <g tabindex="0"><title>Update 3: coupon A gets 6% of traffic</title><rect x="187.4" y="31.0" width="40.1" height="7.9" rx="2" style="fill:var(--muted);opacity:.75" /></g>
  <g tabindex="0"><title>Update 4: coupon B gets 86% of traffic</title><rect x="229.6" y="55.5" width="40.1" height="143.5" rx="2" style="fill:var(--fig-a);opacity:.75" /></g>
  <g tabindex="0"><title>Update 4: coupon C gets 6% of traffic</title><rect x="229.6" y="45.2" width="40.1" height="8.3" rx="2" style="fill:var(--fig-b);opacity:.75" /></g>
  <g tabindex="0"><title>Update 4: coupon A gets 8% of traffic</title><rect x="229.6" y="31.0" width="40.1" height="12.2" rx="2" style="fill:var(--muted);opacity:.75" /></g>
  <g tabindex="0"><title>Update 5: coupon B gets 89% of traffic</title><rect x="271.7" y="50.3" width="40.1" height="148.7" rx="2" style="fill:var(--fig-a);opacity:.75" /></g>
  <g tabindex="0"><title>Update 5: coupon C gets 5% of traffic</title><rect x="271.7" y="41.8" width="40.1" height="6.5" rx="2" style="fill:var(--fig-b);opacity:.75" /></g>
  <g tabindex="0"><title>Update 5: coupon A gets 6% of traffic</title><rect x="271.7" y="31.0" width="40.1" height="8.8" rx="2" style="fill:var(--muted);opacity:.75" /></g>
  <g tabindex="0"><title>Update 6: coupon B gets 91% of traffic</title><rect x="313.9" y="46.9" width="40.1" height="152.1" rx="2" style="fill:var(--fig-a);opacity:.75" /></g>
  <g tabindex="0"><title>Update 6: coupon C gets 5% of traffic</title><rect x="313.9" y="39.0" width="40.1" height="6.0" rx="2" style="fill:var(--fig-b);opacity:.75" /></g>
  <g tabindex="0"><title>Update 6: coupon A gets 5% of traffic</title><rect x="313.9" y="31.0" width="40.1" height="6.0" rx="2" style="fill:var(--muted);opacity:.75" /></g>
  <g tabindex="0"><title>Update 7: coupon B gets 91% of traffic</title><rect x="356.0" y="47.0" width="40.1" height="152.0" rx="2" style="fill:var(--fig-a);opacity:.75" /></g>
  <g tabindex="0"><title>Update 7: coupon C gets 5% of traffic</title><rect x="356.0" y="39.0" width="40.1" height="6.0" rx="2" style="fill:var(--fig-b);opacity:.75" /></g>
  <g tabindex="0"><title>Update 7: coupon A gets 5% of traffic</title><rect x="356.0" y="31.0" width="40.1" height="6.0" rx="2" style="fill:var(--muted);opacity:.75" /></g>
  <g tabindex="0"><title>Update 8: coupon B gets 91% of traffic</title><rect x="398.1" y="46.7" width="40.1" height="152.3" rx="2" style="fill:var(--fig-a);opacity:.75" /></g>
  <g tabindex="0"><title>Update 8: coupon C gets 5% of traffic</title><rect x="398.1" y="38.8" width="40.1" height="5.8" rx="2" style="fill:var(--fig-b);opacity:.75" /></g>
  <g tabindex="0"><title>Update 8: coupon A gets 5% of traffic</title><rect x="398.1" y="31.0" width="40.1" height="5.8" rx="2" style="fill:var(--muted);opacity:.75" /></g>
  <g tabindex="0"><title>Update 9: coupon B gets 88% of traffic</title><rect x="440.3" y="52.1" width="40.1" height="146.9" rx="2" style="fill:var(--fig-a);opacity:.75" /></g>
  <g tabindex="0"><title>Update 9: coupon C gets 5% of traffic</title><rect x="440.3" y="43.9" width="40.1" height="6.2" rx="2" style="fill:var(--fig-b);opacity:.75" /></g>
  <g tabindex="0"><title>Update 9: coupon A gets 8% of traffic</title><rect x="440.3" y="31.0" width="40.1" height="10.9" rx="2" style="fill:var(--muted);opacity:.75" /></g>
  <g tabindex="0"><title>Update 10: coupon B gets 89% of traffic</title><rect x="482.4" y="49.7" width="40.1" height="149.3" rx="2" style="fill:var(--fig-a);opacity:.75" /></g>
  <g tabindex="0"><title>Update 10: coupon C gets 5% of traffic</title><rect x="482.4" y="41.6" width="40.1" height="6.1" rx="2" style="fill:var(--fig-b);opacity:.75" /></g>
  <g tabindex="0"><title>Update 10: coupon A gets 6% of traffic</title><rect x="482.4" y="31.0" width="40.1" height="8.6" rx="2" style="fill:var(--muted);opacity:.75" /></g>
  <g tabindex="0"><title>Update 11: coupon B gets 89% of traffic</title><rect x="524.6" y="49.1" width="40.1" height="149.9" rx="2" style="fill:var(--fig-a);opacity:.75" /></g>
  <g tabindex="0"><title>Update 11: coupon C gets 5% of traffic</title><rect x="524.6" y="41.0" width="40.1" height="6.1" rx="2" style="fill:var(--fig-b);opacity:.75" /></g>
  <g tabindex="0"><title>Update 11: coupon A gets 6% of traffic</title><rect x="524.6" y="31.0" width="40.1" height="8.0" rx="2" style="fill:var(--muted);opacity:.75" /></g>
  <g tabindex="0"><title>Update 12: coupon B gets 91% of traffic</title><rect x="566.7" y="47.1" width="40.1" height="151.9" rx="2" style="fill:var(--fig-a);opacity:.75" /></g>
  <g tabindex="0"><title>Update 12: coupon C gets 5% of traffic</title><rect x="566.7" y="39.0" width="40.1" height="6.0" rx="2" style="fill:var(--fig-b);opacity:.75" /></g>
  <g tabindex="0"><title>Update 12: coupon A gets 5% of traffic</title><rect x="566.7" y="31.0" width="40.1" height="6.0" rx="2" style="fill:var(--muted);opacity:.75" /></g>
  <g tabindex="0"><title>Update 13: coupon B gets 91% of traffic</title><rect x="608.9" y="47.0" width="40.1" height="152.0" rx="2" style="fill:var(--fig-a);opacity:.75" /></g>
  <g tabindex="0"><title>Update 13: coupon C gets 5% of traffic</title><rect x="608.9" y="39.0" width="40.1" height="6.0" rx="2" style="fill:var(--fig-b);opacity:.75" /></g>
  <g tabindex="0"><title>Update 13: coupon A gets 5% of traffic</title><rect x="608.9" y="31.0" width="40.1" height="6.0" rx="2" style="fill:var(--muted);opacity:.75" /></g>
  <text class="m" x="81.1" y="218" text-anchor="middle">0</text>
  <text class="m" x="165.4" y="218" text-anchor="middle">2</text>
  <text class="m" x="249.6" y="218" text-anchor="middle">4</text>
  <text class="m" x="333.9" y="218" text-anchor="middle">6</text>
  <text class="m" x="418.2" y="218" text-anchor="middle">8</text>
  <text class="m" x="502.5" y="218" text-anchor="middle">10</text>
  <text class="m" x="586.8" y="218" text-anchor="middle">12</text>
  <text class="m" x="650" y="234" text-anchor="end">update (every few hours)</text>
  <text class="m" x="52" y="204" text-anchor="end">0%</text>
  <text class="m" x="52" y="119.0" text-anchor="end">50%</text>
  <text class="m" x="52" y="34" text-anchor="end">100%</text>
  <rect x="60" y="8" width="10" height="10" rx="2" style="fill:var(--fig-a)" /><text class="m" x="76" y="17">B</text>
  <rect x="100" y="8" width="10" height="10" rx="2" style="fill:var(--fig-b)" /><text class="m" x="116" y="17">C</text>
  <rect x="140" y="8" width="10" height="10" rx="2" style="fill:var(--muted)" /><text class="m" x="156" y="17">A</text>
  <text class="m" x="650" y="17" text-anchor="end">share of traffic per update · simulated</text>
</svg>
<figcaption>A simulated run with true rates of 6%, 10% and 7.5%. The split starts even, wavers between B and C while C's curve is still wide, then settles on B. A 5% floor keeps the losers in the test.</figcaption>
</figure>

<p>The early updates are the interesting part. After the first batch, B and C look about equally good and split the traffic between them. As C collects more data its curve narrows below B’s, and traffic moves to B. A drops out almost immediately.</p>

<p>The payoff is measured as <strong>regret</strong>: conversions lost compared with already knowing the best coupon.</p>

<figure class="fig">
<svg viewBox="0 0 680 244" role="img" aria-labelledby="t3t t3d">
  <title id="t3t">Cumulative regret, fixed split versus Thompson sampling</title>
  <desc id="t3d">Two lines of conversions lost compared with always showing the best coupon. A fixed even split loses 76 by the end; Thompson sampling loses 21.</desc>
  <line class="grid" x1="60" x2="650" y1="200" y2="200" />
  <g tabindex="0"><title>Fixed A/B/C split: 76 conversions lost by the end</title><path d="M60.0,189.6 L105.4,179.1 L150.8,168.7 L196.2,158.3 L241.5,147.9 L286.9,137.4 L332.3,127.0 L377.7,116.6 L423.1,106.1 L468.5,95.7 L513.8,85.3 L559.2,74.9 L604.6,64.4 L650.0,54.0" class="ln dash" style="stroke-width:2" /></g>
  <g tabindex="0"><title>Thompson sampling: 21 conversions lost by the end</title><path d="M60.0,189.6 L105.4,182.8 L150.8,181.3 L196.2,175.9 L241.5,173.6 L286.9,171.9 L332.3,170.5 L377.7,169.1 L423.1,167.7 L468.5,165.7 L513.8,164.0 L559.2,162.4 L604.6,161.0 L650.0,159.6" class="sa" /></g>
  <text class="m" x="650.0" y="46.0" text-anchor="end">fixed even split · 76 lost</text>
  <text class="ta" x="650.0" y="151.6" text-anchor="end">Thompson sampling · 21 lost</text>
  <text class="m" x="60" y="18">conversions lost versus always showing the best coupon · simulated, 3,500 users</text>
  <text class="m" x="60.0" y="218" text-anchor="middle">0</text>
  <text class="m" x="150.8" y="218" text-anchor="middle">2</text>
  <text class="m" x="241.5" y="218" text-anchor="middle">4</text>
  <text class="m" x="332.3" y="218" text-anchor="middle">6</text>
  <text class="m" x="423.1" y="218" text-anchor="middle">8</text>
  <text class="m" x="513.8" y="218" text-anchor="middle">10</text>
  <text class="m" x="604.6" y="218" text-anchor="middle">12</text>
  <text class="m" x="650" y="234" text-anchor="end">update</text>
</svg>
<figcaption>Regret: conversions lost compared with a world where you already knew the best coupon. A fixed split keeps paying for its losers at the same rate; the bandit stops paying once it's confident.</figcaption>
</figure>

<p>Both lines start the same, because both begin with an even split. After that, the fixed split keeps losing conversions at the same rate until the test ends, while the bandit’s losses flatten once it’s confident.</p>

<h2 id="what-production-adds">What production adds</h2>

<p>The algorithm is ten lines. Most of the work is everything around it:</p>

<ul>
  <li><strong>Define the reward carefully.</strong> A bandit optimises exactly what you give it. Our reward was a purchase, not a click on the coupon: the two can favour different coupons. The window for counting a purchase matters too; too short, and it under-counts people who take a while to decide.</li>
  <li><strong>Keep a floor.</strong> If a losing variant’s share can reach zero, the bandit can never notice that it has improved. A small minimum share keeps every arm observable.</li>
  <li><strong>Expect the world to change.</strong> Coupon performance drifts with seasons, weekdays and campaigns. Old data can be discounted, or the counts computed over a recent window, so the bandit can change its mind.</li>
  <li><strong>Decide what happens when the split moves.</strong> If users are hashed into buckets by split, changing the split moves some users between variants. Either that’s acceptable for the product, or first assignments need to be stored.</li>
  <li><strong>Don’t use it to measure.</strong> A bandit is built to earn while it learns, not to produce a clean estimate of the difference between variants. Because the losing arms get little traffic, their estimates stay noisy, and adaptive allocation biases naive confidence intervals. If the question is “how much better is B, and is it significant?”, run an A/B test.</li>
</ul>

<h2 id="when-to-reach-for-a-bandit">When to reach for a bandit</h2>

<p>Use an A/B test when you need to <em>learn</em>: a product decision, an effect size, something you’ll report. Use a bandit when you need to <em>earn</em> while choosing among options with fast, measurable feedback: coupons, banners, notification copy, ranking tweaks. If the options change often, or the test would otherwise run for weeks while showing users a clearly worse option, a bandit usually pays for itself.</p>

<h2 id="references">References</h2>

<ol>
  <li>Thompson, W. R. <em>On the Likelihood That One Unknown Probability Exceeds Another in View of the Evidence of Two Samples</em>. Biometrika, 1933. <a href="https://doi.org/10.1093/biomet/25.3-4.285">https://doi.org/10.1093/biomet/25.3-4.285</a> — the original idea.</li>
  <li>Russo, D. et al. <em>A Tutorial on Thompson Sampling</em>. Foundations and Trends in Machine Learning, 2018. <a href="https://arxiv.org/abs/1707.02038">https://arxiv.org/abs/1707.02038</a></li>
  <li>Lattimore, T. and Szepesvári, C. <em>Bandit Algorithms</em>. Cambridge University Press, 2020. <a href="https://www.cambridge.org/core/books/bandit-algorithms/8E39FD004E6CE036680F90DD0C6F09FC">https://www.cambridge.org/core/books/bandit-algorithms/8E39FD004E6CE036680F90DD0C6F09FC</a></li>
  <li>Chapelle, O. and Li, L. <em>An Empirical Evaluation of Thompson Sampling</em>. Advances in Neural Information Processing Systems 24 (NIPS), 2011. <a href="https://papers.nips.cc/paper/2011/hash/e53a0a2978c28872a4505bdb51db06dc-Abstract.html">https://papers.nips.cc/paper/2011/hash/e53a0a2978c28872a4505bdb51db06dc-Abstract.html</a></li>
  <li>Agrawal, S. and Goyal, N. <em>Analysis of Thompson Sampling for the Multi-armed Bandit Problem</em>. arXiv, 2011 (rev. 2012). <a href="https://arxiv.org/abs/1111.1797">https://arxiv.org/abs/1111.1797</a> — regret bounds.</li>
  <li>Nie, X. et al. <em>Why Adaptively Collected Data Have Negative Bias and How to Correct for It</em>. AISTATS, 2018. <a href="https://arxiv.org/abs/1708.01977">https://arxiv.org/abs/1708.01977</a></li>
</ol>]]></content><author><name>Fadhil Mochammad</name></author><category term="exp" /><summary type="html"><![CDATA[Say you have three discount coupons and want to know which one gets the most people to buy. The textbook answer is an A/B test: split traffic evenly, wait until the result is significant, then ship the winner. It works, but for the whole test a third of your users see each losing coupon, even after the data has made it fairly obvious they’re losing.]]></summary></entry><entry><title type="html">One definition of conversion</title><link href="https://fadhilmch.github.io/posts/one-definition-of-conversion/" rel="alternate" type="text/html" title="One definition of conversion" /><published>2026-02-28T00:00:00+01:00</published><updated>2026-02-28T00:00:00+01:00</updated><id>https://fadhilmch.github.io/posts/one-definition-of-conversion</id><content type="html" xml:base="https://fadhilmch.github.io/posts/one-definition-of-conversion/"><![CDATA[<p>Two dashboards show a metric called “conversion” for the same week. The numbers differ. Someone asks which one is right, and the next hour goes on reading two pieces of SQL side by side. By the end nobody has learned anything about the product. The conversation has become a debate about queries.</p>

<p>I designed a metric semantic layer for an experimentation platform that served dozens of product teams: a centralised metadata schema, plus a query engine on BigQuery that computed metrics from it. The catalogue held more than 100 reusable metrics, and every product team’s experiments used them. This post is about what that work taught me. Most of it was about people, and the schema was the smaller part.</p>

<h2 id="what-a-semantic-layer-is">What a semantic layer is</h2>

<p>A <strong>semantic layer</strong> is a place where the meaning of a metric is written down once, in a form that both people and programs can read. Instead of each dashboard or experiment analysis carrying its own SQL for “conversion”, they all refer to one definition, and a query engine turns that definition into the query that runs against the warehouse.</p>

<p>Vendors have plenty to say about this. I’m going to skip the theory and describe the problem it solves.</p>

<h2 id="how-drift-happens">How drift happens</h2>

<p>Nobody sets out to define a metric wrongly. Drift happens because every team has a reasonable local reason:</p>

<ul>
  <li>One team counts a conversion when an order is created. Another counts it when the order is paid.</li>
  <li>One divides by all sessions. Another divides by sessions from logged-in users.</li>
  <li>One excludes internal test traffic. Another never got round to it.</li>
  <li>One counts per user per day. Another counts per session.</li>
</ul>

<p>Each choice is defensible. The trouble is that all of them carry the same name, and a name is what people compare.</p>

<figure class="fig">
<svg viewBox="0 0 680 346" role="img" aria-labelledby="odc1t odc1d">
  <title id="odc1t">Four local definitions of conversion versus one governed definition</title>
  <desc id="odc1d">Illustrative. Four dashboards each compute conversion with their own SQL and show four different numbers: 3.1, 4.8, 2.2 and 5.6 percent. In the governed version, one definition feeds both experiment analysis and reports, which show the same number.</desc>
  <defs><marker id="odc1a" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="arrow" d="M0,0 L10,5 L0,10 z" /></marker></defs>
  <text class="h" x="6" y="16">BEFORE · EACH DASHBOARD OWNS A COPY</text>
  <rect class="box" x="6" y="28" width="155" height="76" rx="6" style="stroke:var(--fig-b);stroke-width:2" />
  <text class="t" x="16" y="48">Dashboard A</text><text class="m" x="16" y="66">orders / sessions</text><text class="tb" x="16" y="90">3.1%</text>
  <rect class="box" x="175" y="28" width="155" height="76" rx="6" style="stroke:var(--fig-b);stroke-width:2" />
  <text class="t" x="185" y="48">Dashboard B</text><text class="m" x="185" y="66">paid / visitors</text><text class="tb" x="185" y="90">4.8%</text>
  <rect class="box" x="344" y="28" width="155" height="76" rx="6" style="stroke:var(--fig-b);stroke-width:2" />
  <text class="t" x="354" y="48">Experiment X</text><text class="m" x="354" y="66">orders / logged-in</text><text class="tb" x="354" y="90">2.2%</text>
  <rect class="box" x="513" y="28" width="161" height="76" rx="6" style="stroke:var(--fig-b);stroke-width:2" />
  <text class="t" x="523" y="48">Weekly report</text><text class="m" x="523" y="66">orders / searches</text><text class="tb" x="523" y="90">5.6%</text>
  <text class="tb" x="340" y="128" text-anchor="middle">same name, four queries, four numbers</text>
  <line class="grid" x1="6" x2="674" y1="146" y2="146" />
  <text class="h" x="6" y="170">AFTER · ONE DEFINITION, MANY CONSUMERS</text>
  <rect class="box" x="220" y="184" width="240" height="60" rx="6" style="stroke:var(--fig-a);stroke-width:2" />
  <text class="t" x="234" y="208">conversion, v3</text><text class="m" x="234" y="226">one owner, one definition</text>
  <rect class="box" x="60" y="278" width="200" height="50" rx="6" />
  <text class="t" x="72" y="299">Experiment analysis</text><text class="ta" x="72" y="317">3.4%</text>
  <rect class="box" x="420" y="278" width="200" height="50" rx="6" />
  <text class="t" x="432" y="299">Dashboards and reports</text><text class="ta" x="432" y="317">3.4%</text>
  <path class="sa" d="M290,244 L190,276" marker-end="url(#odc1a)" />
  <path class="sa" d="M390,244 L490,276" marker-end="url(#odc1a)" />
  <text class="m" x="340" y="272" text-anchor="middle">same number</text>
</svg>
<figcaption>Illustrative numbers, not real data. The four figures above differ only because the queries differ. Once every consumer reads the same definition, the argument about the number disappears and the argument about the product can start.</figcaption>
</figure>

<p>For an experimentation platform this costs more than for a dashboard. An experiment reports the effect of a change on a metric. If the metric can mean different things in different teams, results are not comparable, and a headline like “this change lifted conversion” cannot be checked against the last one.</p>

<h2 id="what-a-definition-has-to-say">What a definition has to say</h2>

<p>A definition that both a person and a program can rely on has to answer a few questions explicitly. This is the anatomy I’d expect from any semantic layer:</p>

<figure class="fig">
<svg viewBox="0 0 680 306" role="img" aria-labelledby="odc2t odc2d">
  <title id="odc2t">Anatomy of a metric definition</title>
  <desc id="odc2d">Illustrative example of a metric definition with a name, owner, description, grain, numerator, denominator, filters, and version, each with a note on why the field exists.</desc>
  <rect class="box" x="6" y="8" width="420" height="290" rx="6" />
  <text class="m" x="20" y="38">name</text><text class="t" x="110" y="38">checkout_conversion</text>
  <text class="m" x="20" y="72">owner</text><text class="t" x="110" y="72">a named team</text>
  <text class="m" x="20" y="106">description</text><text class="t" x="110" y="106">users who complete an order</text>
  <text class="m" x="20" y="140">grain</text><text class="t" x="110" y="140">per user, per day</text>
  <text class="m" x="20" y="174">numerator</text><text class="t" x="110" y="174">users with a paid order</text>
  <text class="m" x="20" y="208">denominator</text><text class="t" x="110" y="208">users with a session</text>
  <text class="m" x="20" y="242">filters</text><text class="t" x="110" y="242">exclude internal traffic</text>
  <text class="m" x="20" y="276">version</text><text class="t" x="110" y="276">3</text>
  <path class="ln" d="M426,72 L448,72" /><text class="ta" x="454" y="76">who answers questions</text>
  <path class="ln" d="M426,140 L448,140" /><text class="ta" x="454" y="144">what one row means</text>
  <path class="ln" d="M426,191 L448,191" /><text class="ta" x="454" y="195">the arithmetic, exactly</text>
  <path class="ln" d="M426,242 L448,242" /><text class="ta" x="454" y="246">what is left out</text>
  <path class="ln" d="M426,276 L448,276" /><text class="ta" x="454" y="280">what changed, and when</text>
</svg>
<figcaption>An illustrative definition, not a real one. Each field exists because leaving it out is how two teams end up with two numbers.</figcaption>
</figure>

<ul>
  <li><strong>Name and description</strong> in plain language, so a product manager can tell whether it is the metric they want.</li>
  <li><strong>Owner</strong>, a person or a team that answers questions and approves changes.</li>
  <li><strong>Grain</strong>: what one unit of the metric is. Per user, per session and per order give different answers to the same question.</li>
  <li><strong>Numerator, denominator and filters</strong>: the exact arithmetic, including what is excluded.</li>
  <li><strong>Version</strong>, so a change is visible and old results stay interpretable.</li>
</ul>

<p>The schema had to be precise enough for a program to compute from it. That was the job of the query engine: it took a definition and produced the BigQuery query for it, so the same definition gave the same number wherever it was used.</p>

<figure class="fig">
<svg viewBox="0 0 680 190" role="img" aria-labelledby="odc3t odc3d">
  <title id="odc3t">From definition to number</title>
  <desc id="odc3d">Metric definitions are read by a query engine, which generates BigQuery queries. Experiment analysis, dashboards and ad hoc analysis all get their numbers from that one path.</desc>
  <defs><marker id="odc3a" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="arrow" d="M0,0 L10,5 L0,10 z" /></marker></defs>
  <rect class="box" x="6" y="62" width="150" height="64" rx="6" style="stroke:var(--fig-a);stroke-width:2" />
  <text class="t" x="18" y="88">Definitions</text><text class="m" x="18" y="108">the catalogue</text>
  <rect class="box" x="216" y="62" width="150" height="64" rx="6" />
  <text class="t" x="228" y="88">Query engine</text><text class="m" x="228" y="108">definition to SQL</text>
  <rect class="box" x="426" y="62" width="110" height="64" rx="6" />
  <text class="t" x="438" y="88">BigQuery</text><text class="m" x="438" y="108">runs it</text>
  <path class="sa" d="M156,94 L214,94" marker-end="url(#odc3a)" />
  <path class="sa" d="M366,94 L424,94" marker-end="url(#odc3a)" />
  <path class="ln" d="M536,84 L582,44" marker-end="url(#odc3a)" />
  <path class="ln" d="M536,94 L582,94" marker-end="url(#odc3a)" />
  <path class="ln" d="M536,104 L582,144" marker-end="url(#odc3a)" />
  <text class="t" x="590" y="48">Experiments</text>
  <text class="t" x="590" y="98">Dashboards</text>
  <text class="t" x="590" y="148">Ad hoc work</text>
</svg>
<figcaption>Every consumer gets its number through the same path, so there is no second place for a definition to live.</figcaption>
</figure>

<h2 id="ownership-matters-more-than-the-schema">Ownership matters more than the schema</h2>

<p>You can build the schema and engine and still end up with two versions of conversion, one governed and one not. The technical work leaves open the harder question: who decides what “conversion” means?</p>

<p>In my experience the answer has to be a named owner for each metric. Not a committee, and not “the data team” in general. The person who owns a definition can decide when a request to change it is right, and can say no when a team wants a local variant that would fragment the meaning. Without that, the catalogue turns into a list of opinions.</p>

<p>A few practices follow, and they are the ones I’d insist on:</p>

<ol>
  <li><strong>Every metric has an owner in the definition itself.</strong> If you can’t say who owns it, don’t publish it.</li>
  <li><strong>Changes are reviewed and versioned.</strong> A change to a definition changes past comparisons, so it should be visible and deliberate. Old versions stay readable.</li>
  <li><strong>Disagreements get resolved in the definition.</strong> If two teams need different things, that usually means two metrics with two names, not one metric with a silent fork.</li>
  <li><strong>Deprecation is explicit.</strong> A metric nobody owns any more should be marked so, rather than lingering.</li>
</ol>

<h2 id="adoption-is-the-real-test">Adoption is the real test</h2>

<p>A catalogue that nobody uses is documentation. What made the difference to me was that the governed definition had to be the <em>easiest</em> option, not just the correct one. If a team can write its own SQL in ten minutes but must file a request and wait a week to use a shared metric, the local copy wins every time, and it wins for good reasons.</p>

<p>That pointed the design at a few things. Some of these I did and some are what I’d emphasise now, and I’ll keep to the general principle rather than the detail of any one system:</p>

<ul>
  <li><strong>Meet people where they already work.</strong> The natural home for a governed metric was the experimentation workflow, because that is where teams pick the metrics an experiment is judged on. If choosing a shared metric is a step in a task people already do, it costs less than writing a query.</li>
  <li><strong>Cover the metrics people argue about first.</strong> A catalogue with the thirty metrics that cause the most disagreement is worth more than one with three hundred that nobody trusts. The catalogue grew past 100 reusable metrics, and the growth mattered less than whether it contained the ones teams actually reached for.</li>
  <li><strong>Make search and description good.</strong> People pick the metric they can find and understand. A clear name and a plain description save a lot of local reinvention.</li>
  <li><strong>Treat requests as a signal.</strong> When a team wants something the catalogue doesn’t have, either the catalogue has a gap or an existing metric is described badly. Both are fixable, and both are better than letting the team quietly go its own way.</li>
  <li><strong>Make the first use painless.</strong> The moment someone tries a governed metric and gets a number they trust is the moment they stop writing their own. Anything that makes that first attempt fail, such as a confusing name or a missing filter, sends them back to local SQL.</li>
</ul>

<p>Every product team ended up using the catalogue in its experiments. I’d credit that to the definitions being useful and reachable, and I don’t think a mandate would have produced the same result. A rule can make people register a metric. Only a metric that saves them work makes them keep using it.</p>

<h2 id="ownership-needs-a-process-not-just-a-field">Ownership needs a process, not just a field</h2>

<p>An owner field is a start. What it needs behind it is a path for changing a definition that people can follow without heroics. A common shape looks like this:</p>

<ol>
  <li>Someone proposes a change to a definition, with the reason and an example of the difference in the number.</li>
  <li>The owner reviews it, and any team that depends on the metric is told.</li>
  <li>The new version is published alongside the old one for a while, so results can be compared across the change.</li>
  <li>The old version is retired on a date everyone knows.</li>
</ol>

<p>None of this is complicated. The point is that changing what “conversion” means is a decision that affects other people’s results, and the process makes it visible. Otherwise the definition changes silently, and the next dashboard comparison is a debate again.</p>

<h2 id="what-id-watch-for">What I’d watch for</h2>

<p>If I were starting again, I’d check for these early:</p>

<ul>
  <li><strong>Over-modelling.</strong> It is easy to design a schema that can express everything, and then nobody can fill it in. Start with what the common metrics need.</li>
  <li><strong>Definitions without tests.</strong> A definition should be checkable against a known result, so a change to the engine can’t silently change a number.</li>
  <li><strong>Local variants hiding in plain sight.</strong> Watch what people compute outside the layer. It shows you the gaps.</li>
  <li><strong>Ownership drifting.</strong> People move teams. An owner field that is out of date is worse than none, because it looks like governance.</li>
</ul>

<h2 id="what-i-took-from-it">What I took from it</h2>

<p>The schema and the engine mattered, but a shared definition only becomes a working practice through reuse, and reuse depends on people. Someone has to own each definition, changes have to be visible, and the shared version has to be easier to reach than a copy. Get those right and a dashboard comparison goes back to being a conversation about the product.</p>

<h2 id="references">References</h2>

<ol>
  <li>Chang, R. <em>How Airbnb Achieved Metric Consistency at Scale</em>. The Airbnb Tech Blog. <a href="https://medium.com/airbnb-engineering/how-airbnb-achieved-metric-consistency-at-scale-f23cc53dea70">https://medium.com/airbnb-engineering/how-airbnb-achieved-metric-consistency-at-scale-f23cc53dea70</a> — Minerva, a metric platform with one definition per metric</li>
  <li>dbt Labs. <em>dbt Semantic Layer</em>. dbt documentation. <a href="https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl">https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl</a> — centralised metric definitions, reused across tools</li>
  <li>dbt Labs. <em>About MetricFlow</em>. dbt documentation. <a href="https://docs.getdbt.com/docs/build/about-metricflow">https://docs.getdbt.com/docs/build/about-metricflow</a> — generating SQL from metric definitions</li>
  <li>Kohavi, R., Tang, D., and Xu, Y. <em>Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing</em>. Cambridge University Press, 2020. <a href="https://experimentguide.com/">https://experimentguide.com/</a></li>
  <li>Google Cloud. <em>BigQuery overview</em>. Google Cloud documentation. <a href="https://docs.cloud.google.com/bigquery/docs/introduction">https://docs.cloud.google.com/bigquery/docs/introduction</a></li>
</ol>]]></content><author><name>Fadhil Mochammad</name></author><category term="exp" /><category term="data" /><summary type="html"><![CDATA[Two dashboards show a metric called “conversion” for the same week. The numbers differ. Someone asks which one is right, and the next hour goes on reading two pieces of SQL side by side. By the end nobody has learned anything about the product. The conversation has become a debate about queries.]]></summary></entry><entry><title type="html">The exposure triangle, but for learning rates</title><link href="https://fadhilmch.github.io/posts/the-exposure-triangle-but-for-learning-rates/" rel="alternate" type="text/html" title="The exposure triangle, but for learning rates" /><published>2025-10-18T00:00:00+02:00</published><updated>2025-10-18T00:00:00+02:00</updated><id>https://fadhilmch.github.io/posts/the-exposure-triangle-but-for-learning-rates</id><content type="html" xml:base="https://fadhilmch.github.io/posts/the-exposure-triangle-but-for-learning-rates/"><![CDATA[<p>When I take photos, I keep running into the same annoyance. I open the aperture to let in more light, and the background goes soft. I slow the shutter to let in more light, and a moving subject smears. I raise the ISO to get more light without either, and the picture gets grainy. Nothing is free.</p>

<p>Photographers call this the <strong>exposure triangle</strong>. Neural-network training has its own version of that problem, with three knobs that also pull against each other: learning rate, batch size and warmup. This post lays the two side by side, shows what the training knobs do in a small simulation, and spends a fair amount of time on where the comparison stops being useful.</p>

<h2 id="the-photography-version">The photography version</h2>

<p>A camera sensor needs a certain amount of light to produce a well-exposed picture. Three controls set how much it gets, and each one changes something else about the image:</p>

<ul>
  <li><strong>Aperture</strong> is the size of the opening in the lens. A wide opening lets in more light but keeps only a thin slice of the scene in focus, which is the shallow depth of field that blurs backgrounds.</li>
  <li><strong>Shutter speed</strong> is how long the sensor collects light. A long exposure gathers more, but anything that moves during it, including your own hands, turns into blur.</li>
  <li><strong>ISO</strong> is how strongly the sensor’s signal is amplified. Raising it brightens a dim scene without changing the opening or the exposure time, but it amplifies noise along with the signal, so the picture gets grainy.</li>
</ul>

<figure class="fig">
<svg viewBox="0 0 680 300" role="img" aria-labelledby="ex1t ex1d">
  <title id="ex1t">The exposure triangle</title>
  <desc id="ex1d">A triangle with aperture, shutter speed and ISO at its corners. Aperture controls light and depth of field, and a wide opening gives shallow focus. Shutter speed controls how long light is collected, and slow speeds give motion blur. ISO amplifies the signal, and high values add noise. Each knob has a cost in orange.</desc>
  <path class="ln" d="M115,40 L20,190 L210,190 Z" style="fill:var(--panel)" />
  <circle class="fa" cx="115" cy="40" r="6" /><text class="t" x="115" y="24" text-anchor="middle">Aperture</text>
  <circle class="fa" cx="20" cy="190" r="6" /><text class="t" x="20" y="214" text-anchor="start">Shutter</text>
  <circle class="fa" cx="210" cy="190" r="6" /><text class="t" x="210" y="214" text-anchor="end">ISO</text>
  <text class="m" x="115" y="244" text-anchor="middle">one brightness budget,</text><text class="m" x="115" y="260" text-anchor="middle">three ways to spend it</text>
  <rect class="box" x="250" y="10" width="420" height="82" rx="6" />
  <text class="t" x="264" y="32">Aperture · size of the opening</text>
  <text class="m" x="264" y="52">more light through, or less</text>
  <text class="tb" x="264" y="74">cost: wide gives a thin slice of focus</text>
  <rect class="box" x="250" y="104" width="420" height="82" rx="6" />
  <text class="t" x="264" y="126">Shutter · how long the sensor collects</text>
  <text class="m" x="264" y="146">longer means more light</text>
  <text class="tb" x="264" y="168">cost: slow gives motion blur, shake</text>
  <rect class="box" x="250" y="198" width="420" height="82" rx="6" />
  <text class="t" x="264" y="220">ISO · amplification after the fact</text>
  <text class="m" x="264" y="240">brightens a dim scene</text>
  <text class="tb" x="264" y="262">cost: more grain and noise</text>
</svg>
<figcaption>Every stop of brightness has to come from somewhere. You choose which cost you would rather pay.</figcaption>
</figure>

<p>Brightness is measured in <strong>stops</strong>, and each stop is a doubling or halving of the light. That gives the triangle its useful property: the controls trade against each other one stop at a time. Open the aperture by one stop and you can halve the exposure time for the same brightness. The picture is equally bright, but you’ve swapped a shallower focus for less motion blur.</p>

<p>The point isn’t that any corner is better. The right setting depends on what you care about in the scene, and a good photographer decides that first.</p>

<h2 id="the-training-version">The training version</h2>

<p>Training a neural network has a comparable shape. You have a budget for how much progress each step can make safely, and three settings that spend it:</p>

<ul>
  <li><strong>Learning rate</strong> sets how far each parameter update moves. Too large and the loss oscillates or diverges. Too small and training crawls.</li>
  <li><strong>Batch size</strong> is how many examples the gradient is averaged over at each step. A small batch gives a noisy estimate of the direction. A large one gives a cleaner estimate but costs more memory and compute for each step.</li>
  <li><strong>Warmup</strong> means starting with a very small learning rate and ramping it up over the first steps. It keeps the early updates gentle, while the model and the optimiser’s statistics are still in a poor state, so that you can use a bigger learning rate afterwards.</li>
</ul>

<figure class="fig">
<svg viewBox="0 0 680 300" role="img" aria-labelledby="ex2t ex2d">
  <title id="ex2t">The training triangle</title>
  <desc id="ex2d">A triangle with learning rate, batch size and warmup at its corners, matched to aperture, shutter and ISO. Learning rate sets how far each update moves and a large one can diverge. Batch size sets how many examples each gradient averages, and small batches are noisy. Warmup ramps the learning rate up from a small value; it costs a few slow early steps. Each cost is in orange.</desc>
  <path class="ln" d="M115,40 L20,190 L210,190 Z" style="fill:var(--panel)" />
  <circle class="fa" cx="115" cy="40" r="6" /><text class="t" x="115" y="24" text-anchor="middle">Learning rate</text>
  <circle class="fa" cx="20" cy="190" r="6" /><text class="t" x="20" y="214" text-anchor="start">Batch size</text>
  <circle class="fa" cx="210" cy="190" r="6" /><text class="t" x="210" y="214" text-anchor="end">Warmup</text>
  <text class="m" x="115" y="244" text-anchor="middle">one stability budget,</text><text class="m" x="115" y="260" text-anchor="middle">three ways to spend it</text>
  <rect class="box" x="250" y="10" width="420" height="82" rx="6" />
  <text class="t" x="264" y="32">Learning rate ~ aperture</text>
  <text class="m" x="264" y="52">how far each update moves</text>
  <text class="tb" x="264" y="74">cost: too big diverges, too small crawls</text>
  <rect class="box" x="250" y="104" width="420" height="82" rx="6" />
  <text class="t" x="264" y="126">Batch size ~ shutter</text>
  <text class="m" x="264" y="146">how many examples each gradient averages</text>
  <text class="tb" x="264" y="168">cost: small is noisy, large costs compute</text>
  <rect class="box" x="250" y="198" width="420" height="82" rx="6" />
  <text class="t" x="264" y="220">Warmup ~ ISO</text>
  <text class="m" x="264" y="240">a gentle start that lets you push harder later</text>
  <text class="tb" x="264" y="262">cost: slow early steps, one more knob</text>
</svg>
<figcaption>Each photographic knob is paired with a training knob. The first two pairs are fairly natural. The third is looser, and I say why below.</figcaption>
</figure>

<p>The pairing I’m using is aperture with learning rate, shutter with batch size, and ISO with warmup. I’ll explain each pair in turn, and mark how strong I think it is.</p>

<p><strong>Aperture and learning rate (strong).</strong> Both are the main lever. They decide how big a step you take per unit of effort, and both have a sweet spot between too little and too much.</p>

<p><strong>Shutter and batch size (fairly strong).</strong> A long exposure collects more photons, which averages out random fluctuation. A big batch collects more examples, which averages out sampling noise. In both, more collection reduces noise and costs you something: time on the camera side, compute on the training side.</p>

<p><strong>ISO and warmup (loose).</strong> This is the pair I’m least sure about. ISO is a late-stage amplifier: you reach for it when the other two are already stretched. Warmup plays a similar supporting role, because it’s a modest, cheap adjustment that lets the other two be pushed further. But ISO amplifies noise, and warmup does the opposite: it protects you from instability. I’d call this a pairing of roles, not of mechanisms.</p>

<h2 id="what-the-knobs-do-in-a-simulation">What the knobs do, in a simulation</h2>

<p>Words are easy here, so I ran the experiment. The setup is deliberately tiny: gradient descent on a two-dimensional quadratic bowl, with random noise added to each gradient to stand in for sampling a mini-batch. It’s pure Python, and it’s illustrative rather than a real model.</p>

<figure class="fig">
<svg viewBox="0 0 680 262" role="img" aria-labelledby="lr3t lr3d">
  <title id="lr3t">Simulated loss curves for different learning rates and batch sizes</title>
  <desc id="lr3d">Two log-scale loss charts over 80 gradient descent steps on a noisy two-dimensional quadratic. Left, at batch size 64: a learning rate of 0.005 falls slowly and is still around 2 at step 80; 0.10 drops to about 0.09 by step 20 and then sits at a noise floor near 0.003; 0.22 blows up and leaves the chart within a few steps. Right, at learning rate 0.10: batch size 1 settles at a noisy floor near 0.2, batch size 8 near 0.03, and batch size 64 near 0.003.</desc>
  <text class="h" x="50" y="20">LEARNING RATE · BATCH 64</text><text class="h" x="390" y="20">BATCH SIZE · LEARNING RATE 0.10</text>
  <line class="grid" x1="50" x2="320" y1="208.0" y2="208.0" /><text class="m" x="44" y="212.0" text-anchor="end">1e-4</text>
  <line class="grid" x1="50" x2="320" y1="165.1" y2="165.1" /><text class="m" x="44" y="169.1" text-anchor="end">1e-2</text>
  <line class="grid" x1="50" x2="320" y1="122.3" y2="122.3" /><text class="m" x="44" y="126.3" text-anchor="end">1</text>
  <line class="grid" x1="50" x2="320" y1="79.4" y2="79.4" /><text class="m" x="44" y="83.4" text-anchor="end">1e2</text>
  <text class="m" x="50.0" y="224" text-anchor="middle">0</text>
  <text class="m" x="185.0" y="224" text-anchor="middle">40</text>
  <text class="m" x="320.0" y="224" text-anchor="middle">80</text>
  <text class="m" x="320" y="240" text-anchor="end">step</text>
  <g tabindex="0"><title>lr 0.005 (small): loss 49.5 at step 0, 2.03 at step 80</title><path class="ln" d="M50.0,86.0 L53.4,86.8 L56.8,87.7 L60.1,88.6 L63.5,89.4 L66.9,90.3 L70.2,91.1 L73.6,91.9 L77.0,92.7 L80.4,93.5 L83.8,94.3 L87.1,95.1 L90.5,95.8 L93.9,96.6 L97.2,97.3 L100.6,98.1 L104.0,98.8 L107.4,99.4 L110.8,100.1 L114.1,100.7 L117.5,101.4 L120.9,102.0 L124.2,102.6 L127.6,103.1 L131.0,103.7 L134.4,104.2 L137.8,104.7 L141.1,105.2 L144.5,105.7 L147.9,106.2 L151.2,106.6 L154.6,107.0 L158.0,107.4 L161.4,107.7 L164.8,108.1 L168.1,108.4 L171.5,108.8 L174.9,109.0 L178.2,109.3 L181.6,109.6 L185.0,109.9 L188.4,110.2 L191.8,110.4 L195.1,110.7 L198.5,110.9 L201.9,111.1 L205.2,111.3 L208.6,111.5 L212.0,111.7 L215.4,111.9 L218.8,112.0 L222.1,112.2 L225.5,112.4 L228.9,112.5 L232.2,112.7 L235.6,112.8 L239.0,113.0 L242.4,113.1 L245.8,113.2 L249.1,113.4 L252.5,113.5 L255.9,113.6 L259.2,113.8 L262.6,113.9 L266.0,114.0 L269.4,114.1 L272.8,114.2 L276.1,114.4 L279.5,114.5 L282.9,114.6 L286.2,114.7 L289.6,114.8 L293.0,114.9 L296.4,115.0 L299.8,115.1 L303.1,115.2 L306.5,115.3 L309.9,115.5 L313.2,115.5 L316.6,115.6 L320.0,115.7" /><path d="M50.0,86.0 L53.4,86.8 L56.8,87.7 L60.1,88.6 L63.5,89.4 L66.9,90.3 L70.2,91.1 L73.6,91.9 L77.0,92.7 L80.4,93.5 L83.8,94.3 L87.1,95.1 L90.5,95.8 L93.9,96.6 L97.2,97.3 L100.6,98.1 L104.0,98.8 L107.4,99.4 L110.8,100.1 L114.1,100.7 L117.5,101.4 L120.9,102.0 L124.2,102.6 L127.6,103.1 L131.0,103.7 L134.4,104.2 L137.8,104.7 L141.1,105.2 L144.5,105.7 L147.9,106.2 L151.2,106.6 L154.6,107.0 L158.0,107.4 L161.4,107.7 L164.8,108.1 L168.1,108.4 L171.5,108.8 L174.9,109.0 L178.2,109.3 L181.6,109.6 L185.0,109.9 L188.4,110.2 L191.8,110.4 L195.1,110.7 L198.5,110.9 L201.9,111.1 L205.2,111.3 L208.6,111.5 L212.0,111.7 L215.4,111.9 L218.8,112.0 L222.1,112.2 L225.5,112.4 L228.9,112.5 L232.2,112.7 L235.6,112.8 L239.0,113.0 L242.4,113.1 L245.8,113.2 L249.1,113.4 L252.5,113.5 L255.9,113.6 L259.2,113.8 L262.6,113.9 L266.0,114.0 L269.4,114.1 L272.8,114.2 L276.1,114.4 L279.5,114.5 L282.9,114.6 L286.2,114.7 L289.6,114.8 L293.0,114.9 L296.4,115.0 L299.8,115.1 L303.1,115.2 L306.5,115.3 L309.9,115.5 L313.2,115.5 L316.6,115.6 L320.0,115.7" style="stroke:transparent;stroke-width:12;fill:none" /></g>
  <g tabindex="0"><title>lr 0.10 (good): loss 49.5 at step 0, 0.000764 at step 80</title><path class="sa" d="M50.0,86.0 L53.4,110.2 L56.8,112.1 L60.1,113.9 L63.5,116.1 L66.9,118.3 L70.2,120.4 L73.6,121.8 L77.0,124.0 L80.4,125.2 L83.8,126.9 L87.1,129.0 L90.5,131.2 L93.9,133.3 L97.2,134.5 L100.6,137.0 L104.0,138.6 L107.4,140.4 L110.8,142.9 L114.1,144.1 L117.5,145.2 L120.9,146.7 L124.2,148.2 L127.6,150.2 L131.0,150.0 L134.4,151.4 L137.8,154.7 L141.1,153.3 L144.5,156.1 L147.9,162.6 L151.2,161.3 L154.6,168.1 L158.0,164.1 L161.4,166.6 L164.8,160.7 L168.1,169.9 L171.5,176.6 L174.9,156.6 L178.2,171.3 L181.6,166.9 L185.0,179.2 L188.4,180.3 L191.8,184.7 L195.1,187.0 L198.5,170.4 L201.9,187.6 L205.2,177.1 L208.6,164.6 L212.0,174.4 L215.4,165.2 L218.8,187.8 L222.1,181.6 L225.5,172.9 L228.9,184.8 L232.2,208.0 L235.6,175.8 L239.0,190.1 L242.4,196.3 L245.8,191.2 L249.1,175.3 L252.5,170.0 L255.9,183.3 L259.2,175.4 L262.6,184.4 L266.0,179.3 L269.4,186.7 L272.8,184.8 L276.1,176.3 L279.5,169.8 L282.9,173.4 L286.2,170.5 L289.6,178.8 L293.0,156.4 L296.4,179.5 L299.8,178.5 L303.1,177.9 L306.5,179.4 L309.9,169.8 L313.2,175.2 L316.6,178.6 L320.0,189.1" /><path d="M50.0,86.0 L53.4,110.2 L56.8,112.1 L60.1,113.9 L63.5,116.1 L66.9,118.3 L70.2,120.4 L73.6,121.8 L77.0,124.0 L80.4,125.2 L83.8,126.9 L87.1,129.0 L90.5,131.2 L93.9,133.3 L97.2,134.5 L100.6,137.0 L104.0,138.6 L107.4,140.4 L110.8,142.9 L114.1,144.1 L117.5,145.2 L120.9,146.7 L124.2,148.2 L127.6,150.2 L131.0,150.0 L134.4,151.4 L137.8,154.7 L141.1,153.3 L144.5,156.1 L147.9,162.6 L151.2,161.3 L154.6,168.1 L158.0,164.1 L161.4,166.6 L164.8,160.7 L168.1,169.9 L171.5,176.6 L174.9,156.6 L178.2,171.3 L181.6,166.9 L185.0,179.2 L188.4,180.3 L191.8,184.7 L195.1,187.0 L198.5,170.4 L201.9,187.6 L205.2,177.1 L208.6,164.6 L212.0,174.4 L215.4,165.2 L218.8,187.8 L222.1,181.6 L225.5,172.9 L228.9,184.8 L232.2,208.0 L235.6,175.8 L239.0,190.1 L242.4,196.3 L245.8,191.2 L249.1,175.3 L252.5,170.0 L255.9,183.3 L259.2,175.4 L262.6,184.4 L266.0,179.3 L269.4,186.7 L272.8,184.8 L276.1,176.3 L279.5,169.8 L282.9,173.4 L286.2,170.5 L289.6,178.8 L293.0,156.4 L296.4,179.5 L299.8,178.5 L303.1,177.9 L306.5,179.4 L309.9,169.8 L313.2,175.2 L316.6,178.6 L320.0,189.1" style="stroke:transparent;stroke-width:12;fill:none" /></g>
  <g tabindex="0"><title>lr 0.22 (too large): loss 49.5 at step 0, 1e+06 at step 80</title><path class="sb" d="M50.0,86.0 L53.4,82.9 L56.8,79.7 L60.1,76.4 L63.5,73.1 L66.9,69.7 L70.2,66.4 L73.6,62.9 L77.0,59.5 L80.4,58.0 L83.8,58.0 L87.1,58.0 L90.5,58.0 L93.9,58.0 L97.2,58.0 L100.6,58.0 L104.0,58.0 L107.4,58.0 L110.8,58.0 L114.1,58.0 L117.5,58.0 L120.9,58.0 L124.2,58.0 L127.6,58.0 L131.0,58.0 L134.4,58.0 L137.8,58.0 L141.1,58.0 L144.5,58.0 L147.9,58.0 L151.2,58.0 L154.6,58.0 L158.0,58.0 L161.4,58.0 L164.8,58.0 L168.1,58.0 L171.5,58.0 L174.9,58.0 L178.2,58.0 L181.6,58.0 L185.0,58.0 L188.4,58.0 L191.8,58.0 L195.1,58.0 L198.5,58.0 L201.9,58.0 L205.2,58.0 L208.6,58.0 L212.0,58.0 L215.4,58.0 L218.8,58.0 L222.1,58.0 L225.5,58.0 L228.9,58.0 L232.2,58.0 L235.6,58.0 L239.0,58.0 L242.4,58.0 L245.8,58.0 L249.1,58.0 L252.5,58.0 L255.9,58.0 L259.2,58.0 L262.6,58.0 L266.0,58.0 L269.4,58.0 L272.8,58.0 L276.1,58.0 L279.5,58.0 L282.9,58.0 L286.2,58.0 L289.6,58.0 L293.0,58.0 L296.4,58.0 L299.8,58.0 L303.1,58.0 L306.5,58.0 L309.9,58.0 L313.2,58.0 L316.6,58.0 L320.0,58.0" /><path d="M50.0,86.0 L53.4,82.9 L56.8,79.7 L60.1,76.4 L63.5,73.1 L66.9,69.7 L70.2,66.4 L73.6,62.9 L77.0,59.5 L80.4,58.0 L83.8,58.0 L87.1,58.0 L90.5,58.0 L93.9,58.0 L97.2,58.0 L100.6,58.0 L104.0,58.0 L107.4,58.0 L110.8,58.0 L114.1,58.0 L117.5,58.0 L120.9,58.0 L124.2,58.0 L127.6,58.0 L131.0,58.0 L134.4,58.0 L137.8,58.0 L141.1,58.0 L144.5,58.0 L147.9,58.0 L151.2,58.0 L154.6,58.0 L158.0,58.0 L161.4,58.0 L164.8,58.0 L168.1,58.0 L171.5,58.0 L174.9,58.0 L178.2,58.0 L181.6,58.0 L185.0,58.0 L188.4,58.0 L191.8,58.0 L195.1,58.0 L198.5,58.0 L201.9,58.0 L205.2,58.0 L208.6,58.0 L212.0,58.0 L215.4,58.0 L218.8,58.0 L222.1,58.0 L225.5,58.0 L228.9,58.0 L232.2,58.0 L235.6,58.0 L239.0,58.0 L242.4,58.0 L245.8,58.0 L249.1,58.0 L252.5,58.0 L255.9,58.0 L259.2,58.0 L262.6,58.0 L266.0,58.0 L269.4,58.0 L272.8,58.0 L276.1,58.0 L279.5,58.0 L282.9,58.0 L286.2,58.0 L289.6,58.0 L293.0,58.0 L296.4,58.0 L299.8,58.0 L303.1,58.0 L306.5,58.0 L309.9,58.0 L313.2,58.0 L316.6,58.0 L320.0,58.0" style="stroke:transparent;stroke-width:12;fill:none" /></g>
  <line class="grid" x1="390" x2="660" y1="208.0" y2="208.0" /><text class="m" x="384" y="212.0" text-anchor="end">1e-4</text>
  <line class="grid" x1="390" x2="660" y1="165.1" y2="165.1" /><text class="m" x="384" y="169.1" text-anchor="end">1e-2</text>
  <line class="grid" x1="390" x2="660" y1="122.3" y2="122.3" /><text class="m" x="384" y="126.3" text-anchor="end">1</text>
  <line class="grid" x1="390" x2="660" y1="79.4" y2="79.4" /><text class="m" x="384" y="83.4" text-anchor="end">1e2</text>
  <text class="m" x="390.0" y="224" text-anchor="middle">0</text>
  <text class="m" x="525.0" y="224" text-anchor="middle">40</text>
  <text class="m" x="660.0" y="224" text-anchor="middle">80</text>
  <text class="m" x="660" y="240" text-anchor="end">step</text>
  <g tabindex="0"><title>batch 1: loss 49.5 at step 0, 0.0486 at step 80</title><path class="sb" d="M390.0,86.0 L393.4,109.8 L396.8,111.5 L400.1,112.0 L403.5,115.8 L406.9,120.1 L410.2,123.2 L413.6,119.6 L417.0,123.5 L420.4,117.2 L423.8,120.5 L427.1,123.7 L430.5,126.1 L433.9,129.9 L437.2,121.9 L440.6,127.6 L444.0,130.3 L447.4,133.3 L450.8,138.3 L454.1,131.6 L457.5,128.5 L460.9,133.7 L464.2,126.8 L467.6,129.2 L471.0,130.8 L474.4,129.6 L477.8,137.5 L481.1,128.7 L484.5,131.3 L487.9,147.5 L491.2,131.7 L494.6,146.0 L498.0,132.5 L501.4,142.4 L504.8,124.1 L508.1,147.2 L511.5,147.4 L514.9,119.5 L518.2,140.0 L521.6,133.3 L525.0,155.8 L528.4,150.6 L531.8,144.2 L535.1,143.9 L538.5,132.6 L541.9,147.7 L545.2,142.2 L548.6,126.2 L552.0,136.6 L555.4,127.3 L558.8,155.8 L562.1,144.4 L565.5,134.6 L568.9,148.9 L572.2,198.5 L575.6,137.6 L579.0,150.1 L582.4,162.1 L585.8,154.8 L589.1,136.4 L592.5,131.0 L595.9,144.1 L599.2,136.2 L602.6,144.5 L606.0,140.0 L609.4,147.0 L612.8,145.2 L616.1,137.2 L619.5,130.7 L622.9,134.3 L626.2,131.6 L629.6,139.7 L633.0,117.6 L636.4,140.5 L639.8,139.6 L643.1,139.0 L646.5,140.5 L649.9,130.9 L653.2,136.3 L656.6,139.8 L660.0,150.4" /><path d="M390.0,86.0 L393.4,109.8 L396.8,111.5 L400.1,112.0 L403.5,115.8 L406.9,120.1 L410.2,123.2 L413.6,119.6 L417.0,123.5 L420.4,117.2 L423.8,120.5 L427.1,123.7 L430.5,126.1 L433.9,129.9 L437.2,121.9 L440.6,127.6 L444.0,130.3 L447.4,133.3 L450.8,138.3 L454.1,131.6 L457.5,128.5 L460.9,133.7 L464.2,126.8 L467.6,129.2 L471.0,130.8 L474.4,129.6 L477.8,137.5 L481.1,128.7 L484.5,131.3 L487.9,147.5 L491.2,131.7 L494.6,146.0 L498.0,132.5 L501.4,142.4 L504.8,124.1 L508.1,147.2 L511.5,147.4 L514.9,119.5 L518.2,140.0 L521.6,133.3 L525.0,155.8 L528.4,150.6 L531.8,144.2 L535.1,143.9 L538.5,132.6 L541.9,147.7 L545.2,142.2 L548.6,126.2 L552.0,136.6 L555.4,127.3 L558.8,155.8 L562.1,144.4 L565.5,134.6 L568.9,148.9 L572.2,198.5 L575.6,137.6 L579.0,150.1 L582.4,162.1 L585.8,154.8 L589.1,136.4 L592.5,131.0 L595.9,144.1 L599.2,136.2 L602.6,144.5 L606.0,140.0 L609.4,147.0 L612.8,145.2 L616.1,137.2 L619.5,130.7 L622.9,134.3 L626.2,131.6 L629.6,139.7 L633.0,117.6 L636.4,140.5 L639.8,139.6 L643.1,139.0 L646.5,140.5 L649.9,130.9 L653.2,136.3 L656.6,139.8 L660.0,150.4" style="stroke:transparent;stroke-width:12;fill:none" /></g>
  <g tabindex="0"><title>batch 8: loss 49.5 at step 0, 0.00609 at step 80</title><path class="ln" d="M390.0,86.0 L393.4,110.1 L396.8,112.0 L400.1,113.4 L403.5,116.0 L406.9,118.8 L410.2,121.1 L413.6,121.4 L417.0,123.9 L420.4,123.3 L423.8,125.0 L427.1,127.4 L430.5,129.9 L433.9,132.5 L437.2,131.4 L440.6,135.0 L444.0,136.5 L447.4,138.2 L450.8,141.7 L454.1,140.8 L457.5,140.2 L460.9,142.4 L464.2,140.9 L467.6,143.2 L471.0,142.9 L474.4,143.1 L477.8,148.5 L481.1,143.4 L484.5,146.3 L487.9,157.9 L491.2,149.2 L494.6,161.5 L498.0,150.6 L501.4,157.2 L504.8,143.2 L508.1,161.0 L511.5,165.5 L514.9,138.5 L518.2,157.5 L521.6,151.3 L525.0,170.0 L528.4,167.7 L531.8,164.6 L535.1,164.7 L538.5,151.8 L541.9,167.9 L545.2,160.6 L548.6,145.5 L552.0,155.8 L555.4,146.5 L558.8,173.2 L562.1,163.4 L565.5,153.9 L568.9,167.5 L572.2,208.0 L575.6,156.9 L579.0,169.8 L582.4,180.2 L585.8,173.6 L589.1,155.8 L592.5,150.4 L595.9,163.6 L599.2,155.7 L602.6,164.2 L606.0,159.5 L609.4,166.6 L612.8,164.8 L616.1,156.7 L619.5,150.2 L622.9,153.8 L626.2,151.0 L629.6,159.1 L633.0,137.0 L636.4,159.9 L639.8,159.0 L643.1,158.4 L646.5,159.9 L649.9,150.3 L653.2,155.7 L656.6,159.2 L660.0,169.8" /><path d="M390.0,86.0 L393.4,110.1 L396.8,112.0 L400.1,113.4 L403.5,116.0 L406.9,118.8 L410.2,121.1 L413.6,121.4 L417.0,123.9 L420.4,123.3 L423.8,125.0 L427.1,127.4 L430.5,129.9 L433.9,132.5 L437.2,131.4 L440.6,135.0 L444.0,136.5 L447.4,138.2 L450.8,141.7 L454.1,140.8 L457.5,140.2 L460.9,142.4 L464.2,140.9 L467.6,143.2 L471.0,142.9 L474.4,143.1 L477.8,148.5 L481.1,143.4 L484.5,146.3 L487.9,157.9 L491.2,149.2 L494.6,161.5 L498.0,150.6 L501.4,157.2 L504.8,143.2 L508.1,161.0 L511.5,165.5 L514.9,138.5 L518.2,157.5 L521.6,151.3 L525.0,170.0 L528.4,167.7 L531.8,164.6 L535.1,164.7 L538.5,151.8 L541.9,167.9 L545.2,160.6 L548.6,145.5 L552.0,155.8 L555.4,146.5 L558.8,173.2 L562.1,163.4 L565.5,153.9 L568.9,167.5 L572.2,208.0 L575.6,156.9 L579.0,169.8 L582.4,180.2 L585.8,173.6 L589.1,155.8 L592.5,150.4 L595.9,163.6 L599.2,155.7 L602.6,164.2 L606.0,159.5 L609.4,166.6 L612.8,164.8 L616.1,156.7 L619.5,150.2 L622.9,153.8 L626.2,151.0 L629.6,159.1 L633.0,137.0 L636.4,159.9 L639.8,159.0 L643.1,158.4 L646.5,159.9 L649.9,150.3 L653.2,155.7 L656.6,159.2 L660.0,169.8" style="stroke:transparent;stroke-width:12;fill:none" /></g>
  <g tabindex="0"><title>batch 64: loss 49.5 at step 0, 0.000764 at step 80</title><path class="sa" d="M390.0,86.0 L393.4,110.2 L396.8,112.1 L400.1,113.9 L403.5,116.1 L406.9,118.3 L410.2,120.4 L413.6,121.8 L417.0,124.0 L420.4,125.2 L423.8,126.9 L427.1,129.0 L430.5,131.2 L433.9,133.3 L437.2,134.5 L440.6,137.0 L444.0,138.6 L447.4,140.4 L450.8,142.9 L454.1,144.1 L457.5,145.2 L460.9,146.7 L464.2,148.2 L467.6,150.2 L471.0,150.0 L474.4,151.4 L477.8,154.7 L481.1,153.3 L484.5,156.1 L487.9,162.6 L491.2,161.3 L494.6,168.1 L498.0,164.1 L501.4,166.6 L504.8,160.7 L508.1,169.9 L511.5,176.6 L514.9,156.6 L518.2,171.3 L521.6,166.9 L525.0,179.2 L528.4,180.3 L531.8,184.7 L535.1,187.0 L538.5,170.4 L541.9,187.6 L545.2,177.1 L548.6,164.6 L552.0,174.4 L555.4,165.2 L558.8,187.8 L562.1,181.6 L565.5,172.9 L568.9,184.8 L572.2,208.0 L575.6,175.8 L579.0,190.1 L582.4,196.3 L585.8,191.2 L589.1,175.3 L592.5,170.0 L595.9,183.3 L599.2,175.4 L602.6,184.4 L606.0,179.3 L609.4,186.7 L612.8,184.8 L616.1,176.3 L619.5,169.8 L622.9,173.4 L626.2,170.5 L629.6,178.8 L633.0,156.4 L636.4,179.5 L639.8,178.5 L643.1,177.9 L646.5,179.4 L649.9,169.8 L653.2,175.2 L656.6,178.6 L660.0,189.1" /><path d="M390.0,86.0 L393.4,110.2 L396.8,112.1 L400.1,113.9 L403.5,116.1 L406.9,118.3 L410.2,120.4 L413.6,121.8 L417.0,124.0 L420.4,125.2 L423.8,126.9 L427.1,129.0 L430.5,131.2 L433.9,133.3 L437.2,134.5 L440.6,137.0 L444.0,138.6 L447.4,140.4 L450.8,142.9 L454.1,144.1 L457.5,145.2 L460.9,146.7 L464.2,148.2 L467.6,150.2 L471.0,150.0 L474.4,151.4 L477.8,154.7 L481.1,153.3 L484.5,156.1 L487.9,162.6 L491.2,161.3 L494.6,168.1 L498.0,164.1 L501.4,166.6 L504.8,160.7 L508.1,169.9 L511.5,176.6 L514.9,156.6 L518.2,171.3 L521.6,166.9 L525.0,179.2 L528.4,180.3 L531.8,184.7 L535.1,187.0 L538.5,170.4 L541.9,187.6 L545.2,177.1 L548.6,164.6 L552.0,174.4 L555.4,165.2 L558.8,187.8 L562.1,181.6 L565.5,172.9 L568.9,184.8 L572.2,208.0 L575.6,175.8 L579.0,190.1 L582.4,196.3 L585.8,191.2 L589.1,175.3 L592.5,170.0 L595.9,183.3 L599.2,175.4 L602.6,184.4 L606.0,179.3 L609.4,186.7 L612.8,184.8 L616.1,176.3 L619.5,169.8 L622.9,173.4 L626.2,170.5 L629.6,178.8 L633.0,156.4 L636.4,179.5 L639.8,178.5 L643.1,177.9 L646.5,179.4 L649.9,169.8 L653.2,175.2 L656.6,178.6 L660.0,189.1" style="stroke:transparent;stroke-width:12;fill:none" /></g>
  <line class="ln" x1="50" x2="66" y1="34" y2="34" /><text class="m" x="71" y="38">0.005</text>
  <line class="sa" x1="116" x2="132" y1="34" y2="34" /><text class="m" x="137" y="38">0.10</text>
  <line class="sb" x1="175" x2="191" y1="34" y2="34" /><text class="m" x="196" y="38">0.22 diverges</text>
  <line class="sb" x1="390" x2="406" y1="34" y2="34" /><text class="m" x="411" y="38">batch 1</text>
  <line class="ln" x1="469" x2="485" y1="34" y2="34" /><text class="m" x="490" y="38">batch 8</text>
  <line class="sa" x1="548" x2="564" y1="34" y2="34" /><text class="m" x="569" y="38">batch 64</text>
</svg>
<figcaption>Simulated, not from a real model: gradient descent on a two-dimensional quadratic with added gradient noise. The largest learning rate that still converges here is 0.2, so 0.22 diverges. Smaller batches average less noise, so the loss stalls higher.</figcaption>
</figure>

<p>On the left, batch size is fixed and only the learning rate changes:</p>

<ul>
  <li><strong>0.005</strong> is safe but slow. After 80 steps it has still only reached a loss of about 2.</li>
  <li><strong>0.10</strong> gets to a low loss in about 20 steps and then hovers at a small noise floor.</li>
  <li><strong>0.22</strong> is only a little larger, and it diverges. In this bowl, the steepest direction sets a hard limit: any learning rate above 2 divided by that direction’s curvature (here 10, so 0.2) makes each step overshoot by more than it corrects.</li>
</ul>

<p>On the right, learning rate is fixed and the batch size changes. The averaging effect is visible. A batch of one gives a noisy trace that settles about two orders of magnitude above where a batch of 64 does. Averaging over more examples shrinks the noise floor, roughly in proportion to one over the batch size.</p>

<p>That last sentence hides the coupling. The height of the noise floor depends on both the learning rate and the batch size, in this simple setting roughly as their ratio. Double the learning rate and you raise the floor. Double the batch and you lower it. That’s the same kind of trade as opening the aperture by a stop and shortening the exposure by a stop, and it’s where the analogy earns its keep.</p>

<h2 id="the-trade-that-carries-over">The trade that carries over</h2>

<p>There is a well-known heuristic from large-batch training that has the same shape as the photographer’s reciprocity. If you multiply the batch size by k, multiply the learning rate by about k as well, and add a warmup phase so that the early steps don’t blow up. It’s often called the linear scaling rule, and it comes from work on training image models with very large batches. It isn’t a law. It works up to some batch size and then stops working, and for adaptive optimisers such as Adam, people often find a square-root scaling fits better.</p>

<p>I still find it worth remembering, because it shows why treating the knobs as independent is a mistake. If you double the batch size to “make training more stable” and leave the learning rate alone, you have changed the operating point. The run may get smoother and slower. If you double the learning rate to “make it faster” and leave everything else alone, you have moved towards the edge in the left panel.</p>

<p>A habit I’d suggest, borrowed from the photographers: when you change one control, name the other one you’re spending. “I’m raising the learning rate, so I’m accepting more noise and less margin before divergence.” “I’m shrinking the batch to fit on the device, so I’ll expect a noisier curve and probably a lower learning rate.” Saying it out loud makes the trade visible.</p>

<h2 id="where-the-analogy-breaks">Where the analogy breaks</h2>

<p>I’ve pushed the comparison as far as I can. These are the places where it stops helping.</p>

<p><strong>Photography has a target; training doesn’t.</strong> An exposure is either about right or it isn’t, and the camera can even show you a light meter. Nothing like that exists for training. There is no learning rate that gives the “correct” amount of progress, only better and worse outcomes on a metric you picked.</p>

<p><strong>The knobs aren’t interchangeable.</strong> In a camera, a stop of aperture and a stop of shutter are worth exactly the same amount of light, so the trade is precise. In training, batch size and learning rate coupling is a rough heuristic, and it depends on the model, the optimiser and the data. I’d treat any specific scaling rule as a starting guess to test.</p>

<p><strong>Warmup isn’t a peer of the other two.</strong> ISO is a number you set once per shot. Warmup is a schedule, with a length and a shape, that runs for the first part of training and then disappears. It matters most for large models and adaptive optimisers, and it does very little in my toy problem, where the loss is a smooth bowl and the only limit is curvature. Also, real training has more knobs than three: weight decay, the schedule after warmup, gradient clipping, the optimiser choice. A triangle undercounts.</p>

<p><strong>The costs are different in kind.</strong> In photography every trade lands on a single image, and you see the result at once. In training the cost shows up later, in a curve you read after minutes or days, and sometimes only in an evaluation you weren’t watching. That delay is the reason I think “name what you’re spending” is such a useful habit for training and less necessary in a camera.</p>

<p><strong>Noise means different things.</strong> Grain in a photo is nearly always unwanted. Noise in gradient descent is sometimes a nuisance and sometimes useful, because it can help a run escape sharp regions. A smaller batch isn’t just a worse estimate.</p>

<h2 id="what-i-take-from-it">What I take from it</h2>

<p>I don’t think the analogy predicts anything. It gives me a way to explain to a colleague, or to myself, why “just change the learning rate” is rarely a single-variable question. It also reminds me to start from what I care about, the way you would with a subject that moves or a scene that’s dim.</p>

<p>When I set up or review a training run, I try to ask three questions in that order. What’s the failure I’m most worried about: divergence, slow progress, or noisy results? Which knob buys me the most protection against it? And which cost am I accepting in exchange? Then I look at the curve, the way I’d look at the back of the camera, and adjust.</p>

<h2 id="references">References</h2>

<ol>
  <li>Bottou, L., Curtis, F. E., and Nocedal, J. <em>Optimization Methods for Large-Scale Machine Learning</em>. SIAM Review, 2018. <a href="https://arxiv.org/abs/1606.04838">https://arxiv.org/abs/1606.04838</a> — stochastic gradient noise and step size</li>
  <li>Smith, S. L. and Le, Q. V. <em>A Bayesian Perspective on Generalization and Stochastic Gradient Descent</em>. ICLR, 2018. <a href="https://arxiv.org/abs/1710.06451">https://arxiv.org/abs/1710.06451</a> — SGD noise scale set by learning rate and batch size</li>
  <li>Goyal, P. et al. <em>Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour</em>. arXiv, 2017. <a href="https://arxiv.org/abs/1706.02677">https://arxiv.org/abs/1706.02677</a> — the linear scaling rule and warmup</li>
  <li>Malladi, S., Lyu, K., Panigrahi, A., and Arora, S. <em>On the SDEs and Scaling Rules for Adaptive Gradient Algorithms</em>. NeurIPS, 2022. <a href="https://arxiv.org/abs/2205.10287">https://arxiv.org/abs/2205.10287</a> — square-root scaling for Adam-style optimisers</li>
  <li>Liu, L. et al. <em>On the Variance of the Adaptive Learning Rate and Beyond</em>. ICLR, 2020. <a href="https://arxiv.org/abs/1908.03265">https://arxiv.org/abs/1908.03265</a> — why warmup helps adaptive optimisers</li>
  <li>Kalra, D. S. and Barkeshli, M. <em>Why Warmup the Learning Rate? Underlying Mechanisms and Improvements</em>. NeurIPS, 2024. <a href="https://arxiv.org/abs/2406.09405">https://arxiv.org/abs/2406.09405</a></li>
</ol>]]></content><author><name>Fadhil Mochammad</name></author><category term="photo" /><summary type="html"><![CDATA[When I take photos, I keep running into the same annoyance. I open the aperture to let in more light, and the background goes soft. I slow the shutter to let in more light, and a moving subject smears. I raise the ISO to get more light without either, and the picture gets grainy. Nothing is free.]]></summary></entry><entry><title type="html">A serving SDK data scientists actually use</title><link href="https://fadhilmch.github.io/posts/a-serving-sdk-data-scientists-actually-use/" rel="alternate" type="text/html" title="A serving SDK data scientists actually use" /><published>2025-06-14T00:00:00+02:00</published><updated>2025-06-14T00:00:00+02:00</updated><id>https://fadhilmch.github.io/posts/a-serving-sdk-data-scientists-actually-use</id><content type="html" xml:base="https://fadhilmch.github.io/posts/a-serving-sdk-data-scientists-actually-use/"><![CDATA[<p>An SDK is an interface. The people calling it are data scientists instead of other engineers, but the same rules apply: a small surface, sensible defaults, and errors that say what to do next. If every model author has to learn cluster setup, authentication, tracing and deployment conventions before shipping a model, the platform is exposing its machinery instead of helping with the task.</p>

<p>I built a serving engine and SDK for an ML platform. It supported approximately 40 model deployments, and the teams behind them did not have to configure the underlying infrastructure. This post is about the design thinking behind it. I’ll describe the shape of the interface and the decisions I’d defend. Where I’m talking about general practice rather than my own system, I’ll say so.</p>

<h2 id="the-problem-a-model-is-not-a-service">The problem: a model is not a service</h2>

<p>A trained model in a notebook is a function: features in, prediction out. A model in production is a service, and the gap between the two is mostly not machine learning. Someone has to:</p>

<ul>
  <li>put the model behind an HTTP or gRPC endpoint</li>
  <li>authenticate callers</li>
  <li>emit traces and metrics so the model can be debugged when it misbehaves</li>
  <li>package it and run it on Kubernetes with sensible resource settings</li>
  <li>handle concurrent requests, including any parallel execution the model needs</li>
  <li>keep the same conventions as every other service in the company</li>
</ul>

<p>None of these are hard individually. Together they are a lot of unfamiliar surface area for someone whose expertise is modelling. A data scientist can learn them, but each one who does spends days on work that isn’t their job, and each ends up with slightly different choices. Ten teams shipping ten models by hand gives you ten ways of doing authentication and ten different sets of dashboards.</p>

<figure class="fig">
<svg viewBox="0 0 680 300" role="img" aria-labelledby="sdk1t sdk1d">
  <title id="sdk1t">Path from notebook to endpoint, before and after an SDK</title>
  <desc id="sdk1d">Illustrative. Without a shared serving path, the model author handles six steps: the model, a web app, authentication, tracing, Kubernetes deployment, and metrics and alerts. With the SDK, the author does two things, wrap the model and deploy, and the platform provides the rest.</desc>
  <defs><marker id="sdk1a" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="arrow" d="M0,0 L10,5 L0,10 z" /></marker></defs>
  <text class="h" x="6" y="16">WITHOUT A SHARED PATH · THE AUTHOR OWNS EVERY STEP</text>
  <rect class="box" x="6" y="30" width="90" height="52" rx="6" />
  <text class="t" x="16" y="52">Notebook</text><text class="m" x="16" y="69">the model</text>
  <rect class="box" x="121" y="30" width="90" height="52" rx="6" />
  <text class="t" x="131" y="52">Web app</text><text class="m" x="131" y="69">a server</text>
  <rect class="box" x="236" y="30" width="90" height="52" rx="6" />
  <text class="t" x="246" y="52">Auth</text><text class="m" x="246" y="69">who calls</text>
  <rect class="box" x="351" y="30" width="90" height="52" rx="6" />
  <text class="t" x="361" y="52">Tracing</text><text class="m" x="361" y="69">requests</text>
  <rect class="box" x="466" y="30" width="90" height="52" rx="6" />
  <text class="t" x="476" y="52">Deploy</text><text class="m" x="476" y="69">Kubernetes</text>
  <rect class="box" x="581" y="30" width="90" height="52" rx="6" />
  <text class="t" x="591" y="52">Metrics</text><text class="m" x="591" y="69">and alerts</text>
  <path class="sb" d="M97,56 L119,56" marker-end="url(#sdk1a)" />
  <path class="sb" d="M212,56 L234,56" marker-end="url(#sdk1a)" />
  <path class="sb" d="M327,56 L349,56" marker-end="url(#sdk1a)" />
  <path class="sb" d="M442,56 L464,56" marker-end="url(#sdk1a)" />
  <path class="sb" d="M557,56 L579,56" marker-end="url(#sdk1a)" />
  <text class="tb" x="340" y="104" text-anchor="middle">five steps that are not modelling, each a chance to diverge from the last team</text>
  <line class="grid" x1="6" x2="674" y1="126" y2="126" />
  <text class="h" x="6" y="152">WITH THE SDK · THE AUTHOR OWNS THE MODEL, THE PLATFORM OWNS THE REST</text>
  <rect class="box" x="6" y="166" width="150" height="52" rx="6" />
  <text class="t" x="16" y="188">Notebook</text><text class="m" x="16" y="205">the model</text>
  <rect class="box" x="216" y="166" width="170" height="52" rx="6" style="stroke:var(--fig-a);stroke-width:2" />
  <text class="t" x="226" y="188">Wrap with SDK</text><text class="m" x="226" y="205">a class and a config</text>
  <rect class="box" x="446" y="166" width="228" height="52" rx="6" style="stroke:var(--fig-a);stroke-width:2" />
  <text class="t" x="456" y="188">Deployed endpoint</text><text class="m" x="456" y="205">auth, traces, metrics included</text>
  <path class="sa" d="M156,192 L214,192" marker-end="url(#sdk1a)" />
  <path class="sa" d="M386,192 L444,192" marker-end="url(#sdk1a)" />
  <text class="ta" x="340" y="248" text-anchor="middle">two steps for the author, and the same steps for every model</text>
  <text class="m" x="340" y="280" text-anchor="middle">Illustrative: the steps a typical hand-rolled deployment needs, not a measured timeline.</text>
</svg>
<figcaption>The point of a serving SDK is to shrink the author's path to the parts that are about the model. The steps still happen. They just happen once, in the platform.</figcaption>
</figure>

<h2 id="what-i-put-behind-the-interface">What I put behind the interface</h2>

<p>The engine took over the parts that are the same for every model:</p>

<ul>
  <li><strong>Authentication.</strong> Callers were checked the same way for every endpoint, so model authors did not write auth code.</li>
  <li><strong>Tracing.</strong> The engine handled it, so model code needed no instrumentation calls.</li>
  <li><strong>Kubernetes orchestration.</strong> The engine took care of running the model on the cluster, so authors did not configure it.</li>
  <li><strong>Parallel execution.</strong> Running work in parallel was the engine’s job.</li>
  <li><strong>Observability.</strong> Metrics and logs came out in one consistent form for every model.</li>
</ul>

<p>The design test I used for each item was simple: would two different data scientists ever have a good reason to make a different choice here? For authentication and tracing the answer is almost always no, so they belong in the platform. For anything where the answer is yes, they belong in the interface, as an option with a default.</p>

<figure class="fig">
<svg viewBox="0 0 680 262" role="img" aria-labelledby="sdk2t sdk2d">
  <title id="sdk2t">What the author writes and what the platform handles</title>
  <desc id="sdk2d">The model author writes model loading, a predict function, input and output schemas, and a small config. The SDK sits between, and the platform provides authentication, tracing, Kubernetes orchestration, parallel execution and observability.</desc>
  <defs><marker id="sdk2a" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="arrow" d="M0,0 L10,5 L0,10 z" /></marker></defs>
  <text class="h" x="6" y="16">THE AUTHOR WRITES</text>
  <text class="h" x="674" y="16" text-anchor="end">THE PLATFORM HANDLES</text>
  <rect class="box" x="6" y="30" width="230" height="200" rx="6" style="stroke:var(--fig-a);stroke-width:2" />
  <text class="t" x="20" y="60">load the model</text>
  <text class="t" x="20" y="92">predict()</text>
  <text class="t" x="20" y="124">input and output schema</text>
  <text class="t" x="20" y="156">resources and options</text>
  <text class="m" x="20" y="206">Python, in the tools they know</text>
  <rect class="box" x="264" y="30" width="152" height="200" rx="6" />
  <text class="h" x="340" y="64" text-anchor="middle">SDK</text>
  <text class="m" x="340" y="98" text-anchor="middle">one contract</text>
  <text class="m" x="340" y="118" text-anchor="middle">defaults</text>
  <text class="m" x="340" y="138" text-anchor="middle">clear errors</text>
  <text class="m" x="340" y="158" text-anchor="middle">any framework</text>
  <rect class="box" x="444" y="30" width="230" height="200" rx="6" />
  <text class="t" x="458" y="60">authentication</text>
  <text class="t" x="458" y="92">tracing</text>
  <text class="t" x="458" y="124">Kubernetes orchestration</text>
  <text class="t" x="458" y="156">parallel execution</text>
  <text class="t" x="458" y="188">observability</text>
  <path class="sa" d="M236,130 L262,130" marker-end="url(#sdk2a)" />
  <path class="sa" d="M416,130 L442,130" marker-end="url(#sdk2a)" />
  <text class="m" x="340" y="252" text-anchor="middle">Illustrative split, following the design described in the text.</text>
</svg>
<figcaption>The author's side is small and stays in Python. The right-hand column is the same for every deployment, which is exactly why it belongs in one place.</figcaption>
</figure>

<h2 id="what-the-author-still-writes">What the author still writes</h2>

<p>An SDK is easy to get wrong by hiding too much. If the only way to use it is a magic decorator that reads the notebook’s global state, nobody can reason about what it does. I wanted the author’s side to be a few explicit things:</p>

<ol>
  <li>How to load the model (files, weights, whatever the framework needs).</li>
  <li>What “predict” does for one request.</li>
  <li>What goes in and what comes out, as a schema.</li>
  <li>A short config for the things that do vary, such as resource requests.</li>
</ol>

<p>Here is the shape of that, as a sketch rather than the real API:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">from</span> <span class="n">serving_sdk</span> <span class="kn">import</span> <span class="n">Model</span><span class="p">,</span> <span class="n">Schema</span><span class="p">,</span> <span class="n">Config</span>

<span class="k">class</span> <span class="nc">ChurnModel</span><span class="p">(</span><span class="n">Model</span><span class="p">):</span>
    <span class="nb">input</span> <span class="o">=</span> <span class="nc">Schema</span><span class="p">(</span><span class="n">user_id</span><span class="o">=</span><span class="nb">str</span><span class="p">,</span> <span class="n">features</span><span class="o">=</span><span class="nb">list</span><span class="p">[</span><span class="nb">float</span><span class="p">])</span>
    <span class="n">output</span> <span class="o">=</span> <span class="nc">Schema</span><span class="p">(</span><span class="n">score</span><span class="o">=</span><span class="nb">float</span><span class="p">)</span>

    <span class="k">def</span> <span class="nf">load</span><span class="p">(</span><span class="n">self</span><span class="p">,</span> <span class="n">path</span><span class="p">):</span>
        <span class="n">self</span><span class="p">.</span><span class="n">model</span> <span class="o">=</span> <span class="nf">load_my_model</span><span class="p">(</span><span class="n">path</span><span class="p">)</span>   <span class="c1"># any framework
</span>
    <span class="k">def</span> <span class="nf">predict</span><span class="p">(</span><span class="n">self</span><span class="p">,</span> <span class="n">request</span><span class="p">):</span>
        <span class="k">return</span> <span class="p">{</span><span class="sh">"</span><span class="s">score</span><span class="sh">"</span><span class="p">:</span> <span class="n">self</span><span class="p">.</span><span class="n">model</span><span class="p">.</span><span class="nf">score</span><span class="p">(</span><span class="n">request</span><span class="p">.</span><span class="n">features</span><span class="p">)}</span>

<span class="n">config</span> <span class="o">=</span> <span class="nc">Config</span><span class="p">(</span><span class="n">cpu</span><span class="o">=</span><span class="sh">"</span><span class="s">1</span><span class="sh">"</span><span class="p">,</span> <span class="n">memory</span><span class="o">=</span><span class="sh">"</span><span class="s">2Gi</span><span class="sh">"</span><span class="p">,</span> <span class="n">replicas</span><span class="o">=</span><span class="mi">2</span><span class="p">)</span>
</code></pre></div></div>

<p>Two things make this a good interface rather than just a short one. First, the class has exactly the three methods and attributes a model needs, so there is little to misuse. Second, nothing in it refers to the framework. The engine calls <code class="language-plaintext highlighter-rouge">load</code> and <code class="language-plaintext highlighter-rouge">predict</code>, and what happens inside is the author’s business. That is how a serving engine stays framework-agnostic: the contract is about requests and responses, not about scikit-learn or PyTorch or anything else.</p>

<h2 id="design-choices-id-defend">Design choices I’d defend</h2>

<p><strong>Make the right thing the default.</strong> Settings that matter for safety or operability, like authentication and tracing, should need no configuration to be on. Turning something off should be deliberate. A default the author must remember to enable will be missed by some teams.</p>

<p><strong>Narrow beats flexible.</strong> It is tempting to offer hooks for every possible need. Each hook is a promise you have to keep as the platform changes, and each one lets teams drift back into bespoke setups. I would rather have a smaller interface that covers most models cleanly, and add an option only when several teams have asked for the same thing.</p>

<p><strong>Errors are part of the interface.</strong> A stack trace from deep inside the engine tells a data scientist nothing. Validation before deployment, with messages such as “the output schema says <code class="language-plaintext highlighter-rouge">score</code> is a float but <code class="language-plaintext highlighter-rouge">predict</code> returned a string”, turns a support conversation into a fix the author makes alone. When people say a platform is “easy to use”, a lot of that is error messages.</p>

<p><strong>Examples are documentation.</strong> A working example for each common model type does more than a reference page. Most people will copy the closest example and edit it, so the examples define the real interface.</p>

<p><strong>Keep the escape hatch honest.</strong> Some models will not fit. The choice is between a documented way out and a quiet fork of the platform. A documented way out, even a clumsy one, keeps those teams visible and lets you learn from them.</p>

<h2 id="where-abstractions-leak">Where abstractions leak</h2>

<p>Every abstraction leaks somewhere, and the useful question is where you want it to. A common place is performance: a model that needs a lot of memory or a slow start-up behaves differently under the platform’s scheduling from how it did in a notebook. When the layer hides Kubernetes completely, the author has no vocabulary for the problem.</p>

<p>The general fix I prefer is to expose the <em>effects</em> rather than the machinery. Let the author declare what the model needs (memory, warm-up time, concurrency) in the config, and surface what happened (a start-up that took too long, a request that timed out) in terms of those settings. They shouldn’t need to know what a pod is to fix it.</p>

<h2 id="a-shared-path-also-standardises-operations">A shared path also standardises operations</h2>

<p>Building one serving path had a second benefit that I underrated at the start. When every deployment is produced the same way, the platform can also say something consistent about all of them. Authentication, tracing and metrics come out in one form, so a single set of dashboards and alerts can cover every model. Governance questions, such as which models are running and who owns them, have one place to be answered.</p>

<p>Nobody had to be asked to follow a standard. It came with the easiest way to deploy.</p>

<h2 id="build-or-adopt">Build or adopt</h2>

<p>Existing tools cover this ground, and the build-versus-adopt question is a fair one. My view is a general one and depends on your situation. If an existing serving framework already matches your company’s authentication, tracing and deployment conventions, adopt it and spend the time elsewhere. If most of the work is fitting those conventions, then the real product is the thin layer that connects a model to them, and that layer is where an in-house SDK earns its keep. Either way, the design questions above still apply: the interface your data scientists see is yours to shape.</p>

<h2 id="what-id-take-from-it">What I’d take from it</h2>

<ul>
  <li>Treat data scientists as users of an interface, and design that interface the way you would a public API.</li>
  <li>Decide what belongs to the platform by asking whether any two authors would legitimately choose differently.</li>
  <li>Keep the author’s side explicit and small, and keep the contract independent of any modelling framework.</li>
  <li>Invest in defaults, examples and error messages. They are most of what people experience.</li>
  <li>Expect leaks, and choose in advance where they show up.</li>
</ul>

<p>The measure of success was not the SDK’s feature list. It was that teams could take a model from a notebook to a running endpoint without becoming infrastructure engineers, and that there was one path to keep secure and observable.</p>

<h2 id="references">References</h2>

<ol>
  <li>Sculley, D. et al. <em>Hidden Technical Debt in Machine Learning Systems</em>. NeurIPS, 2015. <a href="https://papers.nips.cc/paper_files/paper/2015/hash/86df7dcfd896fcaf2674f757a2463eba-Abstract.html">https://papers.nips.cc/paper_files/paper/2015/hash/86df7dcfd896fcaf2674f757a2463eba-Abstract.html</a></li>
  <li>OpenTelemetry. <em>Traces</em>. <a href="https://opentelemetry.io/docs/concepts/signals/traces/">https://opentelemetry.io/docs/concepts/signals/traces/</a> — spans, trace IDs and context propagation</li>
  <li>Kubernetes. <em>Resource Management for Pods and Containers</em>. <a href="https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/">https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/</a> — CPU and memory requests and limits</li>
  <li>KServe. <em>KServe documentation</em>. <a href="https://kserve.github.io/website/">https://kserve.github.io/website/</a> — an open-source model serving platform on Kubernetes</li>
  <li>BentoML. <em>BentoML documentation</em>. <a href="https://docs.bentoml.com/en/latest/">https://docs.bentoml.com/en/latest/</a> — a Python framework for packaging and serving models</li>
</ol>]]></content><author><name>Fadhil Mochammad</name></author><category term="mlops" /><category term="design" /><summary type="html"><![CDATA[An SDK is an interface. The people calling it are data scientists instead of other engineers, but the same rules apply: a small surface, sensible defaults, and errors that say what to do next. If every model author has to learn cluster setup, authentication, tracing and deployment conventions before shipping a model, the platform is exposing its machinery instead of helping with the task.]]></summary></entry><entry><title type="html">SSE on Kubernetes: what a service mesh fixes, and what it doesn’t</title><link href="https://fadhilmch.github.io/posts/sse-through-a-service-mesh/" rel="alternate" type="text/html" title="SSE on Kubernetes: what a service mesh fixes, and what it doesn’t" /><published>2025-03-22T00:00:00+01:00</published><updated>2025-03-22T00:00:00+01:00</updated><id>https://fadhilmch.github.io/posts/sse-through-a-service-mesh</id><content type="html" xml:base="https://fadhilmch.github.io/posts/sse-through-a-service-mesh/"><![CDATA[<p>I want the experimentation service to be event-driven: when someone changes an experiment, the service should tell everyone who needs to know, straight away. When nothing changes, it should stay quiet.</p>

<p>That isn’t how it works today. The services that assign users to variants pull their config: they read it through a cache that events and a TTL keep fresh, and apps ask on every request. Most of those reads get the same answer as last time.</p>

<figure class="fig">
<svg viewBox="0 0 680 170" role="img" aria-labelledby="s9t s9d">
  <title id="s9t">Polling compared with push</title>
  <desc id="s9d">Over ten minutes, a client polling every two minutes checks six times. Five checks return no change. The config changes at minute 6.5, and the poll at minute 8 finds it, so the client is stale for 1.5 minutes. With push, the server sends one message at minute 6.5, the moment the change happens.</desc>
  <line class="grid dash" x1="489" x2="489" y1="26" y2="140" /><text class="m" x="489" y="18" text-anchor="middle">config changes</text>
  <text class="t" x="10" y="56">polling every 2 min</text><text class="m" x="10" y="72">5 of 6 checks: no change</text>
  <line class="grid" x1="190" x2="650" y1="60" y2="60" />
  <g tabindex="0"><title>Polls at 0, 2, 4 and 6 min: no change</title><circle cx="190" cy="60" r="5" fill="var(--muted)" /><circle cx="282" cy="60" r="5" fill="var(--muted)" /><circle cx="374" cy="60" r="5" fill="var(--muted)" /><circle cx="466" cy="60" r="5" fill="var(--muted)" /></g>
  <g tabindex="0"><title>Stale from 6.5 to 8 min</title><line class="sb" x1="489" x2="558" y1="60" y2="60" /></g>
  <text class="tb" x="524" y="44" text-anchor="middle">stale 1.5 min</text>
  <g tabindex="0"><title>Poll at 8 min finds the change</title><circle class="fa" cx="558" cy="60" r="5" /></g>
  <g tabindex="0"><title>Poll at 10 min: no change</title><circle cx="650" cy="60" r="5" fill="var(--muted)" /></g>
  <text class="t" x="10" y="116">push (SSE)</text><text class="m" x="10" y="132">one message, when needed</text>
  <line class="grid" x1="190" x2="650" y1="120" y2="120" />
  <g tabindex="0"><title>Push: sent at 6.5 min, the moment it changes</title><circle class="fa" cx="489" cy="120" r="6" /></g>
  <text class="ta" x="477" y="140" text-anchor="end">sent the moment it changes</text>
  <text class="m" x="190" y="162">0</text><text class="m" x="650" y="162" text-anchor="end">10 min</text>
</svg>
<figcaption>Polling is checking the mailbox every two minutes. Push is a doorbell: nothing happens until there's something to deliver.</figcaption>
</figure>

<p>When I wrote about <a href="/posts/from-polling-to-push/">getting experiment config to every service</a>, push was the most advanced option, and I held back from it. Push means keeping a connection open to every client for hours, and I didn’t know how well that works inside Kubernetes. So I went and learned. This post is what I found, in the order I wish I’d learned it:</p>

<ol>
  <li>What server-sent events (SSE) are.</li>
  <li>Why long-lived connections are tricky.</li>
  <li>What a service mesh is, and what Istio does.</li>
  <li>Each problem again, and whether a mesh fixes it.</li>
</ol>

<h2 id="1-server-sent-events">1. Server-sent events</h2>

<p>SSE is a normal HTTP request whose answer never finishes. The client asks once. The server replies, keeps the connection open, and writes a new message into it whenever something happens. Think of a phone call where only one side talks, and only when there’s news.</p>

<p>This is what the client receives over time:</p>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>id: 42
data: {"experiment":"checkout-cta","status":"paused"}

: ping

id: 43
data: {"experiment":"search-rank","traffic":0.5}
</code></pre></div></div>

<p>Three things in there matter later:</p>

<ul>
  <li><strong>Each message ends with a blank line.</strong> The client handles it as soon as it arrives.</li>
  <li><strong>A line starting with a colon is a comment.</strong> The client ignores it. The server can send one every few seconds just to show the line is still alive: a heartbeat.</li>
  <li><strong>Each message can have an <code class="language-plaintext highlighter-rouge">id</code>.</strong> If the connection drops, the client reconnects and says “the last one I got was 42”. The server then sends whatever came after it. Browsers do this automatically, through a header called <code class="language-plaintext highlighter-rouge">Last-Event-ID</code>.</li>
</ul>

<p>SSE only goes one way, from server to client. That’s all config delivery needs. If you need both directions, WebSockets are the usual choice.</p>

<h2 id="2-why-long-connections-are-tricky">2. Why long connections are tricky</h2>

<p>Most web infrastructure is built for short requests: a question comes in, an answer goes out, done in under a second. An SSE connection can stay open for hours. Between the client and the server there are several middlemen, such as load balancers and proxies, and each one was tuned with short requests in mind.</p>

<p>Four things can go wrong:</p>

<ol>
  <li><strong>A middleman hangs up on a quiet line.</strong> Most proxies close connections that have been silent for a while, often after a minute.</li>
  <li><strong>A middleman holds messages back.</strong> Some proxies collect a response in a buffer before passing it on. For a normal request that’s fine. For SSE, the messages sit in the buffer and arrive late or not at all.</li>
  <li><strong>New servers get no clients.</strong> A client picks a server when it connects and stays there. If you add a server because the others are busy, existing clients don’t move to it.</li>
  <li><strong>Everyone calls back at once.</strong> When a server restarts during a deploy, all its clients lose their connection at the same moment and reconnect together.</li>
</ol>

<p>There’s also a fifth, smaller one: every open connection takes some memory on every machine it passes through.</p>

<h2 id="3-what-a-service-mesh-is">3. What a service mesh is</h2>

<p>With many services, each one needs the same networking features: timeouts, retries, encrypted traffic, and metrics on every call. Without help, every team builds these into its own code, in its own way.</p>

<p>A service mesh takes those features out of the application and puts them into a small proxy next to every copy of every service. An analogy that helped me: each service gets a personal assistant that handles its calls. The service just says “call the push service”. The assistant finds a healthy copy, encrypts the call, retries if it fails, and writes down how long it took. A manager gives every assistant the same rulebook.</p>

<p><strong>Istio</strong> is the most common service mesh on Kubernetes. In Istio’s terms:</p>

<ul>
  <li>The assistants are <strong>Envoy</strong> proxies. Istio adds one to every pod automatically, as an extra container called a <strong>sidecar</strong>, and routes all the pod’s traffic through it.</li>
  <li>The manager is <strong>istiod</strong>. It tells every Envoy where the other services are, what the rules are, and which certificates to use for encryption.</li>
  <li>The front desk is the <strong>ingress gateway</strong>, a standalone Envoy where traffic from outside the cluster comes in.</li>
</ul>

<p>You write the rules as Kubernetes objects, for example “send 10% of requests to version 2” or “time out after 3 seconds”.</p>

<figure class="fig">
<svg viewBox="0 0 680 280" role="img" aria-labelledby="s0t s0d">
  <title id="s0t">Istio's architecture</title>
  <desc id="s0d">istiod, the control plane, pushes configuration and certificates to every Envoy proxy over long-lived gRPC streams. Traffic enters through an ingress gateway, which is a standalone Envoy, and reaches the frontend pod's Envoy sidecar. Calls from the frontend to the push service go from the frontend's Envoy to the push service's Envoy over mutual TLS.</desc>
  <defs><marker id="sa0" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="arrow" d="M0,0 L10,5 L0,10 z" /></marker></defs>
  <rect class="box" x="240" y="10" width="200" height="56" rx="6" /><text class="t" x="252" y="32">istiod</text><text class="m" x="252" y="51">config · certificates</text>
  <text class="m" x="452" y="32">pushes config to every proxy</text><text class="m" x="452" y="48">over long-lived gRPC (xDS)</text>
  <rect class="box" x="10" y="140" width="150" height="56" rx="6" /><text class="t" x="22" y="162">ingress gateway</text><text class="m" x="22" y="181">Envoy at the edge</text>
  <rect class="box" x="190" y="112" width="220" height="100" rx="8" /><text class="m" x="202" y="132">pod: frontend</text>
  <rect class="box" x="202" y="146" width="92" height="44" rx="5" /><text class="t" x="214" y="173">Envoy</text>
  <rect class="box" x="306" y="146" width="92" height="44" rx="5" /><text class="t" x="318" y="173">app</text>
  <rect class="box" x="450" y="112" width="220" height="100" rx="8" /><text class="m" x="462" y="132">pod: push service</text>
  <rect class="box" x="462" y="146" width="92" height="44" rx="5" /><text class="t" x="474" y="173">Envoy</text>
  <rect class="box" x="566" y="146" width="92" height="44" rx="5" /><text class="t" x="578" y="173">app</text>
  <path class="ln dash" d="M300,66 L90,138" marker-end="url(#sa0)" /><path class="ln dash" d="M330,66 L256,110" marker-end="url(#sa0)" /><path class="ln dash" d="M380,66 L500,110" marker-end="url(#sa0)" />
  <g tabindex="0"><title>Requests from outside enter through the gateway</title><path class="sa" d="M160,168 L200,168" marker-end="url(#sa0)" /></g>
  <g tabindex="0"><title>Service-to-service calls go proxy to proxy, over mutual TLS</title><path class="sa" d="M248,190 L248,240 L508,240 L508,192" marker-end="url(#sa0)" /></g>
  <text class="ta" x="378" y="258" text-anchor="middle">mTLS between proxies</text>
</svg>
<figcaption>Dashed lines: istiod handing out the rulebook. Blue lines: real traffic, which always goes from one Envoy to another, never straight between apps.</figcaption>
</figure>

<p>There’s a nice irony here. istiod sends its rules to every Envoy over long-lived connections that stay open all the time. The mesh itself runs on the pattern I was nervous about.</p>

<h2 id="4-the-four-problems-with-a-mesh">4. The four problems, with a mesh</h2>

<p>With Istio, an SSE connection from an app to our push service passes through three middlemen: the cloud load balancer, the ingress gateway, and the Envoy sidecar next to the push service. Each has its own timers.</p>

<figure class="fig">
<svg viewBox="0 0 680 290" role="img" aria-labelledby="s1t s1d">
  <title id="s1t">Every hop on an SSE stream has a timer</title>
  <desc id="s1d">An SSE stream goes from the app or SDK through a cloud load balancer, the Istio ingress gateway, the push service's Envoy sidecar, and into the push service. The load balancer has an idle timeout, 60 seconds by default on AWS's Application Load Balancer. Envoy's route timeout is 15 seconds by default, but Istio turns it off. Envoy's stream idle timeout is 5 minutes. On deploy, Istio drains for 5 seconds by default. Without heartbeats, a quiet stream is cut after 60 seconds; with a comment line every 20 seconds, no idle timer fires.</desc>
  <defs><marker id="sa1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="arrow" d="M0,0 L10,5 L0,10 z" /></marker></defs>
  <rect class="box" x="10" y="20" width="120" height="52" rx="6" /><text class="t" x="22" y="41">app / SDK</text><text class="m" x="22" y="59">reconnects</text>
  <rect class="box" x="147" y="20" width="120" height="52" rx="6" /><text class="t" x="159" y="41">cloud LB</text><text class="m" x="159" y="59">load balancer</text>
  <rect class="box" x="284" y="20" width="120" height="52" rx="6" /><text class="t" x="296" y="41">gateway</text><text class="m" x="296" y="59">Envoy</text>
  <rect class="box" x="421" y="20" width="120" height="52" rx="6" /><text class="t" x="433" y="41">sidecar</text><text class="m" x="433" y="59">Envoy</text>
  <rect class="box" x="558" y="20" width="112" height="52" rx="6" /><text class="t" x="570" y="41">push</text><text class="m" x="570" y="59">holds streams</text>
  <path class="ln" d="M130,46 L145,46" marker-end="url(#sa1)" /><path class="ln" d="M267,46 L282,46" marker-end="url(#sa1)" /><path class="ln" d="M404,46 L419,46" marker-end="url(#sa1)" /><path class="ln" d="M541,46 L556,46" marker-end="url(#sa1)" />
  <text class="ta" x="70" y="98" text-anchor="middle">backoff</text><text class="m" x="70" y="114" text-anchor="middle">and jitter</text>
  <text class="tb" x="207" y="98" text-anchor="middle">idle timeout</text><text class="m" x="207" y="114" text-anchor="middle">60 s on AWS ALB</text>
  <text class="tb" x="344" y="98" text-anchor="middle">route timeout</text><text class="m" x="344" y="114" text-anchor="middle">15 s in Envoy,</text><text class="m" x="344" y="130" text-anchor="middle">off in Istio</text>
  <text class="tb" x="481" y="98" text-anchor="middle">stream idle</text><text class="m" x="481" y="114" text-anchor="middle">5 min in Envoy</text>
  <text class="tb" x="614" y="98" text-anchor="middle">drain on deploy</text><text class="m" x="614" y="114" text-anchor="middle">5 s in Istio</text>
  <line class="grid" x1="10" x2="670" y1="152" y2="152" />
  <text class="h" x="10" y="176">ONE QUIET STREAM, FIRST 5 MINUTES</text>
  <text class="m" x="10" y="208">no heartbeat</text>
  <g tabindex="0"><title>No heartbeat: the load balancer cuts the stream after 60 s idle</title><line class="sb" x1="150" x2="252" y1="204" y2="204" /><path class="sb" d="M246,198 L258,210 M258,198 L246,210" /></g>
  <text class="tb" x="266" y="208">cut after 60 s idle</text>
  <text class="m" x="10" y="248">comment every 20 s</text>
  <g tabindex="0"><title>With a comment line every 20 s, no idle timer fires</title><line class="sa" x1="150" x2="660" y1="244" y2="244" />
  <circle class="fa" cx="184" cy="244" r="4" /><circle class="fa" cx="218" cy="244" r="4" /><circle class="fa" cx="252" cy="244" r="4" /><circle class="fa" cx="286" cy="244" r="4" /><circle class="fa" cx="320" cy="244" r="4" /><circle class="fa" cx="354" cy="244" r="4" /><circle class="fa" cx="388" cy="244" r="4" /><circle class="fa" cx="422" cy="244" r="4" /><circle class="fa" cx="456" cy="244" r="4" /><circle class="fa" cx="490" cy="244" r="4" /><circle class="fa" cx="524" cy="244" r="4" /><circle class="fa" cx="558" cy="244" r="4" /><circle class="fa" cx="592" cy="244" r="4" /><circle class="fa" cx="626" cy="244" r="4" /></g>
  <text class="m" x="150" y="276">0</text><text class="m" x="660" y="276" text-anchor="end">5 min</text>
</svg>
<figcaption>The shortest idle timer on the path decides how long a quiet stream lives. A heartbeat shorter than all of them keeps every one from firing. The defaults shown are examples; check your own.</figcaption>
</figure>

<h3 id="a-middleman-hangs-up-on-a-quiet-line">A middleman hangs up on a quiet line</h3>

<p><strong>Does the mesh fix it?</strong> No, it adds more timers. But the fix is easy.</p>

<p>Every hop has an idle timer. For example, AWS’s load balancer closes connections that have been quiet for 60 seconds, and Envoy closes a stream after five quiet minutes. Envoy also has a 15-second limit on how long a whole response may take, which would cut every SSE connection. Istio switches that limit off by default, but it comes back if you set a timeout on the route yourself.</p>

<p><strong>Fix:</strong> send a heartbeat comment every 15 to 30 seconds. That’s shorter than every idle timer on the path, so none of them fire.</p>

<h3 id="a-middleman-holds-messages-back">A middleman holds messages back</h3>

<p><strong>Does the mesh fix it?</strong> Yes, mostly. Envoy passes messages through as they arrive.</p>

<p>Buffering usually comes from other places. NGINX-based ingress controllers buffer responses unless the server sends the header <code class="language-plaintext highlighter-rouge">X-Accel-Buffering: no</code>. Compression can also hold small messages back until it has enough to compress.</p>

<p><strong>Fix:</strong> turn buffering off for the stream, and don’t compress <code class="language-plaintext highlighter-rouge">text/event-stream</code> responses.</p>

<h3 id="new-servers-get-no-clients">New servers get no clients</h3>

<p><strong>Does the mesh fix it?</strong> No. This one surprised me.</p>

<p>Envoy is smarter than plain Kubernetes about spreading load: it balances every request, instead of every connection. But an SSE connection is one request that never ends. Once it lands on a server, it stays there, mesh or no mesh.</p>

<p><strong>Fix:</strong> have the server close each connection after a while, say every 10 minutes, with a little randomness so they don’t all close together. The client reconnects, and that time it may land on the new server.</p>

<figure class="fig">
<svg viewBox="0 0 680 260" role="img" aria-labelledby="s2t s2d">
  <title id="s2t">Share of streams on a newly added pod</title>
  <desc id="s2d">Simulation of 4,000 streams on three push pods when a fourth pod starts. With no cap on stream lifetime, and streams ending naturally about every three hours, the new pod holds 1% of streams after 10 minutes and 6% after an hour. With streams capped at 10 plus or minus 2 minutes, it holds 21% after 10 minutes and about 25%, an even share, from 15 minutes on.</desc>
  <text class="h" x="10" y="20">SHARE OF STREAMS ON THE NEW POD</text>
  <line class="grid" x1="60" x2="600" y1="210" y2="210" /><text class="m" x="52" y="214" text-anchor="end">0%</text>
  <line class="grid" x1="60" x2="600" y1="153" y2="153" /><text class="m" x="52" y="157" text-anchor="end">10%</text>
  <line class="grid" x1="60" x2="600" y1="97" y2="97" /><text class="m" x="52" y="101" text-anchor="end">20%</text>
  <line class="grid" x1="60" x2="600" y1="40" y2="40" /><text class="m" x="52" y="44" text-anchor="end">30%</text>
  <text class="m" x="60" y="228" text-anchor="middle">0</text>
  <text class="m" x="150" y="228" text-anchor="middle">10</text>
  <text class="m" x="240" y="228" text-anchor="middle">20</text>
  <text class="m" x="330" y="228" text-anchor="middle">30</text>
  <text class="m" x="420" y="228" text-anchor="middle">40</text>
  <text class="m" x="510" y="228" text-anchor="middle">50</text>
  <text class="m" x="600" y="228" text-anchor="middle">60</text>
  <text class="m" x="330" y="248" text-anchor="middle">minutes after the fourth pod starts</text>
  <line class="ln dash" x1="60" x2="600" y1="68" y2="68" />
  <text class="m" x="606" y="72">even: 25%</text>
  <g tabindex="0"><title>Never closed: 1% on the new pod after 10 min, 6% after 60 min</title><path class="sb" d="M60.0,210.0 L69.0,209.6 L78.0,209.3 L87.0,208.4 L96.0,208.3 L105.0,207.6 L114.0,206.7 L123.0,205.9 L132.0,205.0 L141.0,204.5 L150.0,203.9 L159.0,203.2 L168.0,201.9 L177.0,201.2 L186.0,199.5 L195.0,198.9 L204.0,198.2 L213.0,197.4 L222.0,197.1 L231.0,196.5 L240.0,196.3 L249.0,195.7 L258.0,195.1 L267.0,193.7 L276.0,193.4 L285.0,192.9 L294.0,192.6 L303.0,192.0 L312.0,191.6 L321.0,191.7 L330.0,191.2 L339.0,191.2 L348.0,190.7 L357.0,189.6 L366.0,188.6 L375.0,187.3 L384.0,186.8 L393.0,186.1 L402.0,185.5 L411.0,185.1 L420.0,184.4 L429.0,184.5 L438.0,183.9 L447.0,182.8 L456.0,182.1 L465.0,181.7 L474.0,181.0 L483.0,180.4 L492.0,179.3 L501.0,178.7 L510.0,178.7 L519.0,177.8 L528.0,177.6 L537.0,176.7 L546.0,176.0 L555.0,176.0 L564.0,175.6 L573.0,174.9 L582.0,174.0 L591.0,174.0 L600.0,174.0" /></g>
  <g tabindex="0"><title>Closed every 10 ± 2 min: 21% after 10 min, 26% after 20 min</title><path class="sa" d="M60.0,210.0 L69.0,199.5 L78.0,186.5 L87.0,174.4 L96.0,162.5 L105.0,149.1 L114.0,137.3 L123.0,124.9 L132.0,115.4 L141.0,104.3 L150.0,90.4 L159.0,79.0 L168.0,66.2 L177.0,64.8 L186.0,64.5 L195.0,63.8 L204.0,64.9 L213.0,64.6 L222.0,64.9 L231.0,64.6 L240.0,62.7 L249.0,63.9 L258.0,64.9 L267.0,66.6 L276.0,66.2 L285.0,64.1 L294.0,64.1 L303.0,62.8 L312.0,62.9 L321.0,63.8 L330.0,63.4 L339.0,62.4 L348.0,65.5 L357.0,66.8 L366.0,68.0 L375.0,68.6 L384.0,71.4 L393.0,75.0 L402.0,76.1 L411.0,77.0 L420.0,77.7 L429.0,76.6 L438.0,74.8 L447.0,71.3 L456.0,70.3 L465.0,69.3 L474.0,65.8 L483.0,63.1 L492.0,63.5 L501.0,64.9 L510.0,65.9 L519.0,63.5 L528.0,64.1 L537.0,67.2 L546.0,70.0 L555.0,70.9 L564.0,69.3 L573.0,70.2 L582.0,68.9 L591.0,68.3 L600.0,65.2" /></g>
  <text class="tb" x="600" y="164" text-anchor="end">never closed</text>
  <text class="ta" x="140" y="140">closed every ~10 min</text>
</svg>
<figcaption>Simulated, illustrative numbers for a fourth server added to three busy ones. If connections never close, the new server stays almost empty for hours. If each connection closes after about 10 minutes, the load evens out within one cycle.</figcaption>
</figure>

<h3 id="everyone-calls-back-at-once">Everyone calls back at once</h3>

<p><strong>Does the mesh fix it?</strong> Partly.</p>

<p>During a deploy, Istio gives the old server a short grace period, five seconds by default, and then closes its connections. The mesh won’t reconnect for the client, so every client does it on its own. What the mesh does add is protection for the servers: it can limit connections and stop sending traffic to a server that is struggling.</p>

<p><strong>Fix:</strong> clients wait a random few seconds before reconnecting, and send the last <code class="language-plaintext highlighter-rouge">id</code> they saw so they don’t miss anything.</p>

<pre><code class="language-mermaid">sequenceDiagram
  participant C as App
  participant G as Gateway
  participant P1 as Old push pod
  participant P2 as New push pod
  C-&gt;&gt;G: open stream
  G-&gt;&gt;P1: forward
  P1--&gt;&gt;C: message 41
  Note over P1: deploy starts
  P1--&gt;&gt;C: goodbye, stream closes
  Note over C: wait a random 0–5 s
  C-&gt;&gt;G: open stream, last id was 41
  G-&gt;&gt;P2: forward
  P2--&gt;&gt;C: messages 42 and 43, then live
</code></pre>

<p>One trap is specific to meshes. In older setups, the sidecar could shut down before the application, cutting connections before the server could say goodbye. Newer Kubernetes versions let sidecars start first and stop last, and Istio can use that.</p>

<h3 id="and-the-memory-cost">And the memory cost</h3>

<p><strong>Does the mesh fix it?</strong> No, it makes it a bit worse. Each open connection is now also held by the gateway and the sidecar. It’s small per connection, but worth measuring before you have hundreds of thousands.</p>

<h3 id="what-the-mesh-adds">What the mesh adds</h3>

<p>Some things come for free with a mesh: encrypted traffic between services without code changes, metrics on how many streams are open and how long they last, and the option to send a new version of the push service only a small share of new connections at first.</p>

<h2 id="summary">Summary</h2>

<table>
  <thead>
    <tr>
      <th>Problem</th>
      <th>Does a mesh fix it?</th>
      <th>What you do</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Middlemen hang up on quiet lines</td>
      <td>No, adds more timers</td>
      <td>heartbeat every 15–30 s</td>
    </tr>
    <tr>
      <td>Middlemen hold messages back</td>
      <td>Mostly</td>
      <td>buffering off, no compression</td>
    </tr>
    <tr>
      <td>New servers get no clients</td>
      <td>No</td>
      <td>close connections every ~10 min, with randomness</td>
    </tr>
    <tr>
      <td>Everyone calls back at once</td>
      <td>Partly</td>
      <td>random wait, resume from last id</td>
    </tr>
    <tr>
      <td>Memory per connection</td>
      <td>No, a bit worse</td>
      <td>measure it</td>
    </tr>
  </tbody>
</table>

<h2 id="would-i-build-it-now">Would I build it now</h2>

<p>For our config system, with around ten changes a day, not yet. Polling cheaply with ETags, or Firestore listeners, gives us enough freshness with far less to run. And adopting a whole service mesh for one push service would be a lot.</p>

<p>If we needed changes to reach every client within a second, I’d build SSE with the fixes above: heartbeats, connections that close every few minutes, message ids for resuming, random waits on reconnect, buffering off, and polling as a fallback. I’d only use a mesh if the cluster already had one. None of the fixes depend on it.</p>

<p>My worry was right about which problems exist. I overestimated how hard they are: each one has a known fix, and most of the fixes are a few lines of server and client code.</p>

<h2 id="references">References</h2>

<ol>
  <li>WHATWG. <em>HTML Standard: Server-sent events</em>. <a href="https://html.spec.whatwg.org/multipage/server-sent-events.html">https://html.spec.whatwg.org/multipage/server-sent-events.html</a> — the SSE format, <code class="language-plaintext highlighter-rouge">Last-Event-ID</code> and comment heartbeats</li>
  <li>MDN Web Docs. <em>Using server-sent events</em>. Mozilla. <a href="https://developer.mozilla.org/en-US/docs/Web/API/Server-sent_events/Using_server-sent_events">https://developer.mozilla.org/en-US/docs/Web/API/Server-sent_events/Using_server-sent_events</a></li>
  <li>Istio. <em>Architecture</em>. <a href="https://istio.io/latest/docs/ops/deployment/architecture/">https://istio.io/latest/docs/ops/deployment/architecture/</a> — istiod, Envoy sidecars and the data plane</li>
  <li>Amazon Web Services. <em>Application Load Balancers</em>. Elastic Load Balancing documentation. <a href="https://docs.aws.amazon.com/elasticloadbalancing/latest/application/application-load-balancers.html">https://docs.aws.amazon.com/elasticloadbalancing/latest/application/application-load-balancers.html</a> — idle timeout, 60 seconds by default</li>
  <li>Envoy. <em>How do I configure timeouts?</em> Envoy proxy documentation. <a href="https://www.envoyproxy.io/docs/envoy/latest/faq/configuration/timeouts">https://www.envoyproxy.io/docs/envoy/latest/faq/configuration/timeouts</a> — route and stream idle timeouts</li>
  <li>NGINX. <em>Module ngx_http_proxy_module</em>. <a href="https://nginx.org/en/docs/http/ngx_http_proxy_module.html">https://nginx.org/en/docs/http/ngx_http_proxy_module.html</a> — <code class="language-plaintext highlighter-rouge">proxy_buffering</code> and <code class="language-plaintext highlighter-rouge">X-Accel-Buffering</code></li>
  <li>Kubernetes. <em>Sidecar Containers</em>. <a href="https://kubernetes.io/docs/concepts/workloads/pods/sidecar-containers/">https://kubernetes.io/docs/concepts/workloads/pods/sidecar-containers/</a> — sidecar start and stop ordering</li>
</ol>]]></content><author><name>Fadhil Mochammad</name></author><category term="systems" /><category term="exp" /><summary type="html"><![CDATA[I want the experimentation service to be event-driven: when someone changes an experiment, the service should tell everyone who needs to know, straight away. When nothing changes, it should stay quiet.]]></summary></entry></feed>