What I learned tuning retrieval for a multilingual RAG chatbot
Multilingual questions against a multilingual knowledge base: English questions were fine all along, while Swedish, Polish and Indonesian ones quietly lost their documents until each question was translated into the knowledge base's languages.
The chatbot I’ve been building answers questions about company policies, most of which are specific to one country. That makes it a multilingual problem twice over. The questions come in whatever language people like to write in. The knowledge base is multilingual too, and each country’s mix is different.
Every answer depends on retrieval. If the right document isn’t among the few the language model gets to read, no prompt will fix the answer. It will either say it doesn’t know or, worse, write something plausible from the wrong source. So before touching prompts I spent a few weeks on the search step: about two dozen configurations, each scored against the same golden set of questions.
This post covers three countries, each with its own knowledge base:
- Sweden: policy pages in Swedish and in English. Questions arrive in English, Swedish and Indonesian.
- Poland: bilingual pages, Polish and English side by side. Questions arrive in English and Polish.
- United States: pages in English only, and questions in English only.
So there are six combinations of knowledge base and question language, and I’ll call each one a cell. Thinking in cells turned out to matter more than any single setting. In short: English questions were fine all along, the overall average hid that questions in Swedish, Polish and Indonesian were quietly losing their documents, and translating each question into the knowledge base’s languages fixed it.
How I scored each configuration
Every question in the golden set is annotated with the page or pages a good answer needs, and tagged with its language and the country it’s about. I wrote about building that set in Golden datasets without the gold rush.
Two metrics did most of the work:
- Recall@10: the share of the needed pages that appear in the top 10 results. This is the main one, because a page that isn’t retrieved can’t be used.
- MRR (mean reciprocal rank): how high the first correct page ranks. When recall ties, a higher MRR means the right page is nearer the top, where the model is more likely to use it.
I also tracked how many documents each configuration passed on to the answer model, plus latency and cost, which I’ll give as ratios. Each experiment changed one thing from the configuration before it, so every gain or loss has a single cause. The previous post explains why I read these numbers as evidence rather than as a pass/fail gate.
Step 1: the search mode ladder
I started by trying every query mode the search engine offers, from plain keyword matching to hybrid search with a reranker.
Keyword search alone found the right pages about half the time. Vector search, which matches meaning rather than words, jumped to 84%, and hybrid search, which runs both and fuses the results, landed in the same place.
The reranker is a model that reads the question and each candidate page together and scores how well they match. On keyword search it added 20 points of recall. On hybrid search, recall barely moved (84.0% to 83.5%), but MRR jumped from 0.593 to 0.711: the same pages, with the right one much nearer the top.
I kept it, for two reasons. The answer model reads the top results first, so ranking matters. And the reranker’s score is what makes it possible to cut the list with a threshold later. But that flat recall number was hiding something, and it came back in the next step.
Step 2: make it cheaper, then look at the cells
Ten documents per question is a lot for the answer model to read, so the next changes were about passing on less.
- A threshold on the reranker score drops weak candidates instead of always passing ten.
- A country filter limits the search to documents for the country the question is about, plus those that apply everywhere.
Together they cost half a point (83.5% to 83.0%) and cut the average number of documents passed on from 10 to 7.9. Up to this point, every configuration had also appended the country name to the query text (“… in Sweden”). With the filter in place that looked redundant, and it hurt other kinds of questions, so the next change removed it. Recall dropped to 81.0%.
Two and a half points down from the base, for a cleaner query and 6.8 documents. Then I split the results into cells.
English questions held up in every knowledge base. Every country has English pages, so an English question always shares a language with something it can find. Sweden’s English questions stayed around 94%, Poland’s rose 4 points, and the United States’ dipped by 2.5. Nothing dramatic.
The other three cells told a different story:
- Swedish questions against Sweden’s mixed knowledge base fell hardest: 90% with plain hybrid search, 70% with the reranker, 55% at the end. The MRR numbers add a twist. Without the reranker, these questions found their pages but ranked them poorly, with an MRR of just 0.17. The reranker pulled the pages it kept up the list (MRR 0.45), but pushed others out of the top 10 entirely.
- Polish questions against Poland’s bilingual pages slid from 85% to 75%. A smaller loss than Swedish, even though both are questions in the local language.
- Indonesian questions against Sweden, which share no language with any page, actually did well with the reranker (100%, and MRR up from 0.37 to 0.78). They only fell, to 85%, when the country name came out of the query.
I don’t have a complete explanation for the reranker step. The cell that lost most to it, Swedish questions, is one where the question does share a language with some of the pages, while the cell with no shared language gained. So it isn’t simply “non-English questions do worse”. My guess is that the reranker is weaker on Swedish text in particular. The last step is easier to read: removing the appended country name cost all three non-English cells 5 to 15 points, while the English cells as a whole didn’t change. An English word like “Sweden” seems to have been giving those questions an anchor in the English pages.
These cells are small, so one question moves the number by several points. Even so, a 35-point drop for Swedish questions is not something an overall average should be allowed to hide.
Step 3: translate into the knowledge base’s languages
The fix was to stop searching with the question as written, and search in the knowledge base’s languages instead. The walkthrough below goes through it step by step. Steps 1 and 2 recap the single search path, 3 to 5 cover translation and merging, and 6 covers a third path I’ll come back to.
Every question now becomes two queries, one per language the knowledge base is written in:
- An English translation, because every country has English pages.
- A translation into the country’s own language: Swedish for Sweden, Polish for Poland. For the United States, whose pages are all English, both queries end up in English.
The language the person wrote in is ignored. An Indonesian question about Sweden gets English and Swedish queries, not an Indonesian one, because no pages are written in Indonesian and searching for Indonesian text would find nothing. Picking languages from the knowledge base’s side, not the user’s, is the part of this design that matters most.
Both queries run in parallel through the same hybrid search and reranker. The two result lists are merged by keeping each page’s higher score, and the same threshold applies to the merged list.
Overall recall went from 81.0% to 85.5%. The gain landed exactly where the problem was:
- Swedish questions went from 55% to 100%, and their MRR from 0.39 to 0.67. Both problems from step 2, missing pages and poor ranking, improved at once.
- Indonesian questions went from 85% to 100%.
- Polish questions went from 75% to 85%, back to where they started before the reranker.
- English questions stayed within a point of where they were, in all three knowledge bases. For an English question, the English translation is close to the original, and in the United States there’s nothing else to translate into.
Fifteen questions that used to miss now found their page, and fourteen of them were asked in Swedish, Indonesian or Polish. One English question got worse. The model reads 6.4 documents on average, slightly fewer than before. Compared with the version that still had the country name in the query, recall is 2.5 points higher: a proper translation does what the appended word was doing by accident.
What I can’t tell you yet is how the work splits between the two queries. I didn’t run each one on its own, so I don’t know how much of the gain comes from the English query and how much from the local one. That’s the next experiment.
The cost is one translation call per question, which is the only language-model call in retrieval. Retrieval latency a little more than doubled, but it’s still a small part of the time a full answer takes, and cost per question rose by about a quarter.
Merging: keep the best score, don’t fuse ranks
The usual way to merge ranked lists is reciprocal rank fusion (RRF). Each document gets 1/(60 + rank) for every list it appears in, and the totals are summed. It ignores the scores completely, which is its strength when the lists come from systems whose scores aren’t comparable.
Here the scores are comparable, because both paths are scored by the same reranker. Taking the maximum keeps that information: a page one path is confident about stays near the top even if the other path missed it. RRF instead rewards pages that turn up everywhere, so a page that was mediocre on both paths can overtake one that a single path was sure of. Step 5 of the walkthrough shows this happening.
When I tested both, recall was the same (84.0% against 83.7%), but max score ranked the right page higher: MRR 0.706 against 0.637. There’s a practical reason too. A threshold only makes sense on a score with a fixed meaning, and RRF totals don’t have one.
The condition matters, though. If your paths are different retrievers with different score scales, such as raw keyword scores next to vector similarities, the max is meaningless and RRF or score normalisation is the right choice.
HyDE: neutral alone, marginal together
HyDE (Hypothetical Document Embeddings) asks a language model to write a plausible answer, then searches with that passage instead of the question. The idea is that an answer-shaped passage looks more like the documents than a question does.
Used on its own instead of the question, it changed nothing overall.
Overall recall stayed at 83.0%: eight questions gained, seven lost, and retrieval was four times slower. Underneath, Swedish questions gained 5 points and Indonesian questions lost 5. That mix is what made it worth one more try. A technique that finds different pages can be useful as an extra path even if it doesn’t beat the original, so I added it as a third parallel path with the same max-score merge. A path that only adds candidates can’t push good results out, because the merge keeps whichever score is highest.
Next to the two translated queries, it added 0.6 points: 86.1% against 85.5%, three questions gained (two of them in the United States) and one lost. That cost 2.4 times the retrieval latency and 4.6 times the language-model tokens. For a chatbot, where people are waiting for an answer, that isn’t worth it, so the chat setup uses two paths.
The lesson is to judge an extra path against your best setup, not a weak one. Once the translations were in place, most of what HyDE could have caught was already being found.
Every step, every cell
The matrix below puts the whole story in one place. Rows are knowledge bases, columns are question languages, and each cell shows how that combination did. Step through the configurations, or press play, and watch which cells move. Switch to MRR to see the ranking side.
Six cells, six configurations
Rows are knowledge bases, columns are the language the question was asked in. Darker cells score higher.
Configuration used for chat: English and country-language translations, merged by max score.
| Knowledge base | English | Swedish | Polish | Indonesian |
|---|---|---|---|---|
| SwedenSwedish + English pages | 93.5% | 100.0% | not asked | 100.0% |
| PolandPolish + English pages | 83.3% | not asked | 85.0% | not asked |
| United StatesEnglish pages | 65.3% | not asked | not asked | not asked |
Compared with the hybrid-plus-reranker base, the configuration I use for chat finds the right pages more often (85.5% against 83.5%), ranks them about as well (MRR 0.705 against 0.711), and passes on 6 documents instead of 10. The overall gain looks modest because most questions are in English, and English questions had nothing to gain. Swedish and Indonesian questions went from the worst cells to perfect ones, and Polish questions won back what they had lost.
I also checked top-k. Keeping only 5 results cost over 3 points, and keeping 20 added under a point for twice the reading, so 10 stayed.
What’s still open
- The United States sits at 65% to 69% in every configuration. It’s the simplest cell, English questions against English pages, and nothing on the search side moved it. When no retrieval change moves a number, I’d look at the documents next: whether the answer is missing, buried in a long page, or phrased very differently from how people ask.
- Measuring each query on its own. An English-only run and a local-language-only run would show how the work splits between the two translations, cell by cell.
What I’d tell myself at the start
- Map your questions’ languages against your knowledge base’s languages before tuning anything. The cells behave differently, and the same setting can help one and hurt another.
- Split every result by those cells before trusting it. My biggest problem showed up as a three-point dip in the average.
- Judge the reranker by MRR, and check it per language. Its overall recall barely moved, while it cost Swedish questions 20 points.
- Search in the knowledge base’s languages, not the user’s. English plus the country’s own language, whatever the question was written in. Don’t rely on a stray English word in the query to do that job.
- Merge on scores when they share a scale, and keep a threshold. Use RRF when they don’t.
- Try new techniques as extra paths, and measure the gain against your best setup.
References
- Robertson, S. and Zaragoza, H. The Probabilistic Relevance Framework: BM25 and Beyond. Foundations and Trends in Information Retrieval, 2009. https://doi.org/10.1561/1500000019
- Karpukhin, V. et al. Dense Passage Retrieval for Open-Domain Question Answering. EMNLP, 2020. https://arxiv.org/abs/2004.04906
- Nogueira, R. and Cho, K. Passage Re-ranking with BERT. arXiv, 2019. https://arxiv.org/abs/1901.04085: the cross-encoder reranking idea.
- Cormack, G. V., Clarke, C. L. A. and Büttcher, S. Reciprocal Rank Fusion Outperforms Condorcet and Individual Rank Learning Methods. SIGIR, 2009. https://doi.org/10.1145/1571941.1572114
- Gao, L., Ma, X., Lin, J. and Callan, J. Precise Zero-Shot Dense Retrieval without Relevance Labels. ACL, 2023. https://arxiv.org/abs/2212.10496: HyDE.
- Rackauckas, Z. RAG-Fusion: a New Take on Retrieval-Augmented Generation. arXiv, 2024. https://arxiv.org/abs/2402.03367: multiple generated queries merged with RRF.
- Thakur, N. et al. BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models. NeurIPS Datasets and Benchmarks, 2021. https://arxiv.org/abs/2104.08663
- Zhang, X. et al. MIRACL: A Multilingual Retrieval Dataset Covering 18 Diverse Languages. TACL, 2023. https://arxiv.org/abs/2210.09984: a benchmark for retrieval when questions and documents span many languages.
- Lewis, P. et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS, 2020. https://arxiv.org/abs/2005.11401