← All posts

FM synthesis and explore–exploit

Where my initials come from, and why modulation feels like bandits.

My initials are FM. Fadhil Mochammad. So when I first learned that “FM” also stands for frequency modulation, and that it is one of the ways synthesisers make sound, I took it personally. I also make music in my spare time, so the coincidence stuck.

Later, at work, I spent time on multi-armed bandits, the family of algorithms that decide which option to show next while still learning which option is best. At some point the two ideas started to rhyme in my head. This post is my attempt to say exactly how they rhyme, and, more usefully, where they don’t. FM synthesis first, because it’s the less familiar half.

What FM synthesis does

A sine wave is the plainest sound there is: one frequency, no character. FM synthesis makes something richer with two sine waves:

  • The carrier is the wave you hear, at the pitch you play.
  • The modulator is a second wave, usually at a different frequency, that isn’t heard directly. It pushes the carrier’s frequency up and down.

The whole thing fits in one line:

output(t) = sin( 2π · fc · t  +  β · sin( 2π · fm · t ) )

Here fc is the carrier frequency, fm is the modulator frequency, and β (beta) is the modulation index. The index says how hard the modulator pushes. Strictly, this is phase modulation, which is what most FM synthesisers implement in practice, and it sounds the same for our purposes.

Carrier, modulator and FM output waveforms Three stacked waveforms over the same time window. The carrier is a sine wave with 8 cycles. The modulator is a slower sine wave with 2 cycles. The FM output is the carrier with its phase pushed back and forth by the modulator at index 3, so its cycles bunch up and stretch out twice across the window. CARRIER · 8 CYCLES Carrier waveform MODULATOR · 2 CYCLES Modulator waveform FM OUTPUT · INDEX 3 Fm waveform time
Computed from the FM formula with a carrier-to-modulator ratio of 4:1 and index 3. The output stays a sine-like wave, but the modulator squeezes and stretches its cycles.

Look at the bottom panel. It’s still a wave, but its cycles bunch together where the modulator pushes the frequency up and stretch out where it pushes down. That wobble happens 2 times in the window, once per modulator cycle. Play it fast enough and you stop hearing a wobble. You hear a different timbre, the quality that makes a bell sound different from a flute at the same pitch.

What the index does to the sound

The interesting part is what the output is made of. A modulated sine wave is no longer a single frequency. It’s the carrier plus a series of sidebands: extra frequencies at fc ± fm, fc ± 2·fm, fc ± 3·fm, and so on. The strength of the n-th sideband is given by a Bessel function of the first kind, written Jn(β). You don’t need to know how to derive it. You only need to know that it’s a fixed function of the index, and that I computed the values for the figures below directly from its integral definition.

FM sidebands spread as the modulation index rises Three spectra of amplitude against frequency offset from the carrier, in steps of the modulator frequency, computed from Bessel functions. At index 0.5 nearly all the amplitude is in the carrier with two small sidebands. At index 2 the carrier is 0.22 and sidebands out to plus and minus 4 are visible. At index 5 the amplitude is spread across offsets from minus 8 to plus 8 and the carrier is 0.18. INDEX 0.5 offset -9: amplitude 0.000 offset -8: amplitude 0.000 offset -7: amplitude 0.000 offset -6: amplitude 0.000 offset -5: amplitude 0.000 offset -4: amplitude 0.000 offset -3: amplitude 0.003 offset -2: amplitude 0.031 offset -1: amplitude 0.242 offset +0: amplitude 0.938 offset +1: amplitude 0.242 offset +2: amplitude 0.031 offset +3: amplitude 0.003 offset +4: amplitude 0.000 offset +5: amplitude 0.000 offset +6: amplitude 0.000 offset +7: amplitude 0.000 offset +8: amplitude 0.000 offset +9: amplitude 0.000 carrier 0.94, first pair 0.24 INDEX 2 offset -9: amplitude 0.000 offset -8: amplitude 0.000 offset -7: amplitude 0.000 offset -6: amplitude 0.001 offset -5: amplitude 0.007 offset -4: amplitude 0.034 offset -3: amplitude 0.129 offset -2: amplitude 0.353 offset -1: amplitude 0.577 offset +0: amplitude 0.224 offset +1: amplitude 0.577 offset +2: amplitude 0.353 offset +3: amplitude 0.129 offset +4: amplitude 0.034 offset +5: amplitude 0.007 offset +6: amplitude 0.001 offset +7: amplitude 0.000 offset +8: amplitude 0.000 offset +9: amplitude 0.000 carrier 0.22, spread to ±4 INDEX 5 offset -9: amplitude 0.006 offset -8: amplitude 0.018 offset -7: amplitude 0.053 offset -6: amplitude 0.131 offset -5: amplitude 0.261 offset -4: amplitude 0.391 offset -3: amplitude 0.365 offset -2: amplitude 0.047 offset -1: amplitude 0.328 offset +0: amplitude 0.178 offset +1: amplitude 0.328 offset +2: amplitude 0.047 offset +3: amplitude 0.365 offset +4: amplitude 0.391 offset +5: amplitude 0.261 offset +6: amplitude 0.131 offset +7: amplitude 0.053 offset +8: amplitude 0.018 offset +9: amplitude 0.006 -8 -4 fc +4 +8 offset from carrier, in steps of fm carrier 0.18, spread to ±8
Amplitude of each component is |Jn(β)|, the Bessel function of the first kind (computed numerically). Raising the index moves energy out of the carrier into more and more sidebands.

Three things to take from this:

  • At a low index (0.5), almost everything is still the carrier. The sound is nearly a pure sine with a faint colouring.
  • At index 2, the carrier has dropped to about a fifth of its amplitude and the first few sidebands have taken over.
  • At index 5, energy is spread across a dozen or more components, and the sound is bright and dense.

A useful rule of thumb, known as Carson’s rule, says the occupied bandwidth is roughly 2 · fm · (β + 1). Widen the index and you widen the spectrum, in proportion.

The total is conserved, though. If you square each component’s amplitude and add them up, you get exactly 1 for every index. The index doesn’t add energy. It moves it around.

Share of power in the carrier as the index rises Line chart, index from 0 to 8. The carrier's share of total power, J0 squared, starts at 1, falls to zero near index 2.4, and then rises and falls in a shrinking ripple. The remainder, in the sidebands, is one minus that. 0 0.5 1 0 2 4 6 8 modulation index β Power left in the sidebands: 1 minus J0 squared Power in the carrier: J0(β) squared carrier sidebands
Total power is fixed: the carrier's share plus the sidebands' share always sums to 1. The carrier does not fade smoothly, though. It hits zero near β = 2.4 and then comes back.

That figure holds one more surprise. The carrier doesn’t fade in a straight line. Its share hits zero near β = 2.4, so at that setting you would hear the sidebands but not the note’s own frequency, and then it comes back. Small changes in one parameter can produce results that aren’t monotonic, and I like that about FM. It’s part of why FM patches can be hard to predict by intuition alone.

The explore–exploit problem in one paragraph

Now the other half. Suppose you have several options, for instance several versions of a page, and you don’t know which one users prefer. Each time someone arrives you must pick one. If you always show the option that looks best so far, you exploit: you collect the most reward from what you know. If you sometimes show another option, you explore: you give up a little now to learn something you may not know yet.

A multi-armed bandit is a policy for balancing the two. The simplest one is epsilon-greedy: with probability ε pick a random option, otherwise pick the current best. Thompson Sampling is more refined. It draws a plausible value for each option from what you currently believe and picks the highest draw, so options you’re unsure about get shown more often. I’ve built bandit pipelines using Thompson Sampling, and the design keeps coming back to a single question: how much should we still be exploring?

Where the analogy holds

The comparison works if you keep it modest.

  • The carrier is the current best option. It carries most of the signal, and it’s what you’d play if you weren’t curious.
  • The sidebands are the other options that still get some traffic. They’re what makes the output richer than a single pure tone.
  • The modulation index is the exploration rate. One knob controls how far the output spreads from the centre. At a low setting you stay close to the best; at a high setting the weight is spread across many alternatives.
  • The total is fixed. Sideband power and carrier power sum to one, just as traffic shares across options sum to one. Giving more to the alternatives means taking it from the best.
Where the FM and bandit analogy holds and where it stops Left column lists FM ideas and right column the bandit ideas they resemble: carrier to the best current option, sidebands to the other options that get some traffic, modulation index to the exploration rate, and total power fixed to traffic shares summing to one. A bottom row in orange lists what FM lacks: no feedback from outcomes and no learning, so the index never shrinks by itself. FM SYNTHESISBANDIT Carrier best current option, most of the traffic Sidebands other options that still get some traffic Modulation index exploration rate: how far to spread Total power fixed traffic shares sum to one Where it stops FM has no feedback: nothing learns, so the index never shrinks on its own.
The mapping is about shape, not mechanism. FM spreads a signal deterministically; a bandit spreads traffic and then uses the outcomes to narrow it again.

There’s a further parallel I find pleasing. In both worlds a small change in the setting can have a large effect, and in both you can go too far. Push the index too high and the sound turns to noise. Push exploration too high and you spend traffic on options you already had good reason to ignore.

Where it breaks

This is an analogy, and I want to be honest about what it hides.

FM has no feedback. The modulator doesn’t know what the carrier sounded like. Nothing in the formula looks at an outcome and adjusts. A bandit does exactly that: every observed reward changes what it believes and therefore how much it explores. In a healthy bandit, exploration falls as evidence accumulates. In FM the index stays wherever you set it, unless you or an envelope changes it.

FM’s variation is periodic; exploration is statistical. The modulator is a perfect sine wave, so the “wandering” repeats forever. Exploration in a bandit is driven by uncertainty, and it’s spent where uncertainty is highest. FM spreads the sound evenly by rule, not where it’s most valuable to look.

The goal is different. An FM patch is judged by ear. A bandit is judged by regret, the reward you gave up compared with always picking the best option. There isn’t an FM equivalent of “the best arm”, so there’s nothing for exploration to be converging on.

Sidebands are a formula; allocations are a history. Sideband amplitudes come from a closed formula of the index alone. Bandit allocations depend on the whole history of what’s happened, including luck.

So I wouldn’t use FM to design a bandit, or the other way round. What I take from the comparison is a habit of thought.

What I actually take from it

The habit is to look for the one knob that controls how much variation a system is allowed to have, and to ask what it does at the extremes.

  • At zero, FM is a pure tone and a bandit is pure exploitation. Both are clean and predictable, and both are stuck with whatever they started with.
  • At a very high setting, FM is noise and a bandit is a uniform random split. Both are lively and both are useless.
  • The interesting region is in between, and the right value depends on what you are trying to achieve and how much you can afford to lose while finding out.

The instinct isn’t new to me. I studied electrical engineering, where signals, spectra and modulation are everyday vocabulary, and asking “what does this look like in the frequency domain?” becomes a reflex. It’s a small step from there to asking “what does this look like as a distribution over choices?” Both questions are about where the weight sits.

And there’s one more practical point that I hold onto from the experiment side. A bandit’s exploration setting is not a default to be left alone. It is a decision about how much you’re willing to pay for information, and it should be reviewed like any other. The index on an FM patch is the same kind of thing: an artistic decision, made deliberately, that you can hear the consequences of immediately.

A small experiment to try

If you’d like to hear this, you don’t need a synthesiser. A few lines of code in any language with a sine function and an audio library will do:

  1. Pick a carrier of 220 Hz and a modulator of 220 Hz, so the ratio is 1:1.
  2. Render one second of the formula above at β = 0, 0.5, 2 and 5.
  3. Listen in order. The step from 0 to 0.5 barely changes the tone. From 2 to 5, the sound goes from bell-like to harsh.

Then imagine the same four settings as exploration rates on a bandit, and ask which one you’d trust in production. My answer changes with the cost of a wrong choice. That is the part FM can’t teach, and it is the part I keep coming back to.

References

  1. Chowning, J. M. The Synthesis of Complex Audio Spectra by Means of Frequency Modulation. Journal of the Audio Engineering Society 21(7), 1973. https://yamahasynth.com/wp-content/uploads/images/fm_synthesispaper-2.pdf — the original FM synthesis paper (author’s digital copy)
  2. NIST. Digital Library of Mathematical Functions, §10.2: Bessel functions, definitions. https://dlmf.nist.gov/10.2
  3. Carson, J. R. Notes on the Theory of Modulation. Proceedings of the IRE 10(1), 1922. https://zenodo.org/records/1432494 — origin of Carson’s bandwidth rule
  4. Lattimore, T. and Szepesvári, C. Bandit Algorithms. Cambridge University Press, 2020. https://www.cambridge.org/9781108486828
  5. Thompson, W. R. On the Likelihood That One Unknown Probability Exceeds Another in View of the Evidence of Two Samples. Biometrika 25(3–4), 1933. https://academic.oup.com/biomet/article-abstract/25/3-4/285/200862
  6. Russo, D. et al. A Tutorial on Thompson Sampling. Foundations and Trends in Machine Learning 11(1), 2018. https://arxiv.org/abs/1707.02038
  7. Chapelle, O. and Li, L. An Empirical Evaluation of Thompson Sampling. Advances in Neural Information Processing Systems 24, 2011. https://proceedings.neurips.cc/paper/2011/hash/e53a0a2978c28872a4505bdb51db06dc-Abstract.html