← All posts

The exposure triangle, but for learning rates

Aperture, shutter and ISO behave a lot like learning rate, batch size and warmup.

When I take photos, I keep running into the same annoyance. I open the aperture to let in more light, and the background goes soft. I slow the shutter to let in more light, and a moving subject smears. I raise the ISO to get more light without either, and the picture gets grainy. Nothing is free.

Photographers call this the exposure triangle. Neural-network training has its own version of that problem, with three knobs that also pull against each other: learning rate, batch size and warmup. This post lays the two side by side, shows what the training knobs do in a small simulation, and spends a fair amount of time on where the comparison stops being useful.

The photography version

A camera sensor needs a certain amount of light to produce a well-exposed picture. Three controls set how much it gets, and each one changes something else about the image:

  • Aperture is the size of the opening in the lens. A wide opening lets in more light but keeps only a thin slice of the scene in focus, which is the shallow depth of field that blurs backgrounds.
  • Shutter speed is how long the sensor collects light. A long exposure gathers more, but anything that moves during it, including your own hands, turns into blur.
  • ISO is how strongly the sensor’s signal is amplified. Raising it brightens a dim scene without changing the opening or the exposure time, but it amplifies noise along with the signal, so the picture gets grainy.
The exposure triangle A triangle with aperture, shutter speed and ISO at its corners. Aperture controls light and depth of field, and a wide opening gives shallow focus. Shutter speed controls how long light is collected, and slow speeds give motion blur. ISO amplifies the signal, and high values add noise. Each knob has a cost in orange. Aperture Shutter ISO one brightness budget,three ways to spend it Aperture · size of the opening more light through, or less cost: wide gives a thin slice of focus Shutter · how long the sensor collects longer means more light cost: slow gives motion blur, shake ISO · amplification after the fact brightens a dim scene cost: more grain and noise
Every stop of brightness has to come from somewhere. You choose which cost you would rather pay.

Brightness is measured in stops, and each stop is a doubling or halving of the light. That gives the triangle its useful property: the controls trade against each other one stop at a time. Open the aperture by one stop and you can halve the exposure time for the same brightness. The picture is equally bright, but you’ve swapped a shallower focus for less motion blur.

The point isn’t that any corner is better. The right setting depends on what you care about in the scene, and a good photographer decides that first.

The training version

Training a neural network has a comparable shape. You have a budget for how much progress each step can make safely, and three settings that spend it:

  • Learning rate sets how far each parameter update moves. Too large and the loss oscillates or diverges. Too small and training crawls.
  • Batch size is how many examples the gradient is averaged over at each step. A small batch gives a noisy estimate of the direction. A large one gives a cleaner estimate but costs more memory and compute for each step.
  • Warmup means starting with a very small learning rate and ramping it up over the first steps. It keeps the early updates gentle, while the model and the optimiser’s statistics are still in a poor state, so that you can use a bigger learning rate afterwards.
The training triangle A triangle with learning rate, batch size and warmup at its corners, matched to aperture, shutter and ISO. Learning rate sets how far each update moves and a large one can diverge. Batch size sets how many examples each gradient averages, and small batches are noisy. Warmup ramps the learning rate up from a small value; it costs a few slow early steps. Each cost is in orange. Learning rate Batch size Warmup one stability budget,three ways to spend it Learning rate ~ aperture how far each update moves cost: too big diverges, too small crawls Batch size ~ shutter how many examples each gradient averages cost: small is noisy, large costs compute Warmup ~ ISO a gentle start that lets you push harder later cost: slow early steps, one more knob
Each photographic knob is paired with a training knob. The first two pairs are fairly natural. The third is looser, and I say why below.

The pairing I’m using is aperture with learning rate, shutter with batch size, and ISO with warmup. I’ll explain each pair in turn, and mark how strong I think it is.

Aperture and learning rate (strong). Both are the main lever. They decide how big a step you take per unit of effort, and both have a sweet spot between too little and too much.

Shutter and batch size (fairly strong). A long exposure collects more photons, which averages out random fluctuation. A big batch collects more examples, which averages out sampling noise. In both, more collection reduces noise and costs you something: time on the camera side, compute on the training side.

ISO and warmup (loose). This is the pair I’m least sure about. ISO is a late-stage amplifier: you reach for it when the other two are already stretched. Warmup plays a similar supporting role, because it’s a modest, cheap adjustment that lets the other two be pushed further. But ISO amplifies noise, and warmup does the opposite: it protects you from instability. I’d call this a pairing of roles, not of mechanisms.

What the knobs do, in a simulation

Words are easy here, so I ran the experiment. The setup is deliberately tiny: gradient descent on a two-dimensional quadratic bowl, with random noise added to each gradient to stand in for sampling a mini-batch. It’s pure Python, and it’s illustrative rather than a real model.

Simulated loss curves for different learning rates and batch sizes Two log-scale loss charts over 80 gradient descent steps on a noisy two-dimensional quadratic. Left, at batch size 64: a learning rate of 0.005 falls slowly and is still around 2 at step 80; 0.10 drops to about 0.09 by step 20 and then sits at a noise floor near 0.003; 0.22 blows up and leaves the chart within a few steps. Right, at learning rate 0.10: batch size 1 settles at a noisy floor near 0.2, batch size 8 near 0.03, and batch size 64 near 0.003. LEARNING RATE · BATCH 64BATCH SIZE · LEARNING RATE 0.10 1e-4 1e-2 1 1e2 0 40 80 step lr 0.005 (small): loss 49.5 at step 0, 2.03 at step 80 lr 0.10 (good): loss 49.5 at step 0, 0.000764 at step 80 lr 0.22 (too large): loss 49.5 at step 0, 1e+06 at step 80 1e-4 1e-2 1 1e2 0 40 80 step batch 1: loss 49.5 at step 0, 0.0486 at step 80 batch 8: loss 49.5 at step 0, 0.00609 at step 80 batch 64: loss 49.5 at step 0, 0.000764 at step 80 0.005 0.10 0.22 diverges batch 1 batch 8 batch 64
Simulated, not from a real model: gradient descent on a two-dimensional quadratic with added gradient noise. The largest learning rate that still converges here is 0.2, so 0.22 diverges. Smaller batches average less noise, so the loss stalls higher.

On the left, batch size is fixed and only the learning rate changes:

  • 0.005 is safe but slow. After 80 steps it has still only reached a loss of about 2.
  • 0.10 gets to a low loss in about 20 steps and then hovers at a small noise floor.
  • 0.22 is only a little larger, and it diverges. In this bowl, the steepest direction sets a hard limit: any learning rate above 2 divided by that direction’s curvature (here 10, so 0.2) makes each step overshoot by more than it corrects.

On the right, learning rate is fixed and the batch size changes. The averaging effect is visible. A batch of one gives a noisy trace that settles about two orders of magnitude above where a batch of 64 does. Averaging over more examples shrinks the noise floor, roughly in proportion to one over the batch size.

That last sentence hides the coupling. The height of the noise floor depends on both the learning rate and the batch size, in this simple setting roughly as their ratio. Double the learning rate and you raise the floor. Double the batch and you lower it. That’s the same kind of trade as opening the aperture by a stop and shortening the exposure by a stop, and it’s where the analogy earns its keep.

The trade that carries over

There is a well-known heuristic from large-batch training that has the same shape as the photographer’s reciprocity. If you multiply the batch size by k, multiply the learning rate by about k as well, and add a warmup phase so that the early steps don’t blow up. It’s often called the linear scaling rule, and it comes from work on training image models with very large batches. It isn’t a law. It works up to some batch size and then stops working, and for adaptive optimisers such as Adam, people often find a square-root scaling fits better.

I still find it worth remembering, because it shows why treating the knobs as independent is a mistake. If you double the batch size to “make training more stable” and leave the learning rate alone, you have changed the operating point. The run may get smoother and slower. If you double the learning rate to “make it faster” and leave everything else alone, you have moved towards the edge in the left panel.

A habit I’d suggest, borrowed from the photographers: when you change one control, name the other one you’re spending. “I’m raising the learning rate, so I’m accepting more noise and less margin before divergence.” “I’m shrinking the batch to fit on the device, so I’ll expect a noisier curve and probably a lower learning rate.” Saying it out loud makes the trade visible.

Where the analogy breaks

I’ve pushed the comparison as far as I can. These are the places where it stops helping.

Photography has a target; training doesn’t. An exposure is either about right or it isn’t, and the camera can even show you a light meter. Nothing like that exists for training. There is no learning rate that gives the “correct” amount of progress, only better and worse outcomes on a metric you picked.

The knobs aren’t interchangeable. In a camera, a stop of aperture and a stop of shutter are worth exactly the same amount of light, so the trade is precise. In training, batch size and learning rate coupling is a rough heuristic, and it depends on the model, the optimiser and the data. I’d treat any specific scaling rule as a starting guess to test.

Warmup isn’t a peer of the other two. ISO is a number you set once per shot. Warmup is a schedule, with a length and a shape, that runs for the first part of training and then disappears. It matters most for large models and adaptive optimisers, and it does very little in my toy problem, where the loss is a smooth bowl and the only limit is curvature. Also, real training has more knobs than three: weight decay, the schedule after warmup, gradient clipping, the optimiser choice. A triangle undercounts.

The costs are different in kind. In photography every trade lands on a single image, and you see the result at once. In training the cost shows up later, in a curve you read after minutes or days, and sometimes only in an evaluation you weren’t watching. That delay is the reason I think “name what you’re spending” is such a useful habit for training and less necessary in a camera.

Noise means different things. Grain in a photo is nearly always unwanted. Noise in gradient descent is sometimes a nuisance and sometimes useful, because it can help a run escape sharp regions. A smaller batch isn’t just a worse estimate.

What I take from it

I don’t think the analogy predicts anything. It gives me a way to explain to a colleague, or to myself, why “just change the learning rate” is rarely a single-variable question. It also reminds me to start from what I care about, the way you would with a subject that moves or a scene that’s dim.

When I set up or review a training run, I try to ask three questions in that order. What’s the failure I’m most worried about: divergence, slow progress, or noisy results? Which knob buys me the most protection against it? And which cost am I accepting in exchange? Then I look at the curve, the way I’d look at the back of the camera, and adjust.

References

  1. Bottou, L., Curtis, F. E., and Nocedal, J. Optimization Methods for Large-Scale Machine Learning. SIAM Review, 2018. https://arxiv.org/abs/1606.04838 — stochastic gradient noise and step size
  2. Smith, S. L. and Le, Q. V. A Bayesian Perspective on Generalization and Stochastic Gradient Descent. ICLR, 2018. https://arxiv.org/abs/1710.06451 — SGD noise scale set by learning rate and batch size
  3. Goyal, P. et al. Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour. arXiv, 2017. https://arxiv.org/abs/1706.02677 — the linear scaling rule and warmup
  4. Malladi, S., Lyu, K., Panigrahi, A., and Arora, S. On the SDEs and Scaling Rules for Adaptive Gradient Algorithms. NeurIPS, 2022. https://arxiv.org/abs/2205.10287 — square-root scaling for Adam-style optimisers
  5. Liu, L. et al. On the Variance of the Adaptive Learning Rate and Beyond. ICLR, 2020. https://arxiv.org/abs/1908.03265 — why warmup helps adaptive optimisers
  6. Kalra, D. S. and Barkeshli, M. Why Warmup the Learning Rate? Underlying Mechanisms and Improvements. NeurIPS, 2024. https://arxiv.org/abs/2406.09405