← All posts

Probability and distributions, one step at a time

Probability from scratch, with a die and a coin. What a chance really promises, what a distribution is, and why a bell curve is read by area.

You flip a coin and get heads. You flip it again and get heads again. Is the next one “due” to be tails?

It feels like it should be. It isn’t. A fair coin has no memory, so the third flip is still 50/50. That part is easy to remember. What took me much longer was understanding why, and what a number like “50%” actually promises. Every textbook I tried opened with notation, and I usually lost the thread somewhere on the second page.

So this post goes the other way round. We do each calculation by hand first, with a die or a coin, and only write the formula once you’ve already done what it says. I’ve written it for someone who has never studied probability. If you have, I hope it still helps when you need to explain it to someone else.

The route:

  1. Probability as counting, with one die roll.
  2. Combining questions: not, or, and, given.
  3. Distributions: every possible answer at once.
  4. Two numbers that summarise a distribution.
  5. From bars to curves, and the normal distribution.
  6. Why averages end up looking like bells.

Probability is counting, when things are fair

Roll an ordinary six-sided die. What’s the chance of an even number?

You probably already know it’s a half. Here is how you got there, slowly, because the same three moves come back in everything that follows:

  1. List everything that could happen. 1, 2, 3, 4, 5, 6. This list is called the sample space.
  2. Mark the outcomes you care about. 2, 4 and 6. A group of outcomes like this is called an event.
  3. Divide. 3 marked out of 6 in total.
\[P(\text{even}) = \frac{3}{6} = 0.5\]

\(P(\ldots)\) is shorthand for “the probability of …”. 0.5, 50% and 1/2 are the same number written three ways. A probability always sits between 0, which means it can’t happen, and 1, which means it’s certain.

Three of six equally likely die outcomes are even Six equal boxes labelled one through six. Two, four and six are blue; one, three and five are unfilled. Three of six boxes count. 123456 2, 4 and 6 count: 3 out of 6 = 50%
Six equally likely faces, three of them even. Three out of six is a half.

The divide step only works because every face is equally likely. If someone weighted the die so that 6 comes up half the time, counting faces would give the wrong answer, and you’d have to add up each face’s own chance instead. So “count and divide” is a shortcut for fair situations, not a general rule.

What does 50% actually promise?

Less than it sounds like. Roll the die ten times and you might get five evens. You might also get three, or eight. A 50% chance says nothing firm about a short run.

What it does describe is the long run. Roll hundreds or thousands of times, and the share of evens drifts closer and closer to a half. This is called the law of large numbers. You can watch it happen:

Notice what the line doesn’t do. After a run of odd numbers, the die doesn’t produce extra evens to catch up. The early streak just gets diluted by thousands of later rolls until it hardly moves the average. That’s the real answer to the coin question at the top. Tails isn’t due. The streak simply stops mattering as more flips pile up.

Combining questions: not, or, and

Same die. Let’s name two events:

  • A: the roll is even (2, 4 or 6).
  • B: the roll is over 4 (5 or 6).

Not A is whatever is left over: 1, 3 and 5. Something has to happen, so the chances of A and of not-A always add up to 1. That means you can get one from the other:

\[P(\text{not }A) = 1 - P(A) = 1 - \tfrac{3}{6} = \tfrac{3}{6}\]

A or B means at least one of them happens: 2, 4, 5 or 6. That’s 4 out of 6. If you’d simply added 3/6 and 2/6 you would get 5/6, which is wrong, because 6 is in both lists and you counted it twice. So add, then take the double-counted overlap away once:

\[P(A\text{ or }B) = P(A) + P(B) - P(A\text{ and }B) = \tfrac{3}{6} + \tfrac{2}{6} - \tfrac{1}{6} = \tfrac{4}{6}\]

A and B means both happen on the same roll: even and over 4. Only 6 fits, so it’s 1 out of 6.

Not, or, and on one die roll Six boxes per row for the faces 1 to 6. A, even, marks 2, 4 and 6: 3 of 6. B, over 4, marks 5 and 6: 2 of 6. A or B marks 2, 4, 5 and 6: 4 of 6. A and B marks only 6: 1 of 6. One roll, four questionsA: even1234563/6B: over 41234562/6A or B1234564/6A and B1234561/6
Read each row as "which faces count?". The two rows below the line are built from the two above. Six is in both A and B, which is why "or" is 4/6 and not 5/6.

Every rule here is a counting trick. The formulas only exist to save you from listing outcomes when the list gets long.

“Given”: new information shrinks the list

Now a friend rolls the die behind a book and tells you “it’s over 4”. What’s the chance it’s even?

You don’t count out of six any more. Only 5 and 6 are still possible, so that’s your new list. One of the two is even, so the answer is 1/2.

This is conditional probability: the chance of A given that B happened. The vertical bar is read “given”:

\[P(A \mid B) = \frac{P(A\text{ and }B)}{P(B)} = \frac{1/6}{2/6} = \frac{1}{2}\]

The formula does exactly what you just did in your head. The top is “outcomes in both”. The bottom is “outcomes that survived the new information”. Dividing by the bottom is how you switch from counting out of six to counting out of two.

Conditioning shrinks the sample space Without information, 3 of 6 outcomes are even. Given greater than 4, only 5 and 6 remain, so 1 of 2 is even. New information changes the denominator123456Before: 3 even out of 6 = 1/2Given greater than 4: keep only 5, 656After: 1 even out of 2 = 1/2
Before your friend speaks, you count out of six. Afterwards, only 5 and 6 are left, so you count out of two.

One trap worth spotting early: the order matters. “Even, given over 4” is 1/2. Flip it round to “over 4, given even” and the list becomes 2, 4, 6, of which only 6 is over 4, so it’s 1/3. Swapping the two sides of the bar changes the question and the answer. Mixing the two up causes a surprising amount of bad reasoning, in medicine, in court and in data work, so it’s worth getting used to the difference now.

Independence: when the news changes nothing

Look at that first answer again. Before your friend said anything, the chance of even was 1/2. After hearing “over 4”, it was still 1/2. The news didn’t change anything.

When that happens, the two events are independent:

\[P(A \mid B) = P(A)\]

Compare it with the news “it’s over 3”. Now the list is 4, 5, 6, and two of those three are even, so the chance jumps to 2/3. “Even” and “over 3” are not independent on this die. Knowing one tells you something about the other.

Coin flips are the classic independent events. The coin has no memory, so the first flip tells you nothing about the second. That gives a shortcut for “and”: multiply. Half the time the first flip is heads, and in half of those cases the second is heads too. Half of a half is a quarter:

\[P(\text{heads, then heads}) = 0.5 \times 0.5 = 0.25\]

In general, only when A and B are independent:

\[P(A\text{ and }B) = P(A) \times P(B)\]

That condition matters more than it looks. Two visits from the same user, or two people in the same household, are often related. Multiplying their chances as if they weren’t is one of the most common mistakes I see in real data work.

A distribution: every answer at once

So far each question had a single answer. Now let’s ask one with several possible answers.

Flip a coin four times and count the heads. You could get 0, 1, 2, 3 or 4. How likely is each?

The tempting guess is “five possible answers, so 20% each”. Let’s check by listing. Four flips can come out in 2 × 2 × 2 × 2 = 16 different orders (HHHH, HHHT, HHTH, and so on), and with a fair coin each order has the same chance, 1/16. Now sort the orders by how many heads they contain:

Sequences group into heads counts All sixteen equally likely sequences of four fair coin flips grouped into 0, 1, 2, 3 and 4 heads, with counts 1, 4, 6, 4 and 1. 16 equal sequences → 5 unequal counts0 headsTTTT1/161 headsTTTHTTHTTHTTHTTT4/162 headsTTHHTHTHTHHTHTTHHTHTHHTT6/163 headsTHHHHTHHHHTHHHHT4/164 headsHHHH1/16
Every tile is one equally likely order, with chance 1/16. There's one way to get 0 heads but six ways to get 2, so 2 heads is six times as likely.

There is only one way to get 0 heads (TTTT), but there are six ways to get 2. So 2 heads has a chance of 6/16, or 37.5%, while 0 heads has 1/16, about 6%. The five answers are nowhere near equally likely.

That whole table, every possible value together with its chance, is a probability distribution. The thing we’re counting, “heads in four flips”, is called a random variable and is usually written \(X\). Despite the name, it’s just a number whose value depends on chance. The distribution is the complete picture of what \(X\) might turn out to be, and its chances always add up to 1, because one of the answers has to happen.

The binomial: a recipe for counting heads

Listing 16 orders is fine. Listing them for 20 flips, over a million orders, is not. We need a recipe.

Here is how to get the chance of exactly 2 heads in 4 flips without listing everything:

  1. Find the chance of one particular order. HHTT means heads, heads, tails, tails. The flips are independent, so multiply: 0.5 × 0.5 × 0.5 × 0.5 = 1/16.
  2. Count the orders that give 2 heads. There are six: HHTT, HTHT, HTTH, THHT, THTH, TTHH.
  3. Multiply the two. 6 × 1/16 = 6/16.

That’s the whole idea: (number of ways) × (chance of each way). Here it is in general form, for \(n\) flips, \(k\) heads, and a coin that lands heads with chance \(q\):

\[P(X = k) = \binom{n}{k}\, q^k\, (1-q)^{n-k}\]

Piece by piece:

  • \(\binom{n}{k}\), read “n choose k”, is the number of ways to pick which \(k\) of the \(n\) flips are heads. In our example it’s the 6. Spreadsheets have it built in (COMBIN in Excel or Google Sheets).
  • \(q^k\) is \(q\) multiplied by itself \(k\) times: the heads part of one order.
  • \((1-q)^{n-k}\) is the tails part. The other \(n-k\) flips are tails, each with chance \(1-q\).

Put in \(n=4\), \(k=2\), \(q=0.5\) and you get 6 × 0.5² × 0.5² = 6/16, exactly what we counted by hand. The formula doesn’t know anything we didn’t. It just counts faster.

This is the binomial distribution. It fits anything shaped like “\(n\) independent tries, each one a yes or no with the same chance”. Coin flips, but also “how many of 1,000 visitors buy something”, if each visitor buys independently with the same chance. (You’ll also see the name Bernoulli for a single try: 1 with chance \(q\), 0 otherwise. It’s just one flip.)

With a coin that lands heads 70% of the time, the same six orders each have chance 0.7² × 0.3² = 0.0441, so exactly 2 heads becomes 6 × 0.0441 ≈ 26%. The bias pushes the whole distribution towards more heads. Try it:

Two numbers that summarise a distribution

A distribution is a lot of numbers. Most of the time you want just two of them: where is it centred, and how spread out is it?

The centre: expected value

Flip a coin that lands heads 70% of the time, ten times. How many heads on average? Each flip contributes 0.7 of a head on average, so ten flips give 7. In general:

\[E[X] = n \times q\]

\(E[X]\) is read “the expected value of \(X\)”. The name is a bit misleading. It’s not the result you should expect on any particular try. It’s the long-run average if you repeat the whole thing many times. One fair flip has an expected value of 0.5 heads, and no coin has ever landed half-heads.

The spread: standard deviation

Two distributions can share the same centre and still look completely different. Ten fair flips average 5 heads, but do batches usually land between 4 and 6, or all over the place?

The standard deviation (SD) measures that. Roughly, it’s how far a typical result lands from the centre. For the binomial:

\[SD(X) = \sqrt{n\,q\,(1-q)}\]

You don’t need to derive it, but you can check it against your intuition:

  • More flips (bigger \(n\)) leave more room to wander, so the SD grows.
  • \(q(1-q)\) is largest when \(q = 0.5\). A fair coin is the hardest one to predict.
  • A coin that always lands heads (\(q = 1\)) gives an SD of 0. No randomness, no spread.

For ten fair flips, that’s √(10 × 0.5 × 0.5) ≈ 1.58 heads. Most batches land within a couple of heads of 5. The widget above shows the SD for whatever you set.

The thing under the square root, \(n\,q\,(1-q)\), has its own name: the variance. It’s measured in “heads squared”, which isn’t a unit anyone can picture, so we take the square root to get back to plain heads.

From bars to curves

Everything so far had gaps between the possible values: 2 heads or 3 heads, never 2.5. These are called discrete distributions, and each value gets its own bar.

Now think about something like a person’s height, or how long you wait for a bus. Between 170 cm and 171 cm there are infinitely many possible values. You can’t give each of them its own slice of probability, because the slices have to add up to 1 and there are infinitely many.

The way out is to stop asking about exact values and ask about ranges instead. Picture a spinner that can stop anywhere between 0 and 10, with no position favoured. What’s the chance it stops between 2 and 4? That stretch is 2 of the 10 units, so 20%. Between 2 and 3? 10%. Exactly on 3.000…? Zero, because a single point has no width.

So we draw a curve and read probability as area under the curve. For the spinner the “curve” is a flat line at height 0.1, and the area between 2 and 4 is width × height = 2 × 0.1 = 0.2. The total area under the whole curve is always 1.

The height of the curve is called the density. It shows where results are packed tightly, but it isn’t a probability by itself. This was the single thing that confused me most when I first met bell curves: I kept reading the height as a chance. It isn’t. Chance is area.

Discrete mass versus continuous area Top: exact probabilities for heads in four fair flips. Bottom: normal density with the interval within one standard deviation shaded, about 68.3 percent probability. Discrete: probability belongs to a bar01234Heads in 4 fair flipsContinuous: probability is area-1σμ+1σShaded area ≈ 68.3%
Top: separate values, each with its own bar. Bottom: a continuous curve, where you get a chance by adding up the area over a range. About 68% of the area sits in the shaded middle.

The flat spinner model is called the uniform distribution. The fair die is its discrete cousin: six bars of equal height.

The normal distribution

The bell in that picture is the normal distribution. It turns up everywhere, for a reason we’ll get to at the end. It has just two settings:

  • the mean, written \(\mu\) (“mu”), which sets where the centre is;
  • the standard deviation, written \(\sigma\) (“sigma”), which sets how wide it is.

Change the mean and the whole bell slides sideways. Increase \(\sigma\) and the bell gets wider and flatter. It has to get flatter, because the total area has to stay 1.

Measuring in standard deviations

Here is the useful part. Say exam scores follow a normal curve with mean 100 and SD 10. A score of 120 is 20 points above the mean, which is 2 standard deviations above it. A score of 90 is 1 SD below.

That “how many SDs away” number is called a z-score:

\[z = \frac{x - \mu}{\sigma}\]

Subtract the mean to get the distance, then divide by the SD to measure that distance in units of spread. For 120: (120 − 100) / 10 = 2. For 90: −1. A negative z just means below the mean.

Why bother? Because on every normal curve, whatever its mean or SD, the same share of the area sits within the same number of SDs of the centre:

  • about 68% within 1 SD,
  • about 95% within 2 SDs,
  • about 99.7% within 3 SDs.

So without knowing anything else about the exam, you know that roughly 95% of scores land between 80 and 120. That rule of thumb is worth memorising. It only holds for things that really are close to normal, though. Incomes, for example, aren’t: a few very large values stretch the right side out.

Try it. The axis stays fixed, so you can see the bell move and stretch:

You don’t need the equations below to follow the rest of the post. They’re here for when you see them somewhere else and want to know what they’re doing.

The equation that draws the bell

A bell needs three properties. It should be tallest at the centre, drop off on both sides, and have a total area of 1. The normal formula builds those in one at a time:

\[f(x)=\frac{1}{\sigma\sqrt{2\pi}}\,e^{-\frac12\left(\frac{x-\mu}{\sigma}\right)^2}\]

Read it from the inside out:

  1. \(\frac{x-\mu}{\sigma}\) is the z-score from above: how many SDs \(x\) is from the centre.
  2. Squaring it makes being 2 SDs below the centre count the same as 2 SDs above.
  3. \(e^{-z^2/2}\) is 1 at the centre and shrinks quickly as you move away. That’s the bell shape. (\(e \approx 2.718\) and \(\pi \approx 3.14\) are just constants.)
  4. The fraction in front scales the whole thing so the area comes out as exactly 1. It has \(\sigma\) on the bottom, which is why a wider bell is lower.

\(f(x)\) is the density at \(x\), the height of the curve. For mean 0 and SD 1, it’s about 0.399 at the centre and 0.242 one SD away. Those are heights, not 39.9% and 24.2% chances. For a chance, you still need area.

The equation that counts the area

The area between two points is “all the area left of the right end” minus “all the area left of the left end”.

For the standard normal curve (mean 0, SD 1), the area to the left of a point \(z\) has its own symbol, \(\Phi(z)\) (“phi”). About 84.1% of the area is left of 1, and about 15.9% is left of −1. Their difference is 68.3%, the number from the rule of thumb.

For any other normal curve, turn both ends into z-scores first, then subtract:

\[P(a \le X \le b)=\Phi\!\left(\frac{b-\mu}{\sigma}\right)-\Phi\!\left(\frac{a-\mu}{\sigma}\right)\]

For the exam example, 90 to 110 becomes −1 to 1, and you get the same 68%. You’ll also see \(\Phi\) called the cumulative distribution function, or CDF: “cumulative” because it has added up all the area so far.

Why so many things look like bells

Back to the coin, one last time. Instead of counting heads, record the fraction of heads in each batch: 3 heads out of 10 flips is 0.3. Then look at the distribution of that fraction.

With one flip per batch, the fraction is either 0 or 1. Two bars, nothing like a bell. With five flips you get six bars and the shape starts to round off. With 30 or 100 it looks very much like a bell, and the bars bunch up tighter around the middle.

Two separate things are happening there, and both matter a lot in practice.

The shape turns into a bell. Averages of many independent pieces tend towards a normal distribution, even when each piece looks nothing like a bell. Each flip here is just a 0 or a 1. This is the central limit theorem, and it’s the reason the normal curve shows up so often: plenty of real measurements are, in effect, sums or averages of many small independent things. How many pieces you need depends on what you start with. A lopsided coin needs far bigger batches before the bell fits, as you saw.

The bell gets narrower. The more flips you average, the less the average wobbles from batch to batch. The spread of an average has its own name, the standard error (SE):

\[SE = \frac{\sigma}{\sqrt{n}}\]

Here \(\sigma\) is the SD of one single observation and \(n\) is how many you average. A single fair flip, recorded as 0 or 1, has \(\sigma = 0.5\). Average 100 flips and the SE is 0.5 / √100 = 0.05, so the fraction of heads typically lands within about 5 percentage points of 50%, and 95% of the time within about two SEs, between 40% and 60%. Average 400 flips and the SE drops to 2.5 points.

Look at the square root. Four times the data only halves the wobble. Getting more precise gets expensive quickly, which is the whole reason A/B tests need so many users.

Standard deviation and standard error are easy to mix up, so to be explicit:

  • Standard deviation is the spread of individual results: single flips, single people.
  • Standard error is the spread of an average, if you repeated the whole batch many times.

Cheat sheet

Idea In one line
Probability Count what you want, divide by everything possible (if all outcomes are equally likely)
Long run A probability describes many tries, not the next few
Not 1 minus the chance that it happens
Or Add the chances, then subtract the overlap
And Multiply, but only if the events are independent
Given New information shrinks the list you count from
Independent Learning one thing doesn’t change the chance of the other
Distribution Every possible value with its chance, adding up to 1
Binomial (Number of ways) × (chance of each way), for n yes/no tries
Expected value The long-run average
Standard deviation How far a typical result lands from the centre
Density Height of a curve. Probability is the area under it
Normal A bell set by its mean and SD. 68 / 95 / 99.7% within 1 / 2 / 3 SDs
Central limit theorem Averages of many independent pieces come out bell-shaped
Standard error How much an average wobbles: σ / √n

Further reading