Deconstructing PlanOut: a coin toss you can repeat
Follow one shopper from a button experiment to the actual SHA-1 bytes, then look at salts, namespaces and the gap between assignment and exposure.
Andi opens a checkout page. We want to test two buttons: A is blue, B is pink. He gets blue today. Tomorrow, he opens the page on another server. He should still get blue.
That sounds like a small requirement. It rules out the most obvious implementation: toss a fresh coin on every request. You could save every user’s choice in a database, but then every server needs access to that record. PlanOut offers another way: make the coin toss repeatable from the inputs.
This is the part of experimentation systems I want to take apart here. One shopper, one button, and the actual numbers from the archived Python implementation. No production traffic or company internals: the example below uses invented IDs.
What PlanOut was trying to separate
The product owns the checkout page. The experiment owns the rule for choosing the button. Those should not become one tangled piece of code.
PlanOut was developed at Facebook as a language and framework for defining online experiments. Its 2014 paper describes how an experiment maps a unit, such as a user ID, to parameter values. A parameter might be a button colour, a piece of text, or a discount. Product code asks for that value and uses it.
In other words, the experiment says “blue”. It does not draw the button, count purchases, or tell us whether blue is better.
Here is a small definition using PlanOut’s Python API:
from planout.experiment import SimpleExperiment
from planout.ops.random import UniformChoice
class CheckoutButton(SimpleExperiment):
def setup(self):
self.salt = "checkout-v1"
def assign(self, params, userid):
params.button = UniformChoice(
choices=["blue", "pink"], unit=userid
)
experiment = CheckoutButton(userid="user-42")
colour = experiment.get("button") # "blue"
Andi’s invented ID is user-42. It is the unit: the thing we randomise. If we used a session ID instead, we’d be randomising sessions, and Andi could get a different button on his next visit. Choosing the unit is a design decision, not a detail to leave to the SDK.
A coin toss without a coin
A hash function turns an input string into a fixed-length output. The same input always gives the same output. Small changes to the input usually produce very different outputs.
PlanOut uses SHA-1 as a mixing function for assignment. This is not a password scheme or a way to hide someone’s group. We want repeatable, well-spread values, not secrecy.
For this example, PlanOut joins three pieces with dots:
- The experiment salt:
checkout-v1. - The parameter salt:
button, supplied from the assigned variable name. - The unit:
user-42.
The complete key is checkout-v1.button.user-42. The default separator is a dot; explicit salt overrides and multi-part units have their own paths in the source.
Try the coin toss
Change the user ID and follow the actual calculation. The salt stays checkout-v1, the parameter stays button. Nothing is sent to a server.
Computing the assignment...
- 1. Full key: salt.parameter.unit
- 2. SHA-1 of the UTF-8 key
- 3. Keep the first 15 hex digits (60 bits)
- 4. Convert hex to an integer
- 5. Remainder selects a choice
SHA-1 produces 40 hexadecimal digits. The Python implementation keeps the first 15. Each hex digit holds four bits, so that prefix gives a 60-bit integer. For Andi:
key: checkout-v1.button.user-42
SHA-1: 3441f9fc514ea2ac8f04b1c44a64d90aa350f4fb
first 15: 3441f9fc514ea2a
integer h: 235347851596851754
h % 2: 0
choices[0]: blue
% means remainder. Dividing an even number by 2 leaves remainder 0; an odd number leaves 1. With two choices, that is the whole UniformChoice decision: choices[h % 2].
A thousand servers can repeat it. None needs a stored row saying “Andi got blue”, as long as they agree on the input and implementation.
The bucket picture, with one important distinction
It is tempting to explain every assignment as “put the hash on a line from zero to one, then split the line in half”. That is a useful picture for weighted choices. It is not the exact code path for UniformChoice.
The archived Python implementation has two different operations:
UniformChoice: use the integer’s remainder to select a list position.WeightedChoice: divide the integer by a scale, then compare it with cumulative weights.
For the second operation, the scale is 2**60 - 1, or 1,152,921,504,606,846,975. Andi’s normalized value is about 0.2041317216. With weights [0.5, 0.5], it lands in blue’s half of the line.
Both operations happen to give Andi blue here. That does not mean they assign every user identically. Replacing one with the other can move users even if both advertise a 50/50 split.
A small source-level detail matters if you want compatibility: PlanOut converts that scale to a floating-point value, and uses <= at the weighted boundaries. The mathematical normalization includes 1 at the maximum hash. Replacing the denominator with 2**60, changing the comparison, or switching to another port without checking its hash behaviour is not a harmless cleanup.
A salt is the name of the shuffle
Think of the salt as the label on a shuffled deck. With the same deck label and the same user ID, draw the same card again. Change the label and you get a new shuffle.
Here are computed results for 10,000 invented users, user-0 through user-9999, using the same key construction and UniformChoice modulo rule:
| Change | Users whose button changed |
|---|---|
Repeat checkout-v1, same choices |
0 / 10,000 |
Change salt to checkout-v2 |
4,984 / 10,000 |
Keep salt, reverse [blue, pink] |
10,000 / 10,000 |
The original split was 5,004 blue and 4,996 pink. These are reproducible results for this toy population, not a promise that every experiment produces those exact counts.
Change the shuffle, not the users
These are the same 200 IDs, user-0 through user-199. Each dot uses SHA-1 and the first 15 hex digits, then h % 2. Change only the salt and see which users cross over.
Salt: checkout-v1 · parameter: button
Computing 200 assignments...
For Andi, the new salt produces prefix 5c87192d9dc8829, integer 416707841066043433, and remainder 1. He moves to pink. Reversing the original choices also gives him pink, but for a different reason: his original index is still 0; index 0 now means pink.
That last row is the trap. Keeping the hash stable is not enough. Keep the meaning of its output stable too.
A new experiment can use a new salt deliberately. A live experiment should not casually rename its salt, change ID formatting, reorder choices, edit weights, or switch hashing algorithms. Any of those can change the experience for people already enrolled. Stateless repeatability is not the same thing as a saved sticky assignment that survives definition changes.
The parameter name acts as a salt too. button and hint get different hash keys for the same user, unless you intentionally share an explicit salt. This keeps two parameters from being perfectly tied to each other by accident.
Namespaces: pick the experiment before the button
Now suppose two teams both want to test the checkout page. One tests the button; the other tests the layout. We do not want Andi in both at the same time.
A namespace is a shared allocation space. In the Python implementation, a user hashes into an integer segment. Experiments reserve disjoint sets of segment IDs. The user’s segment decides which experiment, if any, runs. Only then does that experiment choose its parameter values.
The numbered tiles above are an illustration, not Andi’s computed upstream namespace segment. The real add_experiment() samples from the available segments, so an experiment’s allocation need not be a neat contiguous range.
Mutual exclusion applies inside that namespace: a segment cannot belong to both checkout experiments at once. A separate namespace for the homepage can also enrol Andi. Namespaces do not automatically prevent interactions between experiments on different surfaces.
The namespace name, unit, segment count and allocation history are part of the routing contract. Preserving only the button’s salt will not preserve routing if those change.
Assignment is not exposure
We have computed a button colour. Has Andi seen the button? Not necessarily. The page might fail before rendering it. The application might ask for the colour on a route where checkout is never shown.
PlanOut gives you a logging hook. In the default Python experiment behaviour, the first eligible get() on an instance triggers exposure logging before returning the parameter. The instance remembers that it logged, so repeated get() calls on that same instance do not each emit a new automatic exposure. A new instance is a new situation; there is no global per-user deduplication promise.
That hook records parameter access, not human viewability. Automatic logging can be disabled, and an experiment marked out of experiment does not log an exposure. The base experiment delegates the log to your implementation; SimpleExperiment writes JSON to a file. Getting a production event pipeline right is still your job.
A useful record connects the unit, experiment, parameter values and time to the definition that produced them. Upstream also records a checksum when available. In a production system I would make the definition version, override, surface and event ID explicit, then decide where the application should emit the event that analysis treats as exposure.
A hash can reconstruct an intended choice if the original inputs and definition survive. It cannot reconstruct a missing render event or a fallback that nobody recorded.
What the hash cannot tell you
Run 10,000 repeatable coin tosses
Assign user-0 through user-9999 with checkout-v1.button and h % 2. A 50/50 rule does not promise exactly 5,000 in each group. Run it again: the same inputs give the same counts.
Ready. All computation stays in your browser.
Suppose the button experiment expects an even split, but the logged population contains 700 blue and 300 pink users. That is a sample ratio mismatch, or SRM: the observed group counts are unexpectedly far from the planned allocation.
It is a warning to investigate the experiment, not evidence that pink loses. Maybe assignment changed. Maybe one group has missing events. Maybe the eligibility filter is different on the two paths. Check the population you count, the planned allocation, and the logging before reading an effect estimate. A fixed 50/50 check also does not fit a bandit that intentionally moves traffic.
PlanOut does not perform that investigation or estimate lift for you. Repeatable assignment is one part of trustworthy experimentation. Exposure data, outcome definitions and statistical analysis are separate parts.
Archived does not mean the idea failed
The public repository was archived on June 1, 2021 and is read-only. That is verifiable. An exact public explanation from the maintainers for why it was archived is not established in the sources linked below.
So I would not turn the archive badge into a story about Facebook abandoning experimentation, maintenance costs, or a single successor replacing PlanOut. Its own documentation described the interpreter as one way to define experiments at Facebook, running on top of QuickExperiment. The public library was never the whole internal platform.
What has changed for someone choosing a tool today is the available scope. A language gives you definitions and operators. Modern experimentation products can also provide configuration delivery, allocation controls, exposure handling, metric definitions and analysis. Those are reasons to evaluate a broader platform, not proof of the archive’s motive.
What I would use instead today
There is no official replacement established by the PlanOut sources. The replacement depends on which job you need done.
- For owning the assignment implementation, a small SDK can work, but stable identity, versioned definitions and compatibility tests become your responsibility. GrowthBook’s build-your-own guide is a useful current example of an SDK contract.
- For an experimentation platform, GrowthBook and Statsig are examples worth evaluating. Compare their allocation, logging, warehouse and analysis paths against your requirements. They are alternatives, not drop-in replicas of PlanOut’s byte-level behaviour.
- For adaptive allocation, the policy is another layer. A bandit updates traffic based on rewards; a deterministic hash does not learn which button sells more. I covered that distinction in Thompson sampling, explained with coupons.
If I were moving an existing experiment, I would first make a test set of IDs, salts, definitions, overrides and expected outputs. Run old and new systems against it. Explain every disagreement before moving live traffic. A new SDK producing “roughly half in each group” is not enough if it puts different people in those groups.
The lesson I keep from PlanOut is smaller than a whole platform: separate the experiment definition from product code, make assignment repeatable, and make the boundary between a decision and an observation explicit. Andi getting blue twice is easy to demonstrate. Knowing whether blue helped him is a different system.
References
- Bakshy, Eckles and Bernstein, Designing and Deploying Online Field Experiments, WWW 2014. The design motivation and system boundaries.
- PlanOut repository and release page. Public archival status and date.
- Python source: random operators, assignment, namespaces, and experiment logging. The mechanics here refer to this implementation, not every language port.
- Fabijan et al., Diagnosing Sample Ratio Mismatch, KDD 2019. Why the count mismatch is a diagnostic warning, not a treatment-effect result.
- About PlanOut. Its relationship to QuickExperiment.
- GrowthBook SDK guide, sticky bucketing, and Statsig Experiments. Examples of current alternatives, not evidence of an official succession.