← All posts

Deconstructing PlanOut: a coin toss you can repeat

Follow one shopper from a button experiment to the actual SHA-1 bytes, then look at salts, namespaces and the gap between assignment and exposure.

Andi opens a checkout page. We want to test two buttons: A is blue, B is pink. He gets blue today. Tomorrow, he opens the page on another server. He should still get blue.

That sounds like a small requirement. It rules out the most obvious implementation: toss a fresh coin on every request. You could save every user’s choice in a database, but then every server needs access to that record. PlanOut offers another way: make the coin toss repeatable from the inputs.

This is the part of experimentation systems I want to take apart here. One shopper, one button, and the actual numbers from the archived Python implementation. No production traffic or company internals: the example below uses invented IDs.

What PlanOut was trying to separate

The product owns the checkout page. The experiment owns the rule for choosing the button. Those should not become one tangled piece of code.

PlanOut was developed at Facebook as a language and framework for defining online experiments. Its 2014 paper describes how an experiment maps a unit, such as a user ID, to parameter values. A parameter might be a button colour, a piece of text, or a discount. Product code asks for that value and uses it.

In other words, the experiment says “blue”. It does not draw the button, count purchases, or tell us whether blue is better.

Flow Andi enters the checkout. The experiment chooses blue. The product renders the button. The outcome pipeline counts purchases separately. Andi opens checkoutunit: user-42Experiment definitionchoose button = blueProduct codeuse blue when rendering checkoutOutcome pipelinerecord purchases; analyse later
PlanOut returns parameter values. Rendering and measuring the result live outside that decision.

Here is a small definition using PlanOut’s Python API:

from planout.experiment import SimpleExperiment
from planout.ops.random import UniformChoice

class CheckoutButton(SimpleExperiment):
    def setup(self):
        self.salt = "checkout-v1"

    def assign(self, params, userid):
        params.button = UniformChoice(
            choices=["blue", "pink"], unit=userid
        )

experiment = CheckoutButton(userid="user-42")
colour = experiment.get("button")  # "blue"

Andi’s invented ID is user-42. It is the unit: the thing we randomise. If we used a session ID instead, we’d be randomising sessions, and Andi could get a different button on his next visit. Choosing the unit is a design decision, not a detail to leave to the SDK.

A coin toss without a coin

A hash function turns an input string into a fixed-length output. The same input always gives the same output. Small changes to the input usually produce very different outputs.

PlanOut uses SHA-1 as a mixing function for assignment. This is not a password scheme or a way to hide someone’s group. We want repeatable, well-spread values, not secrecy.

For this example, PlanOut joins three pieces with dots:

  1. The experiment salt: checkout-v1.
  2. The parameter salt: button, supplied from the assigned variable name.
  3. The unit: user-42.

The complete key is checkout-v1.button.user-42. The default separator is a dot; explicit salt overrides and multi-part units have their own paths in the source.

Hash The exact key checkout-v1.button.user-42 is hashed. Its first 15 hex digits become integer 235347851596851754. Remainder modulo two is zero, selecting blue. checkout-v1.button.user-42salt + parameter + unitSHA-1 → keep first 15 hex digits3441f9fc514ea2aConvert hex to integer235347851596851754h % 2 = 0 → choices[0] → bluesame key, same answer
The computed path for Andi. Nothing is freshly drawn on the second request.

SHA-1 produces 40 hexadecimal digits. The Python implementation keeps the first 15. Each hex digit holds four bits, so that prefix gives a 60-bit integer. For Andi:

key:          checkout-v1.button.user-42
SHA-1:        3441f9fc514ea2ac8f04b1c44a64d90aa350f4fb
first 15:     3441f9fc514ea2a
integer h:    235347851596851754
h % 2:        0
choices[0]:   blue

% means remainder. Dividing an even number by 2 leaves remainder 0; an odd number leaves 1. With two choices, that is the whole UniformChoice decision: choices[h % 2].

A thousand servers can repeat it. None needs a stored row saying “Andi got blue”, as long as they agree on the input and implementation.

The bucket picture, with one important distinction

It is tempting to explain every assignment as “put the hash on a line from zero to one, then split the line in half”. That is a useful picture for weighted choices. It is not the exact code path for UniformChoice.

The archived Python implementation has two different operations:

  • UniformChoice: use the integer’s remainder to select a list position.
  • WeightedChoice: divide the integer by a scale, then compare it with cumulative weights.

For the second operation, the scale is 2**60 - 1, or 1,152,921,504,606,846,975. Andi’s normalized value is about 0.2041317216. With weights [0.5, 0.5], it lands in blue’s half of the line.

Bucket UniformChoice selects blue using remainder zero. WeightedChoice instead places normalized value 0.2041 in the first half of a line from zero to one. UNIFORMCHOICE: LIST INDEX0 → blueeven hash1 → pinkodd hashAndi: remainder 0WEIGHTEDCHOICE: CUMULATIVE WEIGHTSbluepink00.51Andi: 0.2041
Two different operators. The example uses the same blue/pink order and equal weights, but matching probabilities do not guarantee matching users.

Both operations happen to give Andi blue here. That does not mean they assign every user identically. Replacing one with the other can move users even if both advertise a 50/50 split.

A small source-level detail matters if you want compatibility: PlanOut converts that scale to a floating-point value, and uses <= at the weighted boundaries. The mathematical normalization includes 1 at the maximum hash. Replacing the denominator with 2**60, changing the comparison, or switching to another port without checking its hash behaviour is not a harmless cleanup.

A salt is the name of the shuffle

Think of the salt as the label on a shuffled deck. With the same deck label and the same user ID, draw the same card again. Change the label and you get a new shuffle.

Here are computed results for 10,000 invented users, user-0 through user-9999, using the same key construction and UniformChoice modulo rule:

Change Users whose button changed
Repeat checkout-v1, same choices 0 / 10,000
Change salt to checkout-v2 4,984 / 10,000
Keep salt, reverse [blue, pink] 10,000 / 10,000

The original split was 5,004 blue and 4,996 pink. These are reproducible results for this toy population, not a promise that every experiment produces those exact counts.

Salt Among ten thousand users, repeating the definition changes zero assignments, changing the experiment salt changes 4984, and reversing the choices changes all ten thousand. USERS CHANGED / 10,000Same salt, same choices0 / 10,000New salt4,984 / 10,000Reversed choices10,000 / 10,000
Computed on invented user-0 through user-9999. Bars share the same scale; the input change, not new coin tosses, moves users.

For Andi, the new salt produces prefix 5c87192d9dc8829, integer 416707841066043433, and remainder 1. He moves to pink. Reversing the original choices also gives him pink, but for a different reason: his original index is still 0; index 0 now means pink.

That last row is the trap. Keeping the hash stable is not enough. Keep the meaning of its output stable too.

A new experiment can use a new salt deliberately. A live experiment should not casually rename its salt, change ID formatting, reorder choices, edit weights, or switch hashing algorithms. Any of those can change the experience for people already enrolled. Stateless repeatability is not the same thing as a saved sticky assignment that survives definition changes.

The parameter name acts as a salt too. button and hint get different hash keys for the same user, unless you intentionally share an explicit salt. This keeps two parameters from being perfectly tied to each other by accident.

Namespaces: pick the experiment before the button

Now suppose two teams both want to test the checkout page. One tests the button; the other tests the layout. We do not want Andi in both at the same time.

A namespace is a shared allocation space. In the Python implementation, a user hashes into an integer segment. Experiments reserve disjoint sets of segment IDs. The user’s segment decides which experiment, if any, runs. Only then does that experiment choose its parameter values.

Namespace Illustrative noncontiguous slots: button experiment owns 0, 3, 7; layout owns 1, 5. Other slots give defaults. A user first routes to one experiment, then gets its parameters. ONE NAMESPACE · 10 ILLUSTRATIVE SLOTS0123456789Button: slots 0, 3, 7Layout: slots 1, 5Other slots: default experienceUser → segment → one experimentthen: experiment → button value
Two decisions: namespace routing, then treatment assignment. Separate namespaces can still overlap in people.

The numbered tiles above are an illustration, not Andi’s computed upstream namespace segment. The real add_experiment() samples from the available segments, so an experiment’s allocation need not be a neat contiguous range.

Mutual exclusion applies inside that namespace: a segment cannot belong to both checkout experiments at once. A separate namespace for the homepage can also enrol Andi. Namespaces do not automatically prevent interactions between experiments on different surfaces.

The namespace name, unit, segment count and allocation history are part of the routing contract. Preserving only the button’s salt will not preserve routing if those change.

Assignment is not exposure

We have computed a button colour. Has Andi seen the button? Not necessarily. The page might fail before rendering it. The application might ask for the colour on a route where checkout is never shown.

PlanOut gives you a logging hook. In the default Python experiment behaviour, the first eligible get() on an instance triggers exposure logging before returning the parameter. The instance remembers that it logged, so repeated get() calls on that same instance do not each emit a new automatic exposure. A new instance is a new situation; there is no global per-user deduplication promise.

Logging Computing assignment is followed by first eligible get, which triggers a log hook. Rendering comes later and can fail. A purchase is another event. Compute assignmentblue is the intended valueFirst eligible get("button")automatic log hook on this instanceRender buttonmay fail; get() is not viewabilityPurchase eventoutcome, not assignment
A parameter-access log is useful evidence, but it is not proof that a person saw the pixels.

That hook records parameter access, not human viewability. Automatic logging can be disabled, and an experiment marked out of experiment does not log an exposure. The base experiment delegates the log to your implementation; SimpleExperiment writes JSON to a file. Getting a production event pipeline right is still your job.

A useful record connects the unit, experiment, parameter values and time to the definition that produced them. Upstream also records a checksum when available. In a production system I would make the definition version, override, surface and event ID explicit, then decide where the application should emit the event that analysis treats as exposure.

A hash can reconstruct an intended choice if the original inputs and definition survive. It cannot reconstruct a missing render event or a fallback that nobody recorded.

What the hash cannot tell you

Suppose the button experiment expects an even split, but the logged population contains 700 blue and 300 pink users. That is a sample ratio mismatch, or SRM: the observed group counts are unexpectedly far from the planned allocation.

It is a warning to investigate the experiment, not evidence that pink loses. Maybe assignment changed. Maybe one group has missing events. Maybe the eligibility filter is different on the two paths. Check the population you count, the planned allocation, and the logging before reading an effect estimate. A fixed 50/50 check also does not fit a bandit that intentionally moves traffic.

PlanOut does not perform that investigation or estimate lift for you. Repeatable assignment is one part of trustworthy experimentation. Exposure data, outcome definitions and statistical analysis are separate parts.

Archived does not mean the idea failed

The public repository was archived on June 1, 2021 and is read-only. That is verifiable. An exact public explanation from the maintainers for why it was archived is not established in the sources linked below.

So I would not turn the archive badge into a story about Facebook abandoning experimentation, maintenance costs, or a single successor replacing PlanOut. Its own documentation described the interpreter as one way to define experiments at Facebook, running on top of QuickExperiment. The public library was never the whole internal platform.

What has changed for someone choosing a tool today is the available scope. A language gives you definitions and operators. Modern experimentation products can also provide configuration delivery, allocation controls, exposure handling, metric definitions and analysis. Those are reasons to evaluate a broader platform, not proof of the archive’s motive.

What I would use instead today

There is no official replacement established by the PlanOut sources. The replacement depends on which job you need done.

  • For owning the assignment implementation, a small SDK can work, but stable identity, versioned definitions and compatibility tests become your responsibility. GrowthBook’s build-your-own guide is a useful current example of an SDK contract.
  • For an experimentation platform, GrowthBook and Statsig are examples worth evaluating. Compare their allocation, logging, warehouse and analysis paths against your requirements. They are alternatives, not drop-in replicas of PlanOut’s byte-level behaviour.
  • For adaptive allocation, the policy is another layer. A bandit updates traffic based on rewards; a deterministic hash does not learn which button sells more. I covered that distinction in Thompson sampling, explained with coupons.

If I were moving an existing experiment, I would first make a test set of IDs, salts, definitions, overrides and expected outputs. Run old and new systems against it. Explain every disagreement before moving live traffic. A new SDK producing “roughly half in each group” is not enough if it puts different people in those groups.

The lesson I keep from PlanOut is smaller than a whole platform: separate the experiment definition from product code, make assignment repeatable, and make the boundary between a decision and an observation explicit. Andi getting blue twice is easy to demonstrate. Knowing whether blue helped him is a different system.

References