← All posts

From polling to push: getting experiment config to every service

The architecture we had, the one we moved to, and the designs I'd weigh today, for the unglamorous job of keeping experiment config fresh everywhere.

An experimentation platform looks like it’s about statistics. A lot of the engineering is about something plainer: making sure that when someone changes an experiment, every service that assigns users to variants finds out, quickly and cheaply.

I led the modernisation of a platform that ran hundreds of experiments a year for dozens of product teams. This post is about one part of that work, the path config takes from where people edit it to where it’s used. We went from polling to event-driven caching, cut the platform’s cloud bill by 95% and API latency by 90% over a year, and ended up with a much clearer design. I’ll finish with the alternatives I’d consider if I were building it again.

Two sides with very different traffic

It helps to split the platform in two. The control plane is where experiments change: people create them, change traffic splits, pause them, and bandits shift traffic toward the winning variant every few hours. Across all running experiments, that was around ten changes a day. The data plane is where config is used: every request that needs to know which variant a user sees. That was more than 20 backend services, each running many pods on Kubernetes, plus every user of the apps.

Control plane and data plane of an experimentation platform People and a bandit updater write experiment config to a config store about ten times a day. Services and apps ask an assignment component for a variant on every request. The open question is how config gets from the store to assignment. CONTROL PLANE · RARE WRITES DATA PLANE · CONSTANT READS Experiment UI / API create, split, pause, ramp Bandit updater re-splits every few hours Config store Firestore Assignment picks a variant per user Services and apps ask on every request ? ~10 changes a day every request, all day
Two sides with very different traffic. This post is about the dashed arrow: how config travels from the store to wherever assignment happens.

A good design for the arrow in the middle has to meet five requirements at once:

  • Fresh: a paused experiment stops serving quickly. That’s the kill switch.
  • Consistent: the same user gets the same variant from every service.
  • Cheap: cost should grow with changes, not with traffic or time.
  • Fast: assignment sits on the request path, so it can’t add much latency.
  • Resilient: if the config store has a bad minute, assignment keeps working with the last known config.

Where we started

The first architecture had two clients, and each failed one requirement badly.

The backend polled. Every pod asked Firestore for the current config every two minutes. That’s 720 polls a day from every pod, to catch about ten changes spread across the whole company. Firestore bills per document read, so the cost was roughly:

pods × documents per poll × polls per day

The rate of change doesn’t appear in that formula. The bill grew with fleet size and the number of hours in a day. Adding a service, adding experiments, or scaling out for a traffic peak all made it bigger, while the answer stayed “no change” nearly every time.

Firestore reads over a day, polling versus events Illustrative. Polling reads rise in a straight line through the day. Event-driven reads stay low and step up slightly at each config change. Polling Events + cache + TTL Firestore reads, cumulative 00:00 06:00 12:00 18:00 24:00 Polling: reads grow with pods × time, whether or not anything changed Events + cache + TTL: a step per config change, plus a slow slope from TTL reloads grows with pods × time grows with changes
The shape of the problem (illustrative, not measured). Polling pays every pod, every interval. Events pay once per change; the slow slope is TTL reloads.

The apps assigned on the device. The app fetched config at start-up and computed each user’s variant locally. That was cheap, but config only refreshed when the app restarted, and a backgrounded mobile app might not restart for days or weeks. When a team paused a broken experiment, users who hadn’t restarted kept seeing it. It also meant the assignment logic existed twice, in the backend and in the apps, and the two had to stay identical.

So the backend paid to stay fresh when it didn’t need to, and the apps stayed stale when it mattered most.

Where we ended up

The redesign moved assignment into one service and changed how that service learns about changes.

Config flow before and after Before: backend services poll Firestore on a schedule and apps fetch config only at start. After: an edit publishes an event to the assignment service, whose pods cache locally and share Redis, reading Firestore only on a miss. BEFORE Backend services20+ services, many pods Mobile and web appsassign on the device Firestorebilled per document read polls on a schedule, all day fetches once, at app start AFTER Experiment edited~10 changes a day Services and appsask the service to assign Assignment servicelocal cache per podTTL as a backstop Redisshared across pods Firestoreread on a miss only event
Before, the cost path ran from every pod to Firestore all day. After, Firestore is only read when a change or a TTL expiry actually needs it.
  1. One assignment service. Backend services and, eventually, the apps ask the service for variants instead of computing them. One implementation, one place to fix bugs, and a kill switch that works for everyone.
  2. A local cache in each pod. Config is read from memory on the request path. That alone removed most Firestore reads and most of the per-request latency.
  3. A change event on every config update. Saving an experiment, or a bandit re-split, publishes an event, and pods invalidate their cache when they receive it. Reads now scale with the number of changes, not the number of seconds.
  4. A TTL as a backstop. If an event is lost or delayed, the entry expires and gets reloaded anyway.
  5. Redis as a shared layer. Because the service ran as many pods, each local cache was its own copy. Two pods could disagree for a while, and every pod reloaded from Firestore separately. Redis between the pods and Firestore gave them one shared copy and turned many reloads into one.

Here’s what happens when someone pauses an experiment:

sequenceDiagram
  participant UI as Experiment UI
  participant FS as Firestore
  participant EV as Change event
  participant P as Assignment pods
  participant R as Redis
  UI->>FS: pause experiment 42
  UI->>EV: publish "42 changed"
  EV->>P: invalidate 42
  Note over P: next request for 42
  P->>R: get config 42
  R-->>P: miss
  P->>FS: read config 42
  P->>R: store config 42 with TTL

The service itself was rewritten in Go, handling requests concurrently with worker pools. Together with serving config from memory instead of a Firestore round trip, that’s where the 90% latency drop came from.

Freshness is the real requirement

Caching trades cost for correctness, so the key question is how stale config is allowed to be. We never had a serious incident from it, but the what-if drove the design. Suppose a team ships a broken experiment, say a variant that fails at checkout, and pauses it. If the pause doesn’t reach the serving path, users keep landing in the broken variant, and every minute of that is lost sales.

How long a paused experiment keeps serving Log time scale. With an event: seconds. If the event is lost: up to the TTL. With assignment on the device: until the app restarts, days or weeks. 1 s1 min 1 h1 day 1 week Event reaches the pod Event lost, TTL expires Assigned on the device Event reaches the pod: seconds seconds Event lost: stale until the TTL expires up to the TTL Client-side assignment: stale until the app restarts, days or weeks until the app restarts
How long a paused experiment keeps serving, on a log scale. The TTL is drawn as minutes for illustration.

With client-side assignment, “a while” meant until each user restarted the app. With events, it means seconds; if an event is lost, one TTL at most. So the TTL should come from the worst staleness you can accept for a kill switch, not from how many reads it saves.

What it changed

  • Cloud cost fell 95% year over year, measured on actual billing with some shared services allocated by estimate. Two things drove it: removing the polling, and right-sizing the workloads once they no longer had to absorb that load.
  • API latency fell 90%, from in-memory config plus the concurrent Go service.
  • An earlier phase of making the API asynchronous had already cut resource use by 50%; I count that separately.
  • The less measurable win was control: one assignment path, and a kill switch that reaches every user within seconds.

What I’d consider now

The design above works, but it has more moving parts than the problem needs: an event channel, Redis, a TTL and a fallback read path, all to keep a small, rarely changing dataset fresh. These are the alternatives I’d weigh today, starting with the smallest change.

1. Let Firestore push the changes

Firestore can already push changes itself. A snapshot listener subscribes to a query and receives an update whenever a matching document changes. Each assignment pod could listen to the config collection and keep the result in memory.

Firestore snapshot listeners Each assignment pod opens a snapshot listener on the Firestore config collection. Firestore pushes changed documents to every pod, which keeps config in memory. Firestore config collection sends only changeddocuments Pod 1 config held in memory Pod 2 config held in memory Pod 3 config held in memory snapshot listeners
Each pod subscribes to the config collection. Firestore sends the full set once on connect, then only the documents that change.

This replaces the change event, the TTL, and arguably Redis: the source of truth notifies the pods directly, so there’s no separate channel to fall out of sync. The costs are different rather than zero. Each listener pays for reading the full result set when it connects and for each changed document after that. A listener that stays disconnected too long is billed as a new query when it reconnects. And every pod holds an open connection to the database. At ten changes a day and tens of pods, that’s cheap. I’d still keep a slow periodic resync, in case a listener silently stops.

2. Poll cheaply with ETags

Polling itself isn’t the problem; paying for a full read on every poll is. HTTP already has the fix. Serve the whole config from one URL, from a storage bucket or a small service, with an ETag: a short fingerprint of the content, usually a hash of it. The client keeps the ETag it last received and sends it back on its next poll in an If-None-Match header. If the config hasn’t changed, the server replies 304 Not Modified with no body. If it has, it sends the new config with its new ETag.

Polling with ETags A publisher writes config.json on every change; the server labels it with an ETag, a hash of its contents. Pods and apps poll with If-None-Match and the last ETag they saw. Usually the answer is 304 Not Modified with no body; after a change it is 200 with the new config and a new ETag. Publisheron every change config.jsonETag = hash(contents) Pods and appsremember the last ETag GET config · If-None-Match: "9f2c" Usual case: nothing changed, empty reply 304 Not Modified · no body After a change: new body and new ETag after a change: 200 · new config · ETag "a71e" most polls:an empty 304
The server fingerprints the config; the client sends back the fingerprint it has. Unchanged config costs a tiny request and an empty reply, which a CDN can answer without touching the origin.

Most polls now cost almost nothing: a tiny request and an empty reply, which a CDN can answer without reaching the origin at all. Because the ETag comes from the content, there’s no version number to keep in step with the data. If the config is identical, the fingerprint is identical. It works for anything that speaks HTTP, including mobile apps. Freshness is still bounded by the poll interval, so for a kill switch the interval has to be short, or paired with a push. If you also want history and one-click rollback, keep a copy of each published config next to the live one; the ETag only handles change detection.

3. Push over long-lived connections

The most advanced option is a service that holds an open stream to every pod and SDK, over server-sent events (SSE) or gRPC streaming, and pushes a diff the moment config changes.

A dedicated config push service A config push service reads from the config store and keeps long-lived streams open to every pod and app SDK, sending diffs when config changes. Config store source of truth Config push service holds open streams sends diffs Pod Pod SDK in an app gRPC / server-sent events
The pattern feature-flag vendors and service meshes use. Freshest, but every arrow is a connection that has to survive load balancers, deploys and scaling.

I considered SSE at the time and held back, mostly out of worry about the network: how long-lived connections behave inside Kubernetes. I didn’t have much reference for it then. Looking back, the worry was reasonable. These are the parts that take real work:

  • Idle timeouts. Load balancers and ingress proxies close connections that go quiet, often after a minute or less. The server has to send heartbeats to keep streams open.
  • Buffering. Some proxies buffer responses, which holds SSE events back until buffering is turned off for that route.
  • No rebalancing. A long-lived connection stays on the pod it first reached. When the push service scales out, existing clients don’t move, so load stays uneven until connections are recycled.
  • Reconnect storms. Every deploy or pod restart drops its connections at once, and all those clients reconnect together. They need backoff with jitter, and a way to catch up on what they missed: SSE’s Last-Event-ID, or simply refetching the full config on reconnect.
  • Cost per connection. Each open stream holds memory and a file descriptor on the server.

None of these is a blocker. Feature-flag vendors and service meshes run exactly this pattern. But each is engineering you have to own, so it’s worth it when you have many clients and a strict freshness requirement, not before.

Side by side

Design Freshness Read cost Moving parts Main risk
What we built seconds; one TTL if an event is lost Firestore on cache misses only event channel, Redis, TTL event and cache paths drifting apart
Firestore listeners seconds initial load + each change, per listener Firestore only open connections; reconnect re-reads
ETag polling the poll interval near zero; mostly empty 304s a config endpoint, optional CDN staleness up to one interval
Push over SSE or gRPC under a second none per read a service holding many streams long-lived connections through Kubernetes networking

For the scale we had, around ten changes a day and a few dozen pods, I’d move the backend to Firestore listeners with a slow resync as a backstop. It deletes the most moving parts and keeps the freshness we needed. Apps would stay on server-side assignment, for the kill switch. If client count or read volume grew by an order of magnitude, I’d move to ETag polling behind a CDN, and add push over long-lived connections only if sub-second freshness became a hard requirement.

What I took from it

If a system reads something far more often than it changes, look at what you’re paying for the reads, in money and in latency. Polling is the easiest design to build and the hardest to notice, because each poll is tiny and the total only shows up on the bill.

The version I’d defend now:

  • make reads free (serve them from memory) and make changes loud (push them);
  • pick the staleness bound from the kill switch, not from cost;
  • keep assignment in one place you control;
  • and before adding infrastructure to push changes, check whether your data store can already push them.

References

  1. Google. Cloud Firestore pricing. Firebase documentation. https://firebase.google.com/docs/firestore/pricing — per-document read billing and listener reconnect charges
  2. Google. Get realtime updates with Cloud Firestore. Firebase documentation. https://firebase.google.com/docs/firestore/query-data/listen
  3. Fielding, R., Nottingham, M. and Reschke, J. HTTP Semantics. RFC 9110, IETF, 2022. https://www.rfc-editor.org/rfc/rfc9110.html — ETag, If-None-Match and 304 Not Modified
  4. Fielding, R., Nottingham, M. and Reschke, J. HTTP Caching. RFC 9111, IETF, 2022. https://www.rfc-editor.org/rfc/rfc9111.html
  5. WHATWG. HTML Living Standard: Server-sent events. https://html.spec.whatwg.org/multipage/server-sent-events.html — Last-Event-ID and reconnection
  6. gRPC authors. Core concepts, architecture and lifecycle. grpc.io. https://grpc.io/docs/what-is-grpc/core-concepts/ — server-streaming RPCs