From polling to push: getting experiment config to every service
The architecture we had, the one we moved to, and the designs I'd weigh today, for the unglamorous job of keeping experiment config fresh everywhere.
An experimentation platform looks like it’s about statistics. A lot of the engineering is about something plainer: making sure that when someone changes an experiment, every service that assigns users to variants finds out, quickly and cheaply.
I led the modernisation of a platform that ran hundreds of experiments a year for dozens of product teams. This post is about one part of that work, the path config takes from where people edit it to where it’s used. We went from polling to event-driven caching, cut the platform’s cloud bill by 95% and API latency by 90% over a year, and ended up with a much clearer design. I’ll finish with the alternatives I’d consider if I were building it again.
Two sides with very different traffic
It helps to split the platform in two. The control plane is where experiments change: people create them, change traffic splits, pause them, and bandits shift traffic toward the winning variant every few hours. Across all running experiments, that was around ten changes a day. The data plane is where config is used: every request that needs to know which variant a user sees. That was more than 20 backend services, each running many pods on Kubernetes, plus every user of the apps.
A good design for the arrow in the middle has to meet five requirements at once:
- Fresh: a paused experiment stops serving quickly. That’s the kill switch.
- Consistent: the same user gets the same variant from every service.
- Cheap: cost should grow with changes, not with traffic or time.
- Fast: assignment sits on the request path, so it can’t add much latency.
- Resilient: if the config store has a bad minute, assignment keeps working with the last known config.
Where we started
The first architecture had two clients, and each failed one requirement badly.
The backend polled. Every pod asked Firestore for the current config every two minutes. That’s 720 polls a day from every pod, to catch about ten changes spread across the whole company. Firestore bills per document read, so the cost was roughly:
pods × documents per poll × polls per day
The rate of change doesn’t appear in that formula. The bill grew with fleet size and the number of hours in a day. Adding a service, adding experiments, or scaling out for a traffic peak all made it bigger, while the answer stayed “no change” nearly every time.
The apps assigned on the device. The app fetched config at start-up and computed each user’s variant locally. That was cheap, but config only refreshed when the app restarted, and a backgrounded mobile app might not restart for days or weeks. When a team paused a broken experiment, users who hadn’t restarted kept seeing it. It also meant the assignment logic existed twice, in the backend and in the apps, and the two had to stay identical.
So the backend paid to stay fresh when it didn’t need to, and the apps stayed stale when it mattered most.
Where we ended up
The redesign moved assignment into one service and changed how that service learns about changes.
- One assignment service. Backend services and, eventually, the apps ask the service for variants instead of computing them. One implementation, one place to fix bugs, and a kill switch that works for everyone.
- A local cache in each pod. Config is read from memory on the request path. That alone removed most Firestore reads and most of the per-request latency.
- A change event on every config update. Saving an experiment, or a bandit re-split, publishes an event, and pods invalidate their cache when they receive it. Reads now scale with the number of changes, not the number of seconds.
- A TTL as a backstop. If an event is lost or delayed, the entry expires and gets reloaded anyway.
- Redis as a shared layer. Because the service ran as many pods, each local cache was its own copy. Two pods could disagree for a while, and every pod reloaded from Firestore separately. Redis between the pods and Firestore gave them one shared copy and turned many reloads into one.
Here’s what happens when someone pauses an experiment:
sequenceDiagram
participant UI as Experiment UI
participant FS as Firestore
participant EV as Change event
participant P as Assignment pods
participant R as Redis
UI->>FS: pause experiment 42
UI->>EV: publish "42 changed"
EV->>P: invalidate 42
Note over P: next request for 42
P->>R: get config 42
R-->>P: miss
P->>FS: read config 42
P->>R: store config 42 with TTL
The service itself was rewritten in Go, handling requests concurrently with worker pools. Together with serving config from memory instead of a Firestore round trip, that’s where the 90% latency drop came from.
Freshness is the real requirement
Caching trades cost for correctness, so the key question is how stale config is allowed to be. We never had a serious incident from it, but the what-if drove the design. Suppose a team ships a broken experiment, say a variant that fails at checkout, and pauses it. If the pause doesn’t reach the serving path, users keep landing in the broken variant, and every minute of that is lost sales.
With client-side assignment, “a while” meant until each user restarted the app. With events, it means seconds; if an event is lost, one TTL at most. So the TTL should come from the worst staleness you can accept for a kill switch, not from how many reads it saves.
What it changed
- Cloud cost fell 95% year over year, measured on actual billing with some shared services allocated by estimate. Two things drove it: removing the polling, and right-sizing the workloads once they no longer had to absorb that load.
- API latency fell 90%, from in-memory config plus the concurrent Go service.
- An earlier phase of making the API asynchronous had already cut resource use by 50%; I count that separately.
- The less measurable win was control: one assignment path, and a kill switch that reaches every user within seconds.
What I’d consider now
The design above works, but it has more moving parts than the problem needs: an event channel, Redis, a TTL and a fallback read path, all to keep a small, rarely changing dataset fresh. These are the alternatives I’d weigh today, starting with the smallest change.
1. Let Firestore push the changes
Firestore can already push changes itself. A snapshot listener subscribes to a query and receives an update whenever a matching document changes. Each assignment pod could listen to the config collection and keep the result in memory.
This replaces the change event, the TTL, and arguably Redis: the source of truth notifies the pods directly, so there’s no separate channel to fall out of sync. The costs are different rather than zero. Each listener pays for reading the full result set when it connects and for each changed document after that. A listener that stays disconnected too long is billed as a new query when it reconnects. And every pod holds an open connection to the database. At ten changes a day and tens of pods, that’s cheap. I’d still keep a slow periodic resync, in case a listener silently stops.
2. Poll cheaply with ETags
Polling itself isn’t the problem; paying for a full read on every poll is. HTTP already has the fix. Serve the whole config from one URL, from a storage bucket or a small service, with an ETag: a short fingerprint of the content, usually a hash of it. The client keeps the ETag it last received and sends it back on its next poll in an If-None-Match header. If the config hasn’t changed, the server replies 304 Not Modified with no body. If it has, it sends the new config with its new ETag.
Most polls now cost almost nothing: a tiny request and an empty reply, which a CDN can answer without reaching the origin at all. Because the ETag comes from the content, there’s no version number to keep in step with the data. If the config is identical, the fingerprint is identical. It works for anything that speaks HTTP, including mobile apps. Freshness is still bounded by the poll interval, so for a kill switch the interval has to be short, or paired with a push. If you also want history and one-click rollback, keep a copy of each published config next to the live one; the ETag only handles change detection.
3. Push over long-lived connections
The most advanced option is a service that holds an open stream to every pod and SDK, over server-sent events (SSE) or gRPC streaming, and pushes a diff the moment config changes.
I considered SSE at the time and held back, mostly out of worry about the network: how long-lived connections behave inside Kubernetes. I didn’t have much reference for it then. Looking back, the worry was reasonable. These are the parts that take real work:
- Idle timeouts. Load balancers and ingress proxies close connections that go quiet, often after a minute or less. The server has to send heartbeats to keep streams open.
- Buffering. Some proxies buffer responses, which holds SSE events back until buffering is turned off for that route.
- No rebalancing. A long-lived connection stays on the pod it first reached. When the push service scales out, existing clients don’t move, so load stays uneven until connections are recycled.
- Reconnect storms. Every deploy or pod restart drops its connections at once, and all those clients reconnect together. They need backoff with jitter, and a way to catch up on what they missed: SSE’s
Last-Event-ID, or simply refetching the full config on reconnect. - Cost per connection. Each open stream holds memory and a file descriptor on the server.
None of these is a blocker. Feature-flag vendors and service meshes run exactly this pattern. But each is engineering you have to own, so it’s worth it when you have many clients and a strict freshness requirement, not before.
Side by side
| Design | Freshness | Read cost | Moving parts | Main risk |
|---|---|---|---|---|
| What we built | seconds; one TTL if an event is lost | Firestore on cache misses only | event channel, Redis, TTL | event and cache paths drifting apart |
| Firestore listeners | seconds | initial load + each change, per listener | Firestore only | open connections; reconnect re-reads |
| ETag polling | the poll interval | near zero; mostly empty 304s | a config endpoint, optional CDN | staleness up to one interval |
| Push over SSE or gRPC | under a second | none per read | a service holding many streams | long-lived connections through Kubernetes networking |
For the scale we had, around ten changes a day and a few dozen pods, I’d move the backend to Firestore listeners with a slow resync as a backstop. It deletes the most moving parts and keeps the freshness we needed. Apps would stay on server-side assignment, for the kill switch. If client count or read volume grew by an order of magnitude, I’d move to ETag polling behind a CDN, and add push over long-lived connections only if sub-second freshness became a hard requirement.
What I took from it
If a system reads something far more often than it changes, look at what you’re paying for the reads, in money and in latency. Polling is the easiest design to build and the hardest to notice, because each poll is tiny and the total only shows up on the bill.
The version I’d defend now:
- make reads free (serve them from memory) and make changes loud (push them);
- pick the staleness bound from the kill switch, not from cost;
- keep assignment in one place you control;
- and before adding infrastructure to push changes, check whether your data store can already push them.
References
- Google. Cloud Firestore pricing. Firebase documentation. https://firebase.google.com/docs/firestore/pricing — per-document read billing and listener reconnect charges
- Google. Get realtime updates with Cloud Firestore. Firebase documentation. https://firebase.google.com/docs/firestore/query-data/listen
- Fielding, R., Nottingham, M. and Reschke, J. HTTP Semantics. RFC 9110, IETF, 2022. https://www.rfc-editor.org/rfc/rfc9110.html — ETag, If-None-Match and 304 Not Modified
- Fielding, R., Nottingham, M. and Reschke, J. HTTP Caching. RFC 9111, IETF, 2022. https://www.rfc-editor.org/rfc/rfc9111.html
- WHATWG. HTML Living Standard: Server-sent events. https://html.spec.whatwg.org/multipage/server-sent-events.html — Last-Event-ID and reconnection
- gRPC authors. Core concepts, architecture and lifecycle. grpc.io. https://grpc.io/docs/what-is-grpc/core-concepts/ — server-streaming RPCs