SSE on Kubernetes: what a service mesh fixes, and what it doesn't
I wanted our experimentation service to push changes instead of being asked for them. What server-sent events are, what a service mesh like Istio does, and which problems it solves for long-lived connections.
I want the experimentation service to be event-driven: when someone changes an experiment, the service should tell everyone who needs to know, straight away. When nothing changes, it should stay quiet.
That isn’t how it works today. The services that assign users to variants pull their config: they read it through a cache that events and a TTL keep fresh, and apps ask on every request. Most of those reads get the same answer as last time.
When I wrote about getting experiment config to every service, push was the most advanced option, and I held back from it. Push means keeping a connection open to every client for hours, and I didn’t know how well that works inside Kubernetes. So I went and learned. This post is what I found, in the order I wish I’d learned it:
- What server-sent events (SSE) are.
- Why long-lived connections are tricky.
- What a service mesh is, and what Istio does.
- Each problem again, and whether a mesh fixes it.
1. Server-sent events
SSE is a normal HTTP request whose answer never finishes. The client asks once. The server replies, keeps the connection open, and writes a new message into it whenever something happens. Think of a phone call where only one side talks, and only when there’s news.
This is what the client receives over time:
id: 42
data: {"experiment":"checkout-cta","status":"paused"}
: ping
id: 43
data: {"experiment":"search-rank","traffic":0.5}
Three things in there matter later:
- Each message ends with a blank line. The client handles it as soon as it arrives.
- A line starting with a colon is a comment. The client ignores it. The server can send one every few seconds just to show the line is still alive: a heartbeat.
- Each message can have an
id. If the connection drops, the client reconnects and says “the last one I got was 42”. The server then sends whatever came after it. Browsers do this automatically, through a header calledLast-Event-ID.
SSE only goes one way, from server to client. That’s all config delivery needs. If you need both directions, WebSockets are the usual choice.
2. Why long connections are tricky
Most web infrastructure is built for short requests: a question comes in, an answer goes out, done in under a second. An SSE connection can stay open for hours. Between the client and the server there are several middlemen, such as load balancers and proxies, and each one was tuned with short requests in mind.
Four things can go wrong:
- A middleman hangs up on a quiet line. Most proxies close connections that have been silent for a while, often after a minute.
- A middleman holds messages back. Some proxies collect a response in a buffer before passing it on. For a normal request that’s fine. For SSE, the messages sit in the buffer and arrive late or not at all.
- New servers get no clients. A client picks a server when it connects and stays there. If you add a server because the others are busy, existing clients don’t move to it.
- Everyone calls back at once. When a server restarts during a deploy, all its clients lose their connection at the same moment and reconnect together.
There’s also a fifth, smaller one: every open connection takes some memory on every machine it passes through.
3. What a service mesh is
With many services, each one needs the same networking features: timeouts, retries, encrypted traffic, and metrics on every call. Without help, every team builds these into its own code, in its own way.
A service mesh takes those features out of the application and puts them into a small proxy next to every copy of every service. An analogy that helped me: each service gets a personal assistant that handles its calls. The service just says “call the push service”. The assistant finds a healthy copy, encrypts the call, retries if it fails, and writes down how long it took. A manager gives every assistant the same rulebook.
Istio is the most common service mesh on Kubernetes. In Istio’s terms:
- The assistants are Envoy proxies. Istio adds one to every pod automatically, as an extra container called a sidecar, and routes all the pod’s traffic through it.
- The manager is istiod. It tells every Envoy where the other services are, what the rules are, and which certificates to use for encryption.
- The front desk is the ingress gateway, a standalone Envoy where traffic from outside the cluster comes in.
You write the rules as Kubernetes objects, for example “send 10% of requests to version 2” or “time out after 3 seconds”.
There’s a nice irony here. istiod sends its rules to every Envoy over long-lived connections that stay open all the time. The mesh itself runs on the pattern I was nervous about.
4. The four problems, with a mesh
With Istio, an SSE connection from an app to our push service passes through three middlemen: the cloud load balancer, the ingress gateway, and the Envoy sidecar next to the push service. Each has its own timers.
A middleman hangs up on a quiet line
Does the mesh fix it? No, it adds more timers. But the fix is easy.
Every hop has an idle timer. For example, AWS’s load balancer closes connections that have been quiet for 60 seconds, and Envoy closes a stream after five quiet minutes. Envoy also has a 15-second limit on how long a whole response may take, which would cut every SSE connection. Istio switches that limit off by default, but it comes back if you set a timeout on the route yourself.
Fix: send a heartbeat comment every 15 to 30 seconds. That’s shorter than every idle timer on the path, so none of them fire.
A middleman holds messages back
Does the mesh fix it? Yes, mostly. Envoy passes messages through as they arrive.
Buffering usually comes from other places. NGINX-based ingress controllers buffer responses unless the server sends the header X-Accel-Buffering: no. Compression can also hold small messages back until it has enough to compress.
Fix: turn buffering off for the stream, and don’t compress text/event-stream responses.
New servers get no clients
Does the mesh fix it? No. This one surprised me.
Envoy is smarter than plain Kubernetes about spreading load: it balances every request, instead of every connection. But an SSE connection is one request that never ends. Once it lands on a server, it stays there, mesh or no mesh.
Fix: have the server close each connection after a while, say every 10 minutes, with a little randomness so they don’t all close together. The client reconnects, and that time it may land on the new server.
Everyone calls back at once
Does the mesh fix it? Partly.
During a deploy, Istio gives the old server a short grace period, five seconds by default, and then closes its connections. The mesh won’t reconnect for the client, so every client does it on its own. What the mesh does add is protection for the servers: it can limit connections and stop sending traffic to a server that is struggling.
Fix: clients wait a random few seconds before reconnecting, and send the last id they saw so they don’t miss anything.
sequenceDiagram
participant C as App
participant G as Gateway
participant P1 as Old push pod
participant P2 as New push pod
C->>G: open stream
G->>P1: forward
P1-->>C: message 41
Note over P1: deploy starts
P1-->>C: goodbye, stream closes
Note over C: wait a random 0–5 s
C->>G: open stream, last id was 41
G->>P2: forward
P2-->>C: messages 42 and 43, then live
One trap is specific to meshes. In older setups, the sidecar could shut down before the application, cutting connections before the server could say goodbye. Newer Kubernetes versions let sidecars start first and stop last, and Istio can use that.
And the memory cost
Does the mesh fix it? No, it makes it a bit worse. Each open connection is now also held by the gateway and the sidecar. It’s small per connection, but worth measuring before you have hundreds of thousands.
What the mesh adds
Some things come for free with a mesh: encrypted traffic between services without code changes, metrics on how many streams are open and how long they last, and the option to send a new version of the push service only a small share of new connections at first.
Summary
| Problem | Does a mesh fix it? | What you do |
|---|---|---|
| Middlemen hang up on quiet lines | No, adds more timers | heartbeat every 15–30 s |
| Middlemen hold messages back | Mostly | buffering off, no compression |
| New servers get no clients | No | close connections every ~10 min, with randomness |
| Everyone calls back at once | Partly | random wait, resume from last id |
| Memory per connection | No, a bit worse | measure it |
Would I build it now
For our config system, with around ten changes a day, not yet. Polling cheaply with ETags, or Firestore listeners, gives us enough freshness with far less to run. And adopting a whole service mesh for one push service would be a lot.
If we needed changes to reach every client within a second, I’d build SSE with the fixes above: heartbeats, connections that close every few minutes, message ids for resuming, random waits on reconnect, buffering off, and polling as a fallback. I’d only use a mesh if the cluster already had one. None of the fixes depend on it.
My worry was right about which problems exist. I overestimated how hard they are: each one has a known fix, and most of the fixes are a few lines of server and client code.
References
- WHATWG. HTML Standard: Server-sent events. https://html.spec.whatwg.org/multipage/server-sent-events.html — the SSE format,
Last-Event-IDand comment heartbeats - MDN Web Docs. Using server-sent events. Mozilla. https://developer.mozilla.org/en-US/docs/Web/API/Server-sent_events/Using_server-sent_events
- Istio. Architecture. https://istio.io/latest/docs/ops/deployment/architecture/ — istiod, Envoy sidecars and the data plane
- Amazon Web Services. Application Load Balancers. Elastic Load Balancing documentation. https://docs.aws.amazon.com/elasticloadbalancing/latest/application/application-load-balancers.html — idle timeout, 60 seconds by default
- Envoy. How do I configure timeouts? Envoy proxy documentation. https://www.envoyproxy.io/docs/envoy/latest/faq/configuration/timeouts — route and stream idle timeouts
- NGINX. Module ngx_http_proxy_module. https://nginx.org/en/docs/http/ngx_http_proxy_module.html —
proxy_bufferingandX-Accel-Buffering - Kubernetes. Sidecar Containers. https://kubernetes.io/docs/concepts/workloads/pods/sidecar-containers/ — sidecar start and stop ordering