← All posts

SSE on Kubernetes: what a service mesh fixes, and what it doesn't

I wanted our experimentation service to push changes instead of being asked for them. What server-sent events are, what a service mesh like Istio does, and which problems it solves for long-lived connections.

I want the experimentation service to be event-driven: when someone changes an experiment, the service should tell everyone who needs to know, straight away. When nothing changes, it should stay quiet.

That isn’t how it works today. The services that assign users to variants pull their config: they read it through a cache that events and a TTL keep fresh, and apps ask on every request. Most of those reads get the same answer as last time.

Polling compared with push Over ten minutes, a client polling every two minutes checks six times. Five checks return no change. The config changes at minute 6.5, and the poll at minute 8 finds it, so the client is stale for 1.5 minutes. With push, the server sends one message at minute 6.5, the moment the change happens. config changes polling every 2 min5 of 6 checks: no change Polls at 0, 2, 4 and 6 min: no change Stale from 6.5 to 8 min stale 1.5 min Poll at 8 min finds the change Poll at 10 min: no change push (SSE)one message, when needed Push: sent at 6.5 min, the moment it changes sent the moment it changes 010 min
Polling is checking the mailbox every two minutes. Push is a doorbell: nothing happens until there's something to deliver.

When I wrote about getting experiment config to every service, push was the most advanced option, and I held back from it. Push means keeping a connection open to every client for hours, and I didn’t know how well that works inside Kubernetes. So I went and learned. This post is what I found, in the order I wish I’d learned it:

  1. What server-sent events (SSE) are.
  2. Why long-lived connections are tricky.
  3. What a service mesh is, and what Istio does.
  4. Each problem again, and whether a mesh fixes it.

1. Server-sent events

SSE is a normal HTTP request whose answer never finishes. The client asks once. The server replies, keeps the connection open, and writes a new message into it whenever something happens. Think of a phone call where only one side talks, and only when there’s news.

This is what the client receives over time:

id: 42
data: {"experiment":"checkout-cta","status":"paused"}

: ping

id: 43
data: {"experiment":"search-rank","traffic":0.5}

Three things in there matter later:

  • Each message ends with a blank line. The client handles it as soon as it arrives.
  • A line starting with a colon is a comment. The client ignores it. The server can send one every few seconds just to show the line is still alive: a heartbeat.
  • Each message can have an id. If the connection drops, the client reconnects and says “the last one I got was 42”. The server then sends whatever came after it. Browsers do this automatically, through a header called Last-Event-ID.

SSE only goes one way, from server to client. That’s all config delivery needs. If you need both directions, WebSockets are the usual choice.

2. Why long connections are tricky

Most web infrastructure is built for short requests: a question comes in, an answer goes out, done in under a second. An SSE connection can stay open for hours. Between the client and the server there are several middlemen, such as load balancers and proxies, and each one was tuned with short requests in mind.

Four things can go wrong:

  1. A middleman hangs up on a quiet line. Most proxies close connections that have been silent for a while, often after a minute.
  2. A middleman holds messages back. Some proxies collect a response in a buffer before passing it on. For a normal request that’s fine. For SSE, the messages sit in the buffer and arrive late or not at all.
  3. New servers get no clients. A client picks a server when it connects and stays there. If you add a server because the others are busy, existing clients don’t move to it.
  4. Everyone calls back at once. When a server restarts during a deploy, all its clients lose their connection at the same moment and reconnect together.

There’s also a fifth, smaller one: every open connection takes some memory on every machine it passes through.

3. What a service mesh is

With many services, each one needs the same networking features: timeouts, retries, encrypted traffic, and metrics on every call. Without help, every team builds these into its own code, in its own way.

A service mesh takes those features out of the application and puts them into a small proxy next to every copy of every service. An analogy that helped me: each service gets a personal assistant that handles its calls. The service just says “call the push service”. The assistant finds a healthy copy, encrypts the call, retries if it fails, and writes down how long it took. A manager gives every assistant the same rulebook.

Istio is the most common service mesh on Kubernetes. In Istio’s terms:

  • The assistants are Envoy proxies. Istio adds one to every pod automatically, as an extra container called a sidecar, and routes all the pod’s traffic through it.
  • The manager is istiod. It tells every Envoy where the other services are, what the rules are, and which certificates to use for encryption.
  • The front desk is the ingress gateway, a standalone Envoy where traffic from outside the cluster comes in.

You write the rules as Kubernetes objects, for example “send 10% of requests to version 2” or “time out after 3 seconds”.

Istio's architecture istiod, the control plane, pushes configuration and certificates to every Envoy proxy over long-lived gRPC streams. Traffic enters through an ingress gateway, which is a standalone Envoy, and reaches the frontend pod's Envoy sidecar. Calls from the frontend to the push service go from the frontend's Envoy to the push service's Envoy over mutual TLS. istiodconfig · certificates pushes config to every proxyover long-lived gRPC (xDS) ingress gatewayEnvoy at the edge pod: frontend Envoy app pod: push service Envoy app Requests from outside enter through the gateway Service-to-service calls go proxy to proxy, over mutual TLS mTLS between proxies
Dashed lines: istiod handing out the rulebook. Blue lines: real traffic, which always goes from one Envoy to another, never straight between apps.

There’s a nice irony here. istiod sends its rules to every Envoy over long-lived connections that stay open all the time. The mesh itself runs on the pattern I was nervous about.

4. The four problems, with a mesh

With Istio, an SSE connection from an app to our push service passes through three middlemen: the cloud load balancer, the ingress gateway, and the Envoy sidecar next to the push service. Each has its own timers.

Every hop on an SSE stream has a timer An SSE stream goes from the app or SDK through a cloud load balancer, the Istio ingress gateway, the push service's Envoy sidecar, and into the push service. The load balancer has an idle timeout, 60 seconds by default on AWS's Application Load Balancer. Envoy's route timeout is 15 seconds by default, but Istio turns it off. Envoy's stream idle timeout is 5 minutes. On deploy, Istio drains for 5 seconds by default. Without heartbeats, a quiet stream is cut after 60 seconds; with a comment line every 20 seconds, no idle timer fires. app / SDKreconnects cloud LBload balancer gatewayEnvoy sidecarEnvoy pushholds streams backoffand jitter idle timeout60 s on AWS ALB route timeout15 s in Envoy,off in Istio stream idle5 min in Envoy drain on deploy5 s in Istio ONE QUIET STREAM, FIRST 5 MINUTES no heartbeat No heartbeat: the load balancer cuts the stream after 60 s idle cut after 60 s idle comment every 20 s With a comment line every 20 s, no idle timer fires 05 min
The shortest idle timer on the path decides how long a quiet stream lives. A heartbeat shorter than all of them keeps every one from firing. The defaults shown are examples; check your own.

A middleman hangs up on a quiet line

Does the mesh fix it? No, it adds more timers. But the fix is easy.

Every hop has an idle timer. For example, AWS’s load balancer closes connections that have been quiet for 60 seconds, and Envoy closes a stream after five quiet minutes. Envoy also has a 15-second limit on how long a whole response may take, which would cut every SSE connection. Istio switches that limit off by default, but it comes back if you set a timeout on the route yourself.

Fix: send a heartbeat comment every 15 to 30 seconds. That’s shorter than every idle timer on the path, so none of them fire.

A middleman holds messages back

Does the mesh fix it? Yes, mostly. Envoy passes messages through as they arrive.

Buffering usually comes from other places. NGINX-based ingress controllers buffer responses unless the server sends the header X-Accel-Buffering: no. Compression can also hold small messages back until it has enough to compress.

Fix: turn buffering off for the stream, and don’t compress text/event-stream responses.

New servers get no clients

Does the mesh fix it? No. This one surprised me.

Envoy is smarter than plain Kubernetes about spreading load: it balances every request, instead of every connection. But an SSE connection is one request that never ends. Once it lands on a server, it stays there, mesh or no mesh.

Fix: have the server close each connection after a while, say every 10 minutes, with a little randomness so they don’t all close together. The client reconnects, and that time it may land on the new server.

Share of streams on a newly added pod Simulation of 4,000 streams on three push pods when a fourth pod starts. With no cap on stream lifetime, and streams ending naturally about every three hours, the new pod holds 1% of streams after 10 minutes and 6% after an hour. With streams capped at 10 plus or minus 2 minutes, it holds 21% after 10 minutes and about 25%, an even share, from 15 minutes on. SHARE OF STREAMS ON THE NEW POD 0% 10% 20% 30% 0 10 20 30 40 50 60 minutes after the fourth pod starts even: 25% Never closed: 1% on the new pod after 10 min, 6% after 60 min Closed every 10 ± 2 min: 21% after 10 min, 26% after 20 min never closed closed every ~10 min
Simulated, illustrative numbers for a fourth server added to three busy ones. If connections never close, the new server stays almost empty for hours. If each connection closes after about 10 minutes, the load evens out within one cycle.

Everyone calls back at once

Does the mesh fix it? Partly.

During a deploy, Istio gives the old server a short grace period, five seconds by default, and then closes its connections. The mesh won’t reconnect for the client, so every client does it on its own. What the mesh does add is protection for the servers: it can limit connections and stop sending traffic to a server that is struggling.

Fix: clients wait a random few seconds before reconnecting, and send the last id they saw so they don’t miss anything.

sequenceDiagram
  participant C as App
  participant G as Gateway
  participant P1 as Old push pod
  participant P2 as New push pod
  C->>G: open stream
  G->>P1: forward
  P1-->>C: message 41
  Note over P1: deploy starts
  P1-->>C: goodbye, stream closes
  Note over C: wait a random 0–5 s
  C->>G: open stream, last id was 41
  G->>P2: forward
  P2-->>C: messages 42 and 43, then live

One trap is specific to meshes. In older setups, the sidecar could shut down before the application, cutting connections before the server could say goodbye. Newer Kubernetes versions let sidecars start first and stop last, and Istio can use that.

And the memory cost

Does the mesh fix it? No, it makes it a bit worse. Each open connection is now also held by the gateway and the sidecar. It’s small per connection, but worth measuring before you have hundreds of thousands.

What the mesh adds

Some things come for free with a mesh: encrypted traffic between services without code changes, metrics on how many streams are open and how long they last, and the option to send a new version of the push service only a small share of new connections at first.

Summary

Problem Does a mesh fix it? What you do
Middlemen hang up on quiet lines No, adds more timers heartbeat every 15–30 s
Middlemen hold messages back Mostly buffering off, no compression
New servers get no clients No close connections every ~10 min, with randomness
Everyone calls back at once Partly random wait, resume from last id
Memory per connection No, a bit worse measure it

Would I build it now

For our config system, with around ten changes a day, not yet. Polling cheaply with ETags, or Firestore listeners, gives us enough freshness with far less to run. And adopting a whole service mesh for one push service would be a lot.

If we needed changes to reach every client within a second, I’d build SSE with the fixes above: heartbeats, connections that close every few minutes, message ids for resuming, random waits on reconnect, buffering off, and polling as a fallback. I’d only use a mesh if the cluster already had one. None of the fixes depend on it.

My worry was right about which problems exist. I overestimated how hard they are: each one has a known fix, and most of the fixes are a few lines of server and client code.

References

  1. WHATWG. HTML Standard: Server-sent events. https://html.spec.whatwg.org/multipage/server-sent-events.html — the SSE format, Last-Event-ID and comment heartbeats
  2. MDN Web Docs. Using server-sent events. Mozilla. https://developer.mozilla.org/en-US/docs/Web/API/Server-sent_events/Using_server-sent_events
  3. Istio. Architecture. https://istio.io/latest/docs/ops/deployment/architecture/ — istiod, Envoy sidecars and the data plane
  4. Amazon Web Services. Application Load Balancers. Elastic Load Balancing documentation. https://docs.aws.amazon.com/elasticloadbalancing/latest/application/application-load-balancers.html — idle timeout, 60 seconds by default
  5. Envoy. How do I configure timeouts? Envoy proxy documentation. https://www.envoyproxy.io/docs/envoy/latest/faq/configuration/timeouts — route and stream idle timeouts
  6. NGINX. Module ngx_http_proxy_module. https://nginx.org/en/docs/http/ngx_http_proxy_module.html — proxy_buffering and X-Accel-Buffering
  7. Kubernetes. Sidecar Containers. https://kubernetes.io/docs/concepts/workloads/pods/sidecar-containers/ — sidecar start and stop ordering