← All posts

What firmware taught me about Go worker pools

Keeping P95 under 100 ms when traffic spikes tenfold, and why bounded concurrency felt familiar from writing interrupt handlers.

Before I worked on ML platforms, I wrote firmware. A smart-card reader on an STM32, sensor acquisition for a drone that streamed readings to a ground station over MAVLink, and small games on LPC and ATmega boards. Years later, building a high-throughput experiment assignment service in Go, I kept reaching for the same few patterns. The hardware couldn’t be more different, but the problem is the same one: work arrives faster than you can always handle it, and you have to decide in advance what happens then.

The firmware version

A microcontroller has kilobytes of RAM and no operating system to save you. Data arrives through interrupts: a byte lands on a UART, the CPU drops what it’s doing and runs a handler. The rule every embedded programmer learns early is that the handler does as little as possible, because while it runs, nothing else can:

#define BUF_SIZE 64
static volatile uint8_t buf[BUF_SIZE];
static volatile uint8_t head, tail;

void USART2_IRQHandler(void) {
    uint8_t byte = USART2->DR;           // read the byte, clear the interrupt
    uint8_t next = (head + 1) % BUF_SIZE;
    if (next != tail) {                  // buffer full? then drop, on purpose
        buf[head] = byte;
        head = next;
    }
}

int main(void) {
    for (;;) {                           // the main loop drains at its own pace
        while (tail != head) {
            handle(buf[tail]);
            tail = (tail + 1) % BUF_SIZE;
        }
        feed_watchdog();
    }
}

Three decisions are baked into those lines. The buffer has a fixed size, because there’s no heap to grow into. When it’s full, the handler drops the byte: a deliberate overload policy, not an accident. And a watchdog resets the chip if the main loop ever hangs.

The Go version

An assignment service answers one question per request: which variant should this user see? Ours ran on Kubernetes with Redis behind it, at around 5,000 requests per second on an average day. Go makes concurrency cheap enough that the tempting design is one goroutine for everything: every request, and every piece of work a request triggers, gets its own. That works right up to a traffic spike. Then goroutines pile up faster than they finish, each one holding memory and waiting on the same Redis connections, and latency climbs for everyone at once.

The fix was the firmware pattern, translated:

Firmware patterns and their Go equivalents Four pairs: an interrupt handler corresponds to an HTTP handler, a ring buffer to a buffered channel, main-loop tasks to worker goroutines, and a watchdog timer to a context deadline. FIRMWARE ON A MICROCONTROLLER GO ASSIGNMENT SERVICE Interrupt handlergrab the byte, push it, return HTTP handlervalidate, enqueue, return ≈ Ring bufferfixed size; full means drop Buffered channelfixed size; full means shed or wait ≈ Main-loop tasksa fixed number, own pace Worker goroutinesa fixed number, own pace ≈ Watchdog timerreset if a task hangs context deadlinegive up if work takes too long ≈
The same four ideas on very different hardware. In both, the fast path does as little as possible, work waits in a fixed-size buffer, and a fixed number of workers drain it.

Channels, briefly

Go’s answer to the ring buffer is the channel: a typed pipe between goroutines, managed by the runtime so you never touch the indices or the locking yourself. Three behaviours cover everything in this post:

Three ways to use a Go channel An unbuffered channel hands a value directly from sender to receiver, and both wait for each other. A buffered channel with four slots lets the sender continue until all slots are full. A select with a default branch tries to send and takes the default branch immediately if the buffer is full. unbufferedmake(chan Job) sender receiver hand-off: both sides wait bufferedmake(chan Job, 4) sender receiver sender waits only when all 4 slots are full select + defaulttry, don't wait sender receiver full → default branch
The three channel behaviours this post relies on. A buffered channel is a ring buffer the runtime manages for you; select with default is the "if full, drop" check.
  • An unbuffered channel is a hand-off. The sender blocks until a receiver takes the value, which is good for synchronising, bad for absorbing bursts.
  • A buffered channel has a fixed number of slots. Sends succeed immediately until it’s full, then block. That’s the ring buffer from the firmware, with the waiting built in.
  • A select with a default branch turns a blocking send into a try: if the channel is full, take the other branch now. That’s the if (next != tail) check.

Two more details make channels a good fit for worker pools. Many goroutines can receive from the same channel, and each value goes to exactly one of them, so workers share a queue without extra locking. And closing a channel ends every worker’s for job := range jobs loop once the queue drains, which gives you graceful shutdown for free.

For work that didn’t need to finish before the response, the handler did the minimum and handed the job to a worker pool: a buffered channel drained by a fixed number of goroutines.

type Pool struct { jobs chan Job }

func NewPool(workers, queue int) *Pool {
    p := &Pool{jobs: make(chan Job, queue)} // fixed-size buffer
    for i := 0; i < workers; i++ {
        go func() {
            for job := range p.jobs { // fixed number of workers
                job.Run()
            }
        }()
    }
    return p
}

// Submit never blocks the request path: a full queue takes the overload path.
func (p *Pool) Submit(job Job) bool {
    select {
    case p.jobs <- job:
        return true
    default:
        return false // count it, log it, fall back
    }
}

Submit is the interrupt handler’s if (next != tail), one level up. The request never waits on a full queue. It takes the overload path immediately, and that path is something you designed, measured and can alert on.

A bounded worker pool Requests reach a cheap handler, which puts work into a fixed-capacity queue drained by four workers. When the queue is full, the handler sheds the request quickly or waits with a deadline. Requests10× in a spike Handlercheap, bounded QUEUE · CAPACITY N worker 1 worker 2 worker 3 worker 4 queue fullshed fast or wait with a deadline N = concurrencylimit, on purpose
A worker pool turns "how much can we do at once?" into a number you choose. When a spike exceeds it, the overload policy is explicit, instead of an ever-growing number of goroutines competing for CPU, memory and connections.

Semaphores and deadlines

Not all work can be deferred. For calls that must finish inside the request, such as a Redis read, the tool is a counting semaphore. It’s exactly the primitive an RTOS gives you, and in Go it’s a buffered channel of empty structs. Pair it with a context deadline and you have the watchdog too:

var sem = make(chan struct{}, 64) // at most 64 concurrent calls

func withLimit(ctx context.Context, fn func(context.Context) error) error {
    select {
    case sem <- struct{}{}: // acquire a permit
        defer func() { <-sem }() // release it
        return fn(ctx)
    case <-ctx.Done(): // no permit before the deadline
        return ctx.Err() // fail fast instead of piling up
    }
}

A request that can’t get a permit before its deadline fails fast, and the caller falls back to a default variant. For an experiment, that means the user sees the control experience, which is always a safe answer. A slow answer that eventually arrives is worse, because by then the page has already rendered.

Did it hold?

Bounded concurrency is easy to claim and easy to test. We ran a 30-minute distributed load test against the service:

Load test numbers About 5,000 requests per second on average in production; more than 50,000 requests per second sustained for 30 minutes in a distributed load test; P95 latency under 100 ms in both. ~5K requests per second,average in production 50K+ requests per second,30-minute load test <100 ms P95 latency, in bothproduction and the test
The numbers, stated carefully: the 50K figure is a 30-minute distributed load test with a standard payload, about ten times average production traffic, not a claim about everyday load.

A test like that is only meaningful if you state its shape: how long, what payload, compared with what baseline. Thirty minutes at ten times average traffic, with P95 still under 100 ms, is evidence that the limits were set well. It doesn’t prove the service can take any spike forever. It shows that when load exceeded capacity, the service degraded the way we’d designed it to.

What carries over

  • Make the limit a number you chose. Firmware forces this because memory is finite. In Go you have to impose it yourself, because goroutines make unbounded concurrency feel free until it isn’t.
  • Keep the fast path fast. The interrupt handler and the HTTP handler both do the minimum, then hand off.
  • Design the overload behaviour. Dropping a byte, shedding a request, or serving the control variant are all fine answers if they’re deliberate, counted and visible. The failure mode to avoid is the undesigned one.
  • Always have a watchdog. A deadline on every call is the difference between one slow dependency and a whole service stuck waiting.

Microcontrollers with kilobytes of memory and a Kubernetes deployment have almost nothing in common, except this: whatever you don’t bound, the next traffic spike will bound for you.

References

  1. MAVLink. MAVLink Developer Guide: Introduction. mavlink.io. https://mavlink.io/en/
  2. Go documentation. Effective Go: Concurrency. go.dev. https://go.dev/doc/effective_go#concurrency — channels, goroutines and semaphores via buffered channels
  3. Go documentation. The Go Programming Language Specification. go.dev. https://go.dev/ref/spec — channel types and select statements
  4. Ajmani, S. Go Concurrency Patterns: Pipelines and cancellation. The Go Blog, 2014. https://go.dev/blog/pipelines
  5. Ajmani, S. Go Concurrency Patterns: Context. The Go Blog, 2014. https://go.dev/blog/context — deadlines and cancellation
  6. Dean, J. and Barroso, L. A. The Tail at Scale. Communications of the ACM 56(2), 2013. https://research.google/pubs/the-tail-at-scale/ — why P95 and tail latency matter