What firmware taught me about Go worker pools
Keeping P95 under 100 ms when traffic spikes tenfold, and why bounded concurrency felt familiar from writing interrupt handlers.
Before I worked on ML platforms, I wrote firmware. A smart-card reader on an STM32, sensor acquisition for a drone that streamed readings to a ground station over MAVLink, and small games on LPC and ATmega boards. Years later, building a high-throughput experiment assignment service in Go, I kept reaching for the same few patterns. The hardware couldn’t be more different, but the problem is the same one: work arrives faster than you can always handle it, and you have to decide in advance what happens then.
The firmware version
A microcontroller has kilobytes of RAM and no operating system to save you. Data arrives through interrupts: a byte lands on a UART, the CPU drops what it’s doing and runs a handler. The rule every embedded programmer learns early is that the handler does as little as possible, because while it runs, nothing else can:
#define BUF_SIZE 64
static volatile uint8_t buf[BUF_SIZE];
static volatile uint8_t head, tail;
void USART2_IRQHandler(void) {
uint8_t byte = USART2->DR; // read the byte, clear the interrupt
uint8_t next = (head + 1) % BUF_SIZE;
if (next != tail) { // buffer full? then drop, on purpose
buf[head] = byte;
head = next;
}
}
int main(void) {
for (;;) { // the main loop drains at its own pace
while (tail != head) {
handle(buf[tail]);
tail = (tail + 1) % BUF_SIZE;
}
feed_watchdog();
}
}
Three decisions are baked into those lines. The buffer has a fixed size, because there’s no heap to grow into. When it’s full, the handler drops the byte: a deliberate overload policy, not an accident. And a watchdog resets the chip if the main loop ever hangs.
The Go version
An assignment service answers one question per request: which variant should this user see? Ours ran on Kubernetes with Redis behind it, at around 5,000 requests per second on an average day. Go makes concurrency cheap enough that the tempting design is one goroutine for everything: every request, and every piece of work a request triggers, gets its own. That works right up to a traffic spike. Then goroutines pile up faster than they finish, each one holding memory and waiting on the same Redis connections, and latency climbs for everyone at once.
The fix was the firmware pattern, translated:
Channels, briefly
Go’s answer to the ring buffer is the channel: a typed pipe between goroutines, managed by the runtime so you never touch the indices or the locking yourself. Three behaviours cover everything in this post:
select with default is the "if full, drop" check.- An unbuffered channel is a hand-off. The sender blocks until a receiver takes the value, which is good for synchronising, bad for absorbing bursts.
- A buffered channel has a fixed number of slots. Sends succeed immediately until it’s full, then block. That’s the ring buffer from the firmware, with the waiting built in.
- A
selectwith adefaultbranch turns a blocking send into a try: if the channel is full, take the other branch now. That’s theif (next != tail)check.
Two more details make channels a good fit for worker pools. Many goroutines can receive from the same channel, and each value goes to exactly one of them, so workers share a queue without extra locking. And closing a channel ends every worker’s for job := range jobs loop once the queue drains, which gives you graceful shutdown for free.
For work that didn’t need to finish before the response, the handler did the minimum and handed the job to a worker pool: a buffered channel drained by a fixed number of goroutines.
type Pool struct { jobs chan Job }
func NewPool(workers, queue int) *Pool {
p := &Pool{jobs: make(chan Job, queue)} // fixed-size buffer
for i := 0; i < workers; i++ {
go func() {
for job := range p.jobs { // fixed number of workers
job.Run()
}
}()
}
return p
}
// Submit never blocks the request path: a full queue takes the overload path.
func (p *Pool) Submit(job Job) bool {
select {
case p.jobs <- job:
return true
default:
return false // count it, log it, fall back
}
}
Submit is the interrupt handler’s if (next != tail), one level up. The request never waits on a full queue. It takes the overload path immediately, and that path is something you designed, measured and can alert on.
Semaphores and deadlines
Not all work can be deferred. For calls that must finish inside the request, such as a Redis read, the tool is a counting semaphore. It’s exactly the primitive an RTOS gives you, and in Go it’s a buffered channel of empty structs. Pair it with a context deadline and you have the watchdog too:
var sem = make(chan struct{}, 64) // at most 64 concurrent calls
func withLimit(ctx context.Context, fn func(context.Context) error) error {
select {
case sem <- struct{}{}: // acquire a permit
defer func() { <-sem }() // release it
return fn(ctx)
case <-ctx.Done(): // no permit before the deadline
return ctx.Err() // fail fast instead of piling up
}
}
A request that can’t get a permit before its deadline fails fast, and the caller falls back to a default variant. For an experiment, that means the user sees the control experience, which is always a safe answer. A slow answer that eventually arrives is worse, because by then the page has already rendered.
Did it hold?
Bounded concurrency is easy to claim and easy to test. We ran a 30-minute distributed load test against the service:
A test like that is only meaningful if you state its shape: how long, what payload, compared with what baseline. Thirty minutes at ten times average traffic, with P95 still under 100 ms, is evidence that the limits were set well. It doesn’t prove the service can take any spike forever. It shows that when load exceeded capacity, the service degraded the way we’d designed it to.
What carries over
- Make the limit a number you chose. Firmware forces this because memory is finite. In Go you have to impose it yourself, because goroutines make unbounded concurrency feel free until it isn’t.
- Keep the fast path fast. The interrupt handler and the HTTP handler both do the minimum, then hand off.
- Design the overload behaviour. Dropping a byte, shedding a request, or serving the control variant are all fine answers if they’re deliberate, counted and visible. The failure mode to avoid is the undesigned one.
- Always have a watchdog. A deadline on every call is the difference between one slow dependency and a whole service stuck waiting.
Microcontrollers with kilobytes of memory and a Kubernetes deployment have almost nothing in common, except this: whatever you don’t bound, the next traffic spike will bound for you.
References
- MAVLink. MAVLink Developer Guide: Introduction. mavlink.io. https://mavlink.io/en/
- Go documentation. Effective Go: Concurrency. go.dev. https://go.dev/doc/effective_go#concurrency — channels, goroutines and semaphores via buffered channels
- Go documentation. The Go Programming Language Specification. go.dev. https://go.dev/ref/spec — channel types and select statements
- Ajmani, S. Go Concurrency Patterns: Pipelines and cancellation. The Go Blog, 2014. https://go.dev/blog/pipelines
- Ajmani, S. Go Concurrency Patterns: Context. The Go Blog, 2014. https://go.dev/blog/context — deadlines and cancellation
- Dean, J. and Barroso, L. A. The Tail at Scale. Communications of the ACM 56(2), 2013. https://research.google/pubs/the-tail-at-scale/ — why P95 and tail latency matter