From a click to a database: how a request crosses the internet
Every step from typing an address to a row in a database, each piece explained with an everyday analogy, then how the same pieces shape regional and global deployments, and the infra and ops practices that follow.
After making sense of Kubernetes, I understood what happens to a request once it reaches our cluster. Everything before that was a blur. Someone types an address, and a fraction of a second later a page appears, built from data in a database in another country. I also kept hearing infra words I couldn’t place: region, zone, VPC, CDN, anycast, active–active, failover.
This post follows one request from start to finish and explains each piece it meets, with an everyday analogy for each. Then it zooms out to the decisions infra and ops teams make with those pieces: where to run things, how to survive failures, and how to operate it all. It ends with the practices I’d now call best practice.
The example throughout: someone in Jakarta opens shop.example.com, and the shop runs in a cloud region in Singapore.
The cast, in one table
Before the details, here is every component in this post with the analogy I’ll use for it. The analogies are about a city and its postal system.
| Component | Analogy | What it does |
|---|---|---|
| Packet | a letter | a small chunk of data with a sender and destination address |
| IP address | a street address | where a machine can be reached |
| Router | a post office | passes each packet one step closer to its destination |
| DNS | the phone book | turns a name like shop.example.com into an IP address |
| TCP | registered mail | numbers packets, confirms delivery, resends lost ones |
| TLS | a sealed envelope and an ID check | encrypts traffic and proves the server is who it says |
| CDN edge | a local convenience store | keeps copies of popular files close to users |
| Anycast | one hotline, answered by the nearest branch | one IP address served from many places at once |
| WAF and DDoS protection | the bouncer | turns away bad traffic before it gets in |
| Load balancer | the restaurant host | sends each guest to a free, working table |
| Region | a city | a cluster of data centres in one metro area |
| Zone | a building in that city | a data centre with its own power and cooling |
| VPC and subnets | a fenced campus with a public lobby and private offices | your private network inside the cloud |
| Firewall rules | the security guard at each door | decide which traffic may reach which machine |
| NAT gateway | the mailroom | lets private machines send mail out without having a public address |
| Kubernetes | the building manager | keeps the right number of workers at their desks |
| Pod | a worker | runs your code |
| Cache | a notepad on the desk | quick answers to repeated questions |
| Database | the filing room | the one true record of everything |
| Replica | a photocopy of the files | a copy for reading or for emergencies |
Part 1: getting there
The internet itself
The internet is thousands of separate networks that have agreed to carry each other’s traffic. Your mobile provider is one network, a cloud provider is another, and they connect at exchange points and through undersea cables.
Everything travels as packets, which are like letters. Each has a sender and a destination IP address, such as 203.0.113.10, the way a letter has a street address. Routers are the post offices: each reads the destination and passes the packet one step closer. Networks tell each other which addresses they can deliver to using a protocol called BGP, which is like post offices sharing their delivery maps.
One fact shaped everything else for me: distance costs time, and nothing makes it free. Light in fibre travels about 200,000 km per second. A packet from Jakarta to Singapore and back covers about 1,800 km, which takes at least 9 ms. To a US data centre and back it’s over 30,000 km, at least 157 ms. Cables don’t run in straight lines, so real numbers are higher.
That there-and-back time is the round-trip time. Most of the rest of this post is about how many round trips a request needs and how far each one travels.
Step 1: DNS, the phone book
Browsers need an IP address, not a name. DNS, the Domain Name System, is the phone book that turns one into the other.
The browser asks a resolver, usually run by your internet provider or a public service like Google’s 8.8.8.8. If the resolver hasn’t looked the name up recently, it walks down a chain, like asking directory enquiries, who refers you to the right regional phone book, which gives you the number:
sequenceDiagram
participant B as Browser
participant R as Resolver
participant Root as Root server
participant TLD as .com server
participant A as example.com's DNS
B->>R: address of shop.example.com?
R->>Root: shop.example.com?
Root-->>R: ask the .com servers
R->>TLD: shop.example.com?
TLD-->>R: ask example.com's DNS
R->>A: shop.example.com?
A-->>R: 203.0.113.10, keep it for 5 minutes
R-->>B: 203.0.113.10
Each answer comes with a TTL (time to live), which says how long it may be remembered. The resolver, the operating system and the browser all keep recent answers, so most lookups take no time at all.
Why ops cares:
- TTL is a trade-off. A short TTL, like 60 seconds, means that when you point the name somewhere new, say during a failover, the world follows within a minute. A long TTL means fewer lookups but slow changes.
- DNS can steer traffic. GeoDNS gives different answers depending on where the question comes from, so users in Jakarta and Frankfurt can be sent to different regions under the same name.
- DNS is a single point of failure if one provider hosts all your records. Big outages have come from exactly this.
Step 2: TCP and TLS, registered mail in a sealed envelope
With an address, the browser opens a connection. Two handshakes happen before any page data moves.
- TCP is registered mail. Packets can be lost or arrive out of order; TCP numbers them, confirms each delivery, resends what’s missing and puts them back in order. Setting it up takes one round trip.
- TLS is a sealed envelope plus an ID check. The server shows a certificate proving it owns the name, and both sides agree on keys for encryption. It’s the “S” in HTTPS. TLS 1.3 takes one more round trip; older versions took two.
Only then does the browser send the actual request, and the answer takes a third round trip. So a first response on a new connection costs at least three round trips, plus the server’s own work.
Why ops cares:
- Reuse connections. Browsers keep them open, HTTP/2 sends many requests over one, and HTTP/3 (running on a newer transport called QUIC) combines the handshakes. Every new connection pays the handshakes again.
- Certificates expire. An expired certificate takes a site down as surely as a crashed server. Automate renewal, for example with Let’s Encrypt or your cloud’s managed certificates, and alert well before expiry.
- Where TLS ends matters. Usually the CDN or load balancer decrypts traffic, which is called TLS termination, so it can route by URL. Traffic behind it is re-encrypted or kept inside a private network.
Step 3: the edge, a convenience store near home
Much of a web page is the same for everyone: images, scripts, stylesheets, fonts. You don’t drive to the central warehouse for a bottle of water; you buy it at the shop on your street. A CDN (content delivery network) is that shop: a provider with servers in hundreds of cities, each keeping copies of your popular files.
For our Jakarta user, the edge is a couple of milliseconds away. The handshakes happen with it, cheaply. On a cache hit it answers directly. On a miss, it fetches the file from the region once, over a connection it already keeps open, and keeps the copy for the next person.
Many CDNs and global load balancers use anycast: the same IP address is announced from many locations at once, and internet routing delivers each user to the nearest one. It’s like one national hotline number that’s answered by whichever branch is closest to the caller.
The edge is also where the bouncer stands. A WAF (web application firewall) blocks known attack patterns, and DDoS protection absorbs floods of fake traffic across the whole edge network, before any of it reaches your servers.
Why ops cares:
- Cache rules are yours to set. Response headers such as
Cache-Controltell the edge how long to keep a file. - Version file names, like
app.3f9c.js. A new release then gets a new name, so nobody sees a stale copy, and you rarely need to clear the cache by hand. - Put all public traffic behind the edge, so the bouncer sees it first.
Part 2: inside the region
Step 4: regions, zones and the load balancer
Requests that need fresh, personal data, like “show my basket”, go on to the region.
A region is a city: a cloud provider’s group of data centres in one metro area, such as Singapore. A zone is one building in that city, with its own power, cooling and network, a few kilometres from the others and linked to them by fast private lines. A round trip between zones takes about a millisecond. If one building loses power, the others keep running.
The request arrives at a load balancer, the restaurant host. It owns the public address, ends the TLS connection, checks which tables (backends) are free and working, and seats the guest at one of them. It learns which backends work through health checks: it asks each one “are you OK?” every few seconds and stops sending traffic to those that don’t answer.
Load balancers come in two kinds that are worth knowing apart:
- Layer 4 balancers look only at addresses and ports. They are fast, and they pick a backend once per connection.
- Layer 7 balancers understand HTTP. They can route by path or hostname, retry a failed request, and balance every request separately.
Cloud providers also offer global load balancers: one anycast address worldwide that sends each user to the nearest healthy region. Regional ones serve a single region.
Step 5: the private network, a fenced campus
Inside the region, your machines live in a VPC (virtual private cloud): a private network that only you use. Think of a fenced campus. It’s split into subnets, which are areas of the campus:
- a public subnet is the lobby: only the load balancer lives there, with an address the internet can reach;
- private subnets are the offices: app servers and databases, with no public address at all.
Firewall rules (called security groups on some clouds) are the guards at each door: “the load balancer may talk to the app on port 8080”, “only the app may talk to the database on port 5432”. Everything else is refused.
Private machines still sometimes need to reach out, for example to download a package or call a partner’s API. A NAT gateway is the mailroom: they hand their outgoing mail to it, it sends it with its own return address, and it passes the replies back. Nobody outside can send mail straight to an office.
Step 6: Kubernetes, the building manager
In our setup, the load balancer hands requests to a Kubernetes cluster whose machines sit in the private subnets, spread across the zones. An ingress reads the hostname and path, a Service picks one of the matching pods (the workers), and the pod runs our code.
The manager’s most useful habit for ops is autoscaling, which is calling in extra staff for the lunch rush:
- the horizontal pod autoscaler adds pods when CPU or request rate goes up, and removes them when it drops;
- the cluster autoscaler adds machines when there’s no room left for new pods.
Step 7: cache and database, the notepad and the filing room
The pod usually needs data. It has two places to get it:
- A cache such as Redis is a notepad on the desk. Recent answers are kept in memory, and reading one takes well under a millisecond.
- The database is the filing room, the single true record. Cache misses and every write go there.
Most databases have one primary, the only copy that accepts writes. Replicas are photocopies:
- a standby in another zone is updated on every write (synchronous replication), so it can take over within a minute or two if the primary’s building fails;
- a read replica in another region is updated a moment later (asynchronous replication). It can serve reads near its users and is a head start if a whole region fails. It may be slightly behind.
This is where distance matters again. A page might need ten database queries, one after another. With the app and database in the same region, that’s ten times about a millisecond. With the app in Singapore and the database in the US, it’s ten times 160 ms, well over a second, before the page can even be built. The app and its database live in the same region, and to serve users far away you copy data closer to them, not only servers.
Step 8: the way back
The response travels the same path backwards: pod, Service, ingress, load balancer, edge, browser. The browser reads the HTML, finds it needs dozens more files, fetches most of them from the edge over connections it already has open, and draws the page.
Part 3: regional or global
Everything so far lives in one region. Where to run things, and how many copies, is one of the main decisions in infra. These are the usual options, from simplest to hardest:
- One zone. Everything in one building. Cheapest and simplest, and fine for experiments. If the building has a bad day, you’re down.
- Several zones in one region. Pods spread across three zones, the database with a standby in a second zone. This survives the most common failures, a machine or a building, at little extra cost. For most services this is the right default.
- Two regions, active–passive. One region serves all traffic. A second one holds a copy of the data and a small, ready setup. If the first region fails, you fail over: promote the replica database, scale up, and point DNS or the global load balancer at the second region. It survives losing a whole city, but switching takes minutes, and anything not yet copied may be lost.
- Two regions, active–active. Both regions serve traffic, and a global load balancer sends users to the nearer one. Users everywhere get low latency, and losing a region only means shifting traffic. The hard part is data: if both regions accept writes, they can conflict. Teams either use a database built for this (such as Google Spanner or CockroachDB), or split users so each person’s data lives in one “home” region.
Two numbers make this choice concrete, and they’re worth agreeing on with the business before building anything:
- RTO (recovery time objective): how long you may be down. “Up again within an hour” is a very different system from “within a minute”.
- RPO (recovery point objective): how much recent data you may lose. Asynchronous copies mean a few seconds of writes can vanish in a failover. If that’s unacceptable, you need synchronous writes, which cost latency.
Other things push you towards more regions:
- Latency. Users far from every region pay the round-trip cost on every request. A region closer to them, or at least an edge that caches more, helps.
- Data residency. Some countries require that personal data about their citizens stays inside the country. That can force a region there, whatever the architecture would prefer.
- Cost. Traffic that leaves a region, or even crosses zones, is billed (egress). A chatty service split across regions can cost more in data transfer than in servers.
Part 4: operating it
Building the path is half the job. Keeping it healthy is the other half.
- Observability is the car’s dashboard. Metrics are the gauges: request rate, error rate, latency. Logs are the trip diary: what happened, line by line. Traces follow one request across every service it touched, showing where the time went.
- SLOs and alerts. Agree on what “healthy” means from the user’s side, for example “99.9% of requests succeed within 300 ms”. Alert when that promise is in danger, not every time a CPU is busy. Someone woken at night should always have something to fix.
- Safe deploys. A rolling deploy replaces pods a few at a time. A canary sends a small share of traffic to the new version first and compares errors. Blue–green runs both versions side by side and switches traffic in one step. All of them need a fast, practised rollback.
- Infrastructure as code. Networks, load balancers, clusters and databases are written as code, for example with Terraform, reviewed like code and applied by a pipeline. It’s the blueprint for the building: you can rebuild it, see every change, and create a second region from the same files.
- Backups and drills. A backup you’ve never restored is a hope, not a backup. Test restores regularly, and practise failing over to another zone or region before you need to.
- Capacity and cost. Autoscaling handles the daily rush. Planning handles the known peaks, like a big sale. Check the bill for surprises, especially data transfer.
- Secrets and access. Keep passwords and keys in a secret manager, not in code. Give every service and person the least access they need.
The whole trip in numbers
| Step | What happens | Typical cost for our Jakarta user |
|---|---|---|
| DNS | name to address | usually zero, it’s remembered |
| TCP and TLS | open an encrypted connection | two round trips, once per connection |
| CDN edge | images, scripts, styles | a few ms on a cache hit |
| Load balancer and cluster | pick a healthy pod | about a millisecond |
| App, cache, database | build the answer | a few to tens of ms, inside the region |
| Back to the browser | response, then more files | one round trip, plus drawing the page |
Best practices for infra and ops
This is the list I’d give my past self.
Latency
- Count round trips, then ask how far each one travels. Distance is the one cost you can’t optimise away.
- Reuse connections, and keep TLS handshakes close to users by ending them at the edge.
- Serve everything that’s the same for every user from a CDN, with versioned file names.
- Keep each app in the same region as its database. To be fast far away, copy data closer, not just servers.
Reliability
- Start with several zones in one region. Spread pods across zones, and keep a standby database in another zone.
- Decide RTO and RPO with the business first. They tell you whether you need a second region, and whether active–passive is enough.
- Put health checks on every backend, and make sure unhealthy ones are really taken out of rotation.
- Keep DNS TTLs short on records you may need to move quickly, and don’t depend on a single DNS provider for critical names.
Security
- Only the edge and load balancers are public. Apps and databases sit in private subnets.
- Allow traffic by rule, deny everything else, and give every service the least access it needs.
- Automate certificate renewal and keep secrets in a secret manager.
- Put a WAF and DDoS protection in front of everything public.
Deploying and changing
- Describe infrastructure as code, review it, and apply it through a pipeline.
- Ship in small steps with canaries or rolling deploys, and keep rollback one command away.
- Autoscale for daily changes in traffic, and plan capacity for known peaks.
Operating
- Measure what users feel, set SLOs on it, and alert on those, not on machine noise.
- Collect metrics, logs and traces from day one; you can’t add them during an outage.
- Test restores and failovers on a schedule. The first time shouldn’t be the real one.
- Watch the bill, especially data moving between zones and regions.
The Kubernetes part I’d spent months learning turned out to be the last few metres of a long trip. Knowing the rest of the route is what makes the infra decisions make sense.
References
- P. Mockapetris. Domain names - implementation and specification (RFC 1035). IETF, 1987. https://www.rfc-editor.org/rfc/rfc1035
- W. Eddy (ed.). Transmission Control Protocol (TCP) (RFC 9293). IETF, 2022. https://www.rfc-editor.org/rfc/rfc9293
- E. Rescorla. The Transport Layer Security (TLS) Protocol Version 1.3 (RFC 8446). IETF, 2018. https://www.rfc-editor.org/rfc/rfc8446
- M. Bishop (ed.). HTTP/3 (RFC 9114). IETF, 2022. https://www.rfc-editor.org/rfc/rfc9114
- J. Abley and K. Lindqvist. Operation of Anycast Services (RFC 4786). IETF, 2006. https://www.rfc-editor.org/rfc/rfc4786
- Google Cloud documentation. Geography and regions. cloud.google.com. https://cloud.google.com/docs/geography-and-regions
- Google Cloud documentation. Cloud Load Balancing overview. cloud.google.com. https://cloud.google.com/load-balancing/docs/load-balancing-overview
- Chris Jones, Jennifer Petoff and Betsy Beyer. Service Level Objectives. In Site Reliability Engineering, O’Reilly Media, 2016. https://sre.google/sre-book/service-level-objectives/