← All posts

From a click to a database: how a request crosses the internet

Every step from typing an address to a row in a database, each piece explained with an everyday analogy, then how the same pieces shape regional and global deployments, and the infra and ops practices that follow.

After making sense of Kubernetes, I understood what happens to a request once it reaches our cluster. Everything before that was a blur. Someone types an address, and a fraction of a second later a page appears, built from data in a database in another country. I also kept hearing infra words I couldn’t place: region, zone, VPC, CDN, anycast, active–active, failover.

This post follows one request from start to finish and explains each piece it meets, with an everyday analogy for each. Then it zooms out to the decisions infra and ops teams make with those pieces: where to run things, how to survive failures, and how to operate it all. It ends with the practices I’d now call best practice.

The example throughout: someone in Jakarta opens shop.example.com, and the shop runs in a cloud region in Singapore.

The path of one request 1: the browser in Jakarta asks a DNS resolver for the shop's address. 2: it fetches static files such as images from a CDN edge in Jakarta. 3: API requests go on to a load balancer in the Singapore region, then into the Kubernetes cluster, then to an app pod. 4: the app reads from a Redis cache and a database primary. The database copies its data to a read replica in another region, Tokyo. DNS resolvername → address browserin Jakarta CDN edgealso in Jakarta 1 2 REGION: SINGAPORE load balancerpublic address Kubernetesingress → Service app podsyour code Rediscache databaseprimary 3: API requests go on to the region3 4 another region: Tokyoread replica, a little behind copies
The whole trip in one picture. The rest of the post walks through it piece by piece.

The cast, in one table

Before the details, here is every component in this post with the analogy I’ll use for it. The analogies are about a city and its postal system.

Component Analogy What it does
Packet a letter a small chunk of data with a sender and destination address
IP address a street address where a machine can be reached
Router a post office passes each packet one step closer to its destination
DNS the phone book turns a name like shop.example.com into an IP address
TCP registered mail numbers packets, confirms delivery, resends lost ones
TLS a sealed envelope and an ID check encrypts traffic and proves the server is who it says
CDN edge a local convenience store keeps copies of popular files close to users
Anycast one hotline, answered by the nearest branch one IP address served from many places at once
WAF and DDoS protection the bouncer turns away bad traffic before it gets in
Load balancer the restaurant host sends each guest to a free, working table
Region a city a cluster of data centres in one metro area
Zone a building in that city a data centre with its own power and cooling
VPC and subnets a fenced campus with a public lobby and private offices your private network inside the cloud
Firewall rules the security guard at each door decide which traffic may reach which machine
NAT gateway the mailroom lets private machines send mail out without having a public address
Kubernetes the building manager keeps the right number of workers at their desks
Pod a worker runs your code
Cache a notepad on the desk quick answers to repeated questions
Database the filing room the one true record of everything
Replica a photocopy of the files a copy for reading or for emergencies

Part 1: getting there

The internet itself

The internet is thousands of separate networks that have agreed to carry each other’s traffic. Your mobile provider is one network, a cloud provider is another, and they connect at exchange points and through undersea cables.

Everything travels as packets, which are like letters. Each has a sender and a destination IP address, such as 203.0.113.10, the way a letter has a street address. Routers are the post offices: each reads the destination and passes the packet one step closer. Networks tell each other which addresses they can deliver to using a protocol called BGP, which is like post offices sharing their delivery maps.

One fact shaped everything else for me: distance costs time, and nothing makes it free. Light in fibre travels about 200,000 km per second. A packet from Jakarta to Singapore and back covers about 1,800 km, which takes at least 9 ms. To a US data centre and back it’s over 30,000 km, at least 157 ms. Cables don’t run in straight lines, so real numbers are higher.

Minimum round trip from Jakarta The theoretical minimum round-trip time from Jakarta through fibre: Singapore, 880 km, 9 ms. Tokyo, 5,780 km, 58 ms. Frankfurt, 11,000 km, 110 ms. Iowa in the US, 15,700 km, 157 ms. Singapore880 km Singapore: 880 km, at least 9 ms round trip 9 ms Tokyo5,780 km Tokyo: 5,780 km, at least 58 ms round trip 58 ms Frankfurt11,000 km Frankfurt: 11,000 km, at least 110 ms round trip 110 ms Iowa, US15,700 km Iowa, US: 15,700 km, at least 157 ms round trip 157 ms 0 ms 50 ms 100 ms 150 ms
Straight-line distance at the speed of light in fibre. This is the floor; real routes add more. Every step below that needs a round trip pays this cost.

That there-and-back time is the round-trip time. Most of the rest of this post is about how many round trips a request needs and how far each one travels.

Step 1: DNS, the phone book

Browsers need an IP address, not a name. DNS, the Domain Name System, is the phone book that turns one into the other.

The browser asks a resolver, usually run by your internet provider or a public service like Google’s 8.8.8.8. If the resolver hasn’t looked the name up recently, it walks down a chain, like asking directory enquiries, who refers you to the right regional phone book, which gives you the number:

sequenceDiagram
  participant B as Browser
  participant R as Resolver
  participant Root as Root server
  participant TLD as .com server
  participant A as example.com's DNS
  B->>R: address of shop.example.com?
  R->>Root: shop.example.com?
  Root-->>R: ask the .com servers
  R->>TLD: shop.example.com?
  TLD-->>R: ask example.com's DNS
  R->>A: shop.example.com?
  A-->>R: 203.0.113.10, keep it for 5 minutes
  R-->>B: 203.0.113.10

Each answer comes with a TTL (time to live), which says how long it may be remembered. The resolver, the operating system and the browser all keep recent answers, so most lookups take no time at all.

Why ops cares:

  • TTL is a trade-off. A short TTL, like 60 seconds, means that when you point the name somewhere new, say during a failover, the world follows within a minute. A long TTL means fewer lookups but slow changes.
  • DNS can steer traffic. GeoDNS gives different answers depending on where the question comes from, so users in Jakarta and Frankfurt can be sent to different regions under the same name.
  • DNS is a single point of failure if one provider hosts all your records. Big outages have come from exactly this.

Step 2: TCP and TLS, registered mail in a sealed envelope

With an address, the browser opens a connection. Two handshakes happen before any page data moves.

  • TCP is registered mail. Packets can be lost or arrive out of order; TCP numbers them, confirms each delivery, resends what’s missing and puts them back in order. Setting it up takes one round trip.
  • TLS is a sealed envelope plus an ID check. The server shows a certificate proving it owns the name, and both sides agree on keys for encryption. It’s the “S” in HTTPS. TLS 1.3 takes one more round trip; older versions took two.

Only then does the browser send the actual request, and the answer takes a third round trip. So a first response on a new connection costs at least three round trips, plus the server’s own work.

Time to the first response on a new connection Three round trips plus server time. A server in Iowa, 160 ms away: 320 ms of handshakes and 210 ms for the request, 530 ms in total. A server in Singapore, 10 ms away: 20 ms of handshakes and 60 ms for the request, 80 ms. A CDN edge in Jakarta, 2 ms away, serving a cached file: 4 ms of handshakes and 3 ms for the request, 7 ms. TCP + TLS handshakes request, response and server work server in Iowa160 ms round trip server in Iowa: handshakes 320 ms (TCP + TLS) server in Iowa: request and response 210 ms, including 50 ms of server work 530 ms server in Singapore10 ms round trip server in Singapore: handshakes 20 ms (TCP + TLS) server in Singapore: request and response 60 ms, including 50 ms of server work 80 ms CDN edge, Jakarta2 ms round trip CDN edge, Jakarta: handshakes 4 ms (TCP + TLS) CDN edge, Jakarta: request and response 3 ms, including 1 ms of server work 7 ms
Illustrative round-trip times, with 50 ms of server work for the app and a cached file at the edge. Distance multiplies: three round trips to Iowa cost more than the whole request to Singapore.

Why ops cares:

  • Reuse connections. Browsers keep them open, HTTP/2 sends many requests over one, and HTTP/3 (running on a newer transport called QUIC) combines the handshakes. Every new connection pays the handshakes again.
  • Certificates expire. An expired certificate takes a site down as surely as a crashed server. Automate renewal, for example with Let’s Encrypt or your cloud’s managed certificates, and alert well before expiry.
  • Where TLS ends matters. Usually the CDN or load balancer decrypts traffic, which is called TLS termination, so it can route by URL. Traffic behind it is re-encrypted or kept inside a private network.

Step 3: the edge, a convenience store near home

Much of a web page is the same for everyone: images, scripts, stylesheets, fonts. You don’t drive to the central warehouse for a bottle of water; you buy it at the shop on your street. A CDN (content delivery network) is that shop: a provider with servers in hundreds of cities, each keeping copies of your popular files.

For our Jakarta user, the edge is a couple of milliseconds away. The handshakes happen with it, cheaply. On a cache hit it answers directly. On a miss, it fetches the file from the region once, over a connection it already keeps open, and keeps the copy for the next person.

Many CDNs and global load balancers use anycast: the same IP address is announced from many locations at once, and internet routing delivers each user to the nearest one. It’s like one national hotline number that’s answered by whichever branch is closest to the caller.

The edge is also where the bouncer stands. A WAF (web application firewall) blocks known attack patterns, and DDoS protection absorbs floods of fake traffic across the whole edge network, before any of it reaches your servers.

Why ops cares:

  • Cache rules are yours to set. Response headers such as Cache-Control tell the edge how long to keep a file.
  • Version file names, like app.3f9c.js. A new release then gets a new name, so nobody sees a stale copy, and you rarely need to clear the cache by hand.
  • Put all public traffic behind the edge, so the bouncer sees it first.

Part 2: inside the region

Step 4: regions, zones and the load balancer

Requests that need fresh, personal data, like “show my basket”, go on to the region.

A region is a city: a cloud provider’s group of data centres in one metro area, such as Singapore. A zone is one building in that city, with its own power, cooling and network, a few kilometres from the others and linked to them by fast private lines. A round trip between zones takes about a millisecond. If one building loses power, the others keep running.

The request arrives at a load balancer, the restaurant host. It owns the public address, ends the TLS connection, checks which tables (backends) are free and working, and seats the guest at one of them. It learns which backends work through health checks: it asks each one “are you OK?” every few seconds and stops sending traffic to those that don’t answer.

A region with three zones A Singapore region with three zones, a, b and c. A load balancer spreads requests across app pods in all three zones. The database primary is in zone a, with a standby copy in zone b that is updated on every write. If zone a fails, the pods in b and c keep serving and the standby becomes the primary. load balancer ZONE A ZONE B ZONE C app pods app pods app pods Database primary in zone a: takes all writesdatabase primaryall writes go here standbytakes over if A fails
Zones are what make a region survive a building going dark. Pods spread across all three; the database keeps a standby copy in a second zone.

Load balancers come in two kinds that are worth knowing apart:

  • Layer 4 balancers look only at addresses and ports. They are fast, and they pick a backend once per connection.
  • Layer 7 balancers understand HTTP. They can route by path or hostname, retry a failed request, and balance every request separately.

Cloud providers also offer global load balancers: one anycast address worldwide that sends each user to the nearest healthy region. Regional ones serve a single region.

Step 5: the private network, a fenced campus

Inside the region, your machines live in a VPC (virtual private cloud): a private network that only you use. Think of a fenced campus. It’s split into subnets, which are areas of the campus:

  • a public subnet is the lobby: only the load balancer lives there, with an address the internet can reach;
  • private subnets are the offices: app servers and databases, with no public address at all.

Firewall rules (called security groups on some clouds) are the guards at each door: “the load balancer may talk to the app on port 8080”, “only the app may talk to the database on port 5432”. Everything else is refused.

Private machines still sometimes need to reach out, for example to download a package or call a partner’s API. A NAT gateway is the mailroom: they hand their outgoing mail to it, it sends it with its own return address, and it passes the replies back. Nobody outside can send mail straight to an office.

Step 6: Kubernetes, the building manager

In our setup, the load balancer hands requests to a Kubernetes cluster whose machines sit in the private subnets, spread across the zones. An ingress reads the hostname and path, a Service picks one of the matching pods (the workers), and the pod runs our code.

The manager’s most useful habit for ops is autoscaling, which is calling in extra staff for the lunch rush:

  • the horizontal pod autoscaler adds pods when CPU or request rate goes up, and removes them when it drops;
  • the cluster autoscaler adds machines when there’s no room left for new pods.

Step 7: cache and database, the notepad and the filing room

The pod usually needs data. It has two places to get it:

  • A cache such as Redis is a notepad on the desk. Recent answers are kept in memory, and reading one takes well under a millisecond.
  • The database is the filing room, the single true record. Cache misses and every write go there.

Most databases have one primary, the only copy that accepts writes. Replicas are photocopies:

  • a standby in another zone is updated on every write (synchronous replication), so it can take over within a minute or two if the primary’s building fails;
  • a read replica in another region is updated a moment later (asynchronous replication). It can serve reads near its users and is a head start if a whole region fails. It may be slightly behind.

This is where distance matters again. A page might need ten database queries, one after another. With the app and database in the same region, that’s ten times about a millisecond. With the app in Singapore and the database in the US, it’s ten times 160 ms, well over a second, before the page can even be built. The app and its database live in the same region, and to serve users far away you copy data closer to them, not only servers.

Step 8: the way back

The response travels the same path backwards: pod, Service, ingress, load balancer, edge, browser. The browser reads the HTML, finds it needs dozens more files, fetches most of them from the edge over connections it already has open, and draws the page.

Part 3: regional or global

Everything so far lives in one region. Where to run things, and how many copies, is one of the main decisions in infra. These are the usual options, from simplest to hardest:

Four ways to deploy One: everything in one zone; if the building fails you are down. Two: several zones in one region; survives a building, not a city. Three: two regions, active and passive; Singapore serves traffic and copies data to a standby in Tokyo; survives a region, with minutes to switch. Four: two regions, active and active, behind a global load balancer; both serve traffic; survives a region, with the hardest data problems. 1. One zonea building fails: you're down Singapore zone aapp, db 2. Zones in one regionsurvives a building, not a city Singapore zone aapp, db zone bapp, copy zone capp 3. Two regions, active–passivesurvives a city; minutes to switch Singaporeserving Tokyostanding by copies data 4. Two regions, active–activesurvives a city; hardest for data global LB Singaporeserving Tokyoserving
Each step up survives a bigger failure and costs more, in money and in complexity. Most services should start at 2.
  1. One zone. Everything in one building. Cheapest and simplest, and fine for experiments. If the building has a bad day, you’re down.
  2. Several zones in one region. Pods spread across three zones, the database with a standby in a second zone. This survives the most common failures, a machine or a building, at little extra cost. For most services this is the right default.
  3. Two regions, active–passive. One region serves all traffic. A second one holds a copy of the data and a small, ready setup. If the first region fails, you fail over: promote the replica database, scale up, and point DNS or the global load balancer at the second region. It survives losing a whole city, but switching takes minutes, and anything not yet copied may be lost.
  4. Two regions, active–active. Both regions serve traffic, and a global load balancer sends users to the nearer one. Users everywhere get low latency, and losing a region only means shifting traffic. The hard part is data: if both regions accept writes, they can conflict. Teams either use a database built for this (such as Google Spanner or CockroachDB), or split users so each person’s data lives in one “home” region.

Two numbers make this choice concrete, and they’re worth agreeing on with the business before building anything:

  • RTO (recovery time objective): how long you may be down. “Up again within an hour” is a very different system from “within a minute”.
  • RPO (recovery point objective): how much recent data you may lose. Asynchronous copies mean a few seconds of writes can vanish in a failover. If that’s unacceptable, you need synchronous writes, which cost latency.

Other things push you towards more regions:

  • Latency. Users far from every region pay the round-trip cost on every request. A region closer to them, or at least an edge that caches more, helps.
  • Data residency. Some countries require that personal data about their citizens stays inside the country. That can force a region there, whatever the architecture would prefer.
  • Cost. Traffic that leaves a region, or even crosses zones, is billed (egress). A chatty service split across regions can cost more in data transfer than in servers.

Part 4: operating it

Building the path is half the job. Keeping it healthy is the other half.

  • Observability is the car’s dashboard. Metrics are the gauges: request rate, error rate, latency. Logs are the trip diary: what happened, line by line. Traces follow one request across every service it touched, showing where the time went.
  • SLOs and alerts. Agree on what “healthy” means from the user’s side, for example “99.9% of requests succeed within 300 ms”. Alert when that promise is in danger, not every time a CPU is busy. Someone woken at night should always have something to fix.
  • Safe deploys. A rolling deploy replaces pods a few at a time. A canary sends a small share of traffic to the new version first and compares errors. Blue–green runs both versions side by side and switches traffic in one step. All of them need a fast, practised rollback.
  • Infrastructure as code. Networks, load balancers, clusters and databases are written as code, for example with Terraform, reviewed like code and applied by a pipeline. It’s the blueprint for the building: you can rebuild it, see every change, and create a second region from the same files.
  • Backups and drills. A backup you’ve never restored is a hope, not a backup. Test restores regularly, and practise failing over to another zone or region before you need to.
  • Capacity and cost. Autoscaling handles the daily rush. Planning handles the known peaks, like a big sale. Check the bill for surprises, especially data transfer.
  • Secrets and access. Keep passwords and keys in a secret manager, not in code. Give every service and person the least access they need.

The whole trip in numbers

Step What happens Typical cost for our Jakarta user
DNS name to address usually zero, it’s remembered
TCP and TLS open an encrypted connection two round trips, once per connection
CDN edge images, scripts, styles a few ms on a cache hit
Load balancer and cluster pick a healthy pod about a millisecond
App, cache, database build the answer a few to tens of ms, inside the region
Back to the browser response, then more files one round trip, plus drawing the page

Best practices for infra and ops

This is the list I’d give my past self.

Latency

  • Count round trips, then ask how far each one travels. Distance is the one cost you can’t optimise away.
  • Reuse connections, and keep TLS handshakes close to users by ending them at the edge.
  • Serve everything that’s the same for every user from a CDN, with versioned file names.
  • Keep each app in the same region as its database. To be fast far away, copy data closer, not just servers.

Reliability

  • Start with several zones in one region. Spread pods across zones, and keep a standby database in another zone.
  • Decide RTO and RPO with the business first. They tell you whether you need a second region, and whether active–passive is enough.
  • Put health checks on every backend, and make sure unhealthy ones are really taken out of rotation.
  • Keep DNS TTLs short on records you may need to move quickly, and don’t depend on a single DNS provider for critical names.

Security

  • Only the edge and load balancers are public. Apps and databases sit in private subnets.
  • Allow traffic by rule, deny everything else, and give every service the least access it needs.
  • Automate certificate renewal and keep secrets in a secret manager.
  • Put a WAF and DDoS protection in front of everything public.

Deploying and changing

  • Describe infrastructure as code, review it, and apply it through a pipeline.
  • Ship in small steps with canaries or rolling deploys, and keep rollback one command away.
  • Autoscale for daily changes in traffic, and plan capacity for known peaks.

Operating

  • Measure what users feel, set SLOs on it, and alert on those, not on machine noise.
  • Collect metrics, logs and traces from day one; you can’t add them during an outage.
  • Test restores and failovers on a schedule. The first time shouldn’t be the real one.
  • Watch the bill, especially data moving between zones and regions.

The Kubernetes part I’d spent months learning turned out to be the last few metres of a long trip. Knowing the rest of the route is what makes the infra decisions make sense.

References

  1. P. Mockapetris. Domain names - implementation and specification (RFC 1035). IETF, 1987. https://www.rfc-editor.org/rfc/rfc1035
  2. W. Eddy (ed.). Transmission Control Protocol (TCP) (RFC 9293). IETF, 2022. https://www.rfc-editor.org/rfc/rfc9293
  3. E. Rescorla. The Transport Layer Security (TLS) Protocol Version 1.3 (RFC 8446). IETF, 2018. https://www.rfc-editor.org/rfc/rfc8446
  4. M. Bishop (ed.). HTTP/3 (RFC 9114). IETF, 2022. https://www.rfc-editor.org/rfc/rfc9114
  5. J. Abley and K. Lindqvist. Operation of Anycast Services (RFC 4786). IETF, 2006. https://www.rfc-editor.org/rfc/rfc4786
  6. Google Cloud documentation. Geography and regions. cloud.google.com. https://cloud.google.com/docs/geography-and-regions
  7. Google Cloud documentation. Cloud Load Balancing overview. cloud.google.com. https://cloud.google.com/load-balancing/docs/load-balancing-overview
  8. Chris Jones, Jennifer Petoff and Betsy Beyer. Service Level Objectives. In Site Reliability Engineering, O’Reilly Media, 2016. https://sre.google/sre-book/service-level-objectives/