← All posts

Making sense of Kubernetes

My first cluster felt like a pile of YAML, charts and unfamiliar names. The one idea underneath it, the parts, namespaces, Helm, and how traffic finds a pod.

I learned Kubernetes because I had to. The services I worked on were deployed on it, mostly through Helm charts, so every change I shipped ended in a helm upgrade. The first time, I copied a chart, changed the image tag, ran the command, and it worked. I had no idea why. Over the following months I kept meeting new nouns: pods, ReplicaSets, Services, namespaces, ingresses, kubelets, CNI, charts, releases. Each one made sense on its own page of the docs. None of them fitted together.

This is the explanation I wish I’d had at the start. I’ve written it about a year in, so it’s the view of someone who uses Kubernetes every day, not someone who runs clusters for a living.

One idea underneath everything

Kubernetes is a database of what you want, plus many small programs that keep making the world match it.

You never tell Kubernetes “start three containers”. You write down “there should be three copies of this container” and store it. A controller then notices the difference between what’s written and what’s running, and acts to close the gap. If a machine dies and takes a copy with it, the same controller sees the gap again and starts a replacement. Nobody has to issue a new command.

The reconcile loop You declare a desired state of three replicas. A controller compares it with the observed state of two running pods and acts by starting one more pod. The loop repeats forever, and every controller runs one for its own kind of object. You declarereplicas: 3 The cluster2 pods running Controllercompare the two desired observed The controller closes the gap: one more pod act: start 1 pod repeat forever; each controller runs this loop for its own kind of object
Kubernetes never runs a script of steps. It keeps comparing what you asked for with what exists, and fixes the difference.

This is called reconciliation, and once I saw it, a lot of Kubernetes behaviour stopped surprising me. Deleting a pod that belongs to a deployment doesn’t get rid of it: a new one appears, because the desired state still says three. Editing a YAML file and applying it again is safe, because you’re replacing the target, not replaying a sequence of commands.

A deployment file is only a description of that target:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: web
spec:
  replicas: 3
  selector:
    matchLabels: {app: web}
  template:
    metadata:
      labels: {app: web}
    spec:
      containers:
        - name: web
          image: registry.example.com/web:2.1
          ports: [{containerPort: 8080}]

The parts of a cluster

A cluster splits into a control plane, which decides, and nodes, which run your containers.

The parts of a Kubernetes cluster The control plane holds the API server, etcd, the scheduler and the controllers. Everything talks to the API server, and only the API server touches etcd. Each node runs a kubelet, kube-proxy, a container runtime and pods. Kubelets watch the API server for pods assigned to their node. CONTROL PLANE API serverthe front door etcdstores all state schedulerpicks a node controllersreconcile loops everything talks to the API server; only the API server reads and writes etcd NODE 1 NODE 2 kubelet kube-proxy kubelet kube-proxy containerdcontainerd Pods on node 1podpod Pods on node 2podpodpod kubelets watch the API server
The control plane decides and records; the nodes do the work. On a managed cluster you only ever see the bottom half.

The control plane:

  • API server. The front door. kubectl, dashboards and every internal component talk to it over HTTP. It checks each request and is the only thing that reads and writes the store.
  • etcd. A small, consistent key-value store holding every object, both the desired state and the reported status. Lose etcd and the cluster forgets what it was supposed to be.
  • Scheduler. Watches for pods that have no node yet and picks one, based on the CPU and memory they request, node labels and spreading rules. It only decides. It doesn’t start anything.
  • Controller manager. A bundle of reconcile loops, one per kind of object: deployments, ReplicaSets, nodes, endpoints, jobs.

Each node runs:

  • kubelet. The node’s agent. It watches the API server for pods assigned to its node, asks the container runtime to start them, runs their health checks and reports their status.
  • Container runtime. containerd or similar, which pulls images and runs containers.
  • kube-proxy. Programs the node’s network rules so that Service addresses work. More on this below, because it was the part I understood last.

On a managed cluster like GKE, the cloud provider runs the control plane and you only see the nodes. That’s why I went months without thinking about etcd.

What took me longest to notice is that components don’t call each other. The scheduler never tells a kubelet to start a pod. It writes “this pod belongs on node 2” to the API server, and node 2’s kubelet, which is watching, picks it up. Every part coordinates by writing to, and watching, the same shared state.

What happens on kubectl apply

Following one deployment through the cluster is what made the parts click for me:

sequenceDiagram
  participant U as kubectl
  participant A as API server
  participant E as etcd
  participant C as Controllers
  participant S as Scheduler
  participant K as kubelet (node 2)
  U->>A: apply deployment.yaml
  A->>E: store the desired state
  A-->>C: watch: new Deployment
  C->>A: create a ReplicaSet, then 3 Pods
  A-->>S: watch: Pods with no node
  S->>A: bind each Pod to a node
  A-->>K: watch: a Pod is assigned to me
  K->>K: pull the image, start containers
  K->>A: status: Running

Every arrow is either a write to the API server or a watch on it. The deployment controller creates a ReplicaSet; the ReplicaSet controller creates the pods. None of them keeps its own list of pods as the truth: each reads the current state back from the API server every time it reconciles.

The objects you actually write

Most of what I write day to day is a handful of object kinds:

Object What it is
Pod One or more containers that share a network address and volumes, scheduled together. The smallest thing Kubernetes runs.
ReplicaSet Keeps N identical pods running. You rarely write one yourself.
Deployment Manages ReplicaSets, so a new version rolls out gradually and can roll back.
Service A stable name and address in front of a changing set of pods.
ConfigMap, Secret Configuration and credentials, mounted as files or environment variables.
Ingress HTTP routing from outside the cluster to Services.
Namespace A folder for objects, used for access control and quotas.

The objects find each other through labels. A Service doesn’t keep a list of its pods. It has a selector, app: web, and it matches whichever pods carry that label at this moment.

Ownership and labels A Deployment named web with three replicas owns a ReplicaSet, which owns three pods labelled app=web, with addresses 10.4.1.7, 10.4.2.3 and 10.4.1.9. A Service named web, with ClusterIP 10.96.0.12 and selector app=web, sends traffic to all three pods because of their label, not because anything links them by name. Deploymentweb · replicas: 3 ReplicaSetweb-7d9f owns podapp=web · 10.4.1.7 podapp=web · 10.4.2.3 podapp=web · 10.4.1.9 Service web10.96.0.12selector app=web The Service matches pods by label
Grey arrows are ownership: delete the Deployment and everything under it goes. Blue arrows are a label match: the Service follows whatever pods carry app=web right now.

When a deployment rolls out a new version, it creates a new ReplicaSet and scales it up while scaling the old one down. The Service never changes. The new pods carry the same label, so traffic follows them.

Namespaces

Every object above except the cluster-wide ones lives in a namespace. A namespace is a folder with a name: two teams can each have a Service called web as long as they’re in different namespaces. When you don’t say which one, you get default, which is why so many tutorials never mention them.

Namespaces are also where most of the cluster’s rules attach:

  • Names and DNS. A Service’s DNS name includes its namespace: web.payments.svc.cluster.local. From inside the same namespace, web is enough. From another one you write web.payments.
  • Access control. Roles and role bindings grant permissions within a namespace, so a team can be allowed to deploy to its own namespace and only read others.
  • Quotas and defaults. A ResourceQuota caps how much CPU, memory and how many objects a namespace can use; a LimitRange sets default requests and limits for pods that don’t declare them.
  • Separation of environments. A common pattern is one namespace per team, or per team and environment, such as payments-staging and payments-prod.

Namespaces don’t isolate the network on their own: by default a pod in one namespace can still reach a pod in any other. Blocking that takes a NetworkPolicy. And some objects belong to the whole cluster rather than to a namespace: nodes, persistent volumes, and namespaces themselves. kubectl needs -n payments, or a default namespace set in its context, to see anything outside default, and forgetting that flag was the most common reason I thought something had disappeared.

Helm: how the YAML actually gets there

In practice I almost never wrote those objects by hand. A real service needs a Deployment, a Service, a ConfigMap, maybe an Ingress and an autoscaler, and each environment needs slightly different versions of all of them: more replicas in production, a different hostname in staging. Copying YAML files per environment goes wrong fast. Helm is the tool we used instead, and it’s often called the package manager for Kubernetes.

A chart is a folder of templates plus default values:

web/
  Chart.yaml          name and version of the chart
  values.yaml         defaults: image tag, replicas, resources
  templates/
    deployment.yaml   Kubernetes YAML with placeholders
    service.yaml
    ingress.yaml

The templates are ordinary Kubernetes objects with placeholders in Go’s template syntax:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: {{ .Release.Name }}
spec:
  replicas: {{ .Values.replicas }}
  template:
    spec:
      containers:
        - name: web
          image: "{{ .Values.image.repository }}:{{ .Values.image.tag }}"

Each environment then only needs a small file with the values that differ:

# values-prod.yaml
replicas: 6
image:
  tag: "2.1"

Running helm upgrade --install web ./web -f values-prod.yaml -n payments merges the defaults with that file, renders the templates into plain Kubernetes YAML, and sends it to the API server, the same way kubectl apply would. From there everything in the sections above happens as before. Helm only decides what the desired state is; the controllers still do the work.

What Helm does A chart's templates are combined with default values and an environment's values file. Helm renders them into plain Kubernetes YAML and sends it to the API server. Each install or upgrade is recorded as a numbered revision of a release, here revisions 1 to 3, so helm rollback can return to revision 2. templates/with placeholders values.yamldefaults values-prod.yamlwhat differs helm upgrademerge, render plain YAMLDeployment, ... API serveras before Helm renders the chart into plain Kubernetes objects RELEASE "web" IN NAMESPACE payments revision 1image 1.9 revision 2image 2.0 revision 3image 2.1, bad helm rollback web 2 re-applies revision 2 helm rollback web 2
Helm is a templating and bookkeeping layer in front of the API server. It records every upgrade as a revision, which is what makes a one-line rollback possible.

The installed copy of a chart is a release, with a name and a namespace. Helm stores each install and upgrade as a numbered revision, so helm history web lists them and helm rollback web 2 puts back what revision 2 rendered. Because of that, releases, not individual objects, became the thing I thought in: “which revision of web is running in production” is a Helm question.

A few things took me a while with Helm:

  • helm template renders a chart locally without installing it. When a chart misbehaved, reading the rendered YAML was almost always faster than reading the templates.
  • Helm doesn’t replace understanding the objects. When a pod crashes, the answer is still in kubectl describe and the logs, not in Helm.
  • Values files are where environments differ, so they’re where most mistakes are: a wrong indent, or a key the template never reads, fails silently.
  • Charts can depend on other charts, which is how a single command installs a database or a monitoring stack from a public chart repository.

Networking, the part I understood last

Kubernetes makes two promises about the network, and hands the job of keeping them to a plugin:

  • Every pod gets its own IP address.
  • Any pod can reach any other pod at that address, on any node, without address translation.

A CNI (Container Network Interface) plugin keeps those promises, for example by giving each node a range of addresses and setting up routes between nodes. On GKE it’s configured for you.

Pod addresses aren’t much use on their own, because pods are replaced all the time and every replacement gets a new one. That’s what a Service is for. Creating one gives you a DNS name, such as web.default.svc.cluster.local for a Service called web in the default namespace, and a virtual address, the ClusterIP, which stays the same for as long as the Service exists.

What surprised me is that the ClusterIP doesn’t belong to any machine. No network interface has it. kube-proxy runs on every node, watches Services and their pods, and writes rules into the node’s kernel (iptables, or IPVS) that say: a packet for 10.96.0.12 port 80 should go to one of these pod addresses instead, picked at random. The rewrite happens on the sending node, before the packet leaves it.

How a request finds a pod A client pod asks CoreDNS for the name web and gets the ClusterIP 10.96.0.12. It sends to that address. Rules on its own node, written by kube-proxy, rewrite the destination to one pod, 10.4.2.3 on node B, chosen at random once per connection. Other web pods such as 10.4.1.7 could have been chosen instead. The ClusterIP exists only as a rule. CoreDNSweb → 10.96.0.12 client podGET http://web 1 look up the name node ruleswritten by kube-proxy 2 send web pod, node B10.4.2.3 3 the destination is rewritten to one pod3 rewrite web pod · 10.4.1.7 one pod per connection,picked at random 10.96.0.12 exists only as a rule on each node; no machine has that address
A Service is a set of rules copied to every node, not a box traffic passes through. The pick happens once, when the connection opens.

The detail that mattered to me later is that the choice happens once per connection. A client that opens a connection and keeps it open, as HTTP/2 and gRPC clients do, sends every request down it to the same pod. With a few long-lived clients, load can end up badly uneven, even though the Service is “balancing” traffic.

Traffic from outside the cluster comes in through a Service of type LoadBalancer, which asks the cloud provider for a load balancer, or through an Ingress, which routes HTTP by host and path to Services.

Where a service mesh comes in

Any discussion of Kubernetes networking eventually mentions Istio, so I read up on what a service mesh adds.

A mesh puts a small proxy, usually Envoy, next to every pod as an extra container, called a sidecar. Rules inside the pod send all its traffic in and out through that proxy, so the application talks plain HTTP and the proxy handles the rest. The mesh’s control plane configures all the proxies.

A service mesh with sidecars Two pods each contain an app container and a proxy container. The app in pod A sends a plain request to its own proxy, which sends it over mutual TLS to the proxy in pod B, which hands it to the app in pod B. A mesh control plane sends configuration to both proxies. mesh control planeistiod, in Istio POD A app proxy POD B proxy app Proxy to proxy, encrypted with mutual TLSmTLS config retries, timeouts, encryption and metrics move from app code into the proxy
The application only ever talks to the proxy next to it. Everything between the two proxies is the mesh's job.

Because the proxy understands HTTP and gRPC, it can do things kube-proxy can’t:

  • balance per request instead of per connection, which fixes the gRPC problem above;
  • retry failed requests and enforce timeouts;
  • encrypt traffic between pods with mutual TLS, without any change to application code;
  • split traffic by percentage, for canary releases;
  • report latency and error rates for every call between services.

The price is an extra proxy on every hop, more memory per pod, and one more system to understand when something breaks. I haven’t used a mesh myself yet. Reading about one still helped me place kube-proxy: it works on connections, and a mesh works on requests.

What made it click

  • Treat YAML as a description of the end state, not a list of steps.
  • Helm writes the YAML; Kubernetes still does the work. When in doubt, render the chart and read what it produced.
  • When something is wrong, read the object’s status and events with kubectl describe before the logs. They tell you which controller gave up, and why.
  • Components coordinate through the API server, never directly.
  • A Service is a set of rules on every node, not a box in the middle.
  • Load is balanced per connection, unless something above kube-proxy does better.

Kubernetes is still big. But it’s one idea applied many times, and seeing that made the rest learnable.

References

  1. Kubernetes documentation. Controllers. kubernetes.io. https://kubernetes.io/docs/concepts/architecture/controller/
  2. Kubernetes documentation. Kubernetes Components. kubernetes.io. https://kubernetes.io/docs/concepts/overview/components/
  3. Kubernetes documentation. Service. kubernetes.io. https://kubernetes.io/docs/concepts/services-networking/service/
  4. Kubernetes documentation. Namespaces. kubernetes.io. https://kubernetes.io/docs/concepts/overview/working-with-objects/namespaces/
  5. Kubernetes documentation. Network Policies. kubernetes.io. https://kubernetes.io/docs/concepts/services-networking/network-policies/
  6. Helm. Charts. helm.sh. https://helm.sh/docs/topics/charts/
  7. Kubernetes documentation. Ingress. kubernetes.io. https://kubernetes.io/docs/concepts/services-networking/ingress/
  8. Istio. Architecture. istio.io. https://istio.io/latest/docs/ops/deployment/architecture/ — sidecar proxies and the mesh control plane.