Making sense of Kubernetes
My first cluster felt like a pile of YAML, charts and unfamiliar names. The one idea underneath it, the parts, namespaces, Helm, and how traffic finds a pod.
I learned Kubernetes because I had to. The services I worked on were deployed on it, mostly through Helm charts, so every change I shipped ended in a helm upgrade. The first time, I copied a chart, changed the image tag, ran the command, and it worked. I had no idea why. Over the following months I kept meeting new nouns: pods, ReplicaSets, Services, namespaces, ingresses, kubelets, CNI, charts, releases. Each one made sense on its own page of the docs. None of them fitted together.
This is the explanation I wish I’d had at the start. I’ve written it about a year in, so it’s the view of someone who uses Kubernetes every day, not someone who runs clusters for a living.
One idea underneath everything
Kubernetes is a database of what you want, plus many small programs that keep making the world match it.
You never tell Kubernetes “start three containers”. You write down “there should be three copies of this container” and store it. A controller then notices the difference between what’s written and what’s running, and acts to close the gap. If a machine dies and takes a copy with it, the same controller sees the gap again and starts a replacement. Nobody has to issue a new command.
This is called reconciliation, and once I saw it, a lot of Kubernetes behaviour stopped surprising me. Deleting a pod that belongs to a deployment doesn’t get rid of it: a new one appears, because the desired state still says three. Editing a YAML file and applying it again is safe, because you’re replacing the target, not replaying a sequence of commands.
A deployment file is only a description of that target:
apiVersion: apps/v1
kind: Deployment
metadata:
name: web
spec:
replicas: 3
selector:
matchLabels: {app: web}
template:
metadata:
labels: {app: web}
spec:
containers:
- name: web
image: registry.example.com/web:2.1
ports: [{containerPort: 8080}]
The parts of a cluster
A cluster splits into a control plane, which decides, and nodes, which run your containers.
The control plane:
- API server. The front door.
kubectl, dashboards and every internal component talk to it over HTTP. It checks each request and is the only thing that reads and writes the store. - etcd. A small, consistent key-value store holding every object, both the desired state and the reported status. Lose etcd and the cluster forgets what it was supposed to be.
- Scheduler. Watches for pods that have no node yet and picks one, based on the CPU and memory they request, node labels and spreading rules. It only decides. It doesn’t start anything.
- Controller manager. A bundle of reconcile loops, one per kind of object: deployments, ReplicaSets, nodes, endpoints, jobs.
Each node runs:
- kubelet. The node’s agent. It watches the API server for pods assigned to its node, asks the container runtime to start them, runs their health checks and reports their status.
- Container runtime. containerd or similar, which pulls images and runs containers.
- kube-proxy. Programs the node’s network rules so that Service addresses work. More on this below, because it was the part I understood last.
On a managed cluster like GKE, the cloud provider runs the control plane and you only see the nodes. That’s why I went months without thinking about etcd.
What took me longest to notice is that components don’t call each other. The scheduler never tells a kubelet to start a pod. It writes “this pod belongs on node 2” to the API server, and node 2’s kubelet, which is watching, picks it up. Every part coordinates by writing to, and watching, the same shared state.
What happens on kubectl apply
Following one deployment through the cluster is what made the parts click for me:
sequenceDiagram
participant U as kubectl
participant A as API server
participant E as etcd
participant C as Controllers
participant S as Scheduler
participant K as kubelet (node 2)
U->>A: apply deployment.yaml
A->>E: store the desired state
A-->>C: watch: new Deployment
C->>A: create a ReplicaSet, then 3 Pods
A-->>S: watch: Pods with no node
S->>A: bind each Pod to a node
A-->>K: watch: a Pod is assigned to me
K->>K: pull the image, start containers
K->>A: status: Running
Every arrow is either a write to the API server or a watch on it. The deployment controller creates a ReplicaSet; the ReplicaSet controller creates the pods. None of them keeps its own list of pods as the truth: each reads the current state back from the API server every time it reconciles.
The objects you actually write
Most of what I write day to day is a handful of object kinds:
| Object | What it is |
|---|---|
| Pod | One or more containers that share a network address and volumes, scheduled together. The smallest thing Kubernetes runs. |
| ReplicaSet | Keeps N identical pods running. You rarely write one yourself. |
| Deployment | Manages ReplicaSets, so a new version rolls out gradually and can roll back. |
| Service | A stable name and address in front of a changing set of pods. |
| ConfigMap, Secret | Configuration and credentials, mounted as files or environment variables. |
| Ingress | HTTP routing from outside the cluster to Services. |
| Namespace | A folder for objects, used for access control and quotas. |
The objects find each other through labels. A Service doesn’t keep a list of its pods. It has a selector, app: web, and it matches whichever pods carry that label at this moment.
When a deployment rolls out a new version, it creates a new ReplicaSet and scales it up while scaling the old one down. The Service never changes. The new pods carry the same label, so traffic follows them.
Namespaces
Every object above except the cluster-wide ones lives in a namespace. A namespace is a folder with a name: two teams can each have a Service called web as long as they’re in different namespaces. When you don’t say which one, you get default, which is why so many tutorials never mention them.
Namespaces are also where most of the cluster’s rules attach:
- Names and DNS. A Service’s DNS name includes its namespace:
web.payments.svc.cluster.local. From inside the same namespace,webis enough. From another one you writeweb.payments. - Access control. Roles and role bindings grant permissions within a namespace, so a team can be allowed to deploy to its own namespace and only read others.
- Quotas and defaults. A ResourceQuota caps how much CPU, memory and how many objects a namespace can use; a LimitRange sets default requests and limits for pods that don’t declare them.
- Separation of environments. A common pattern is one namespace per team, or per team and environment, such as
payments-stagingandpayments-prod.
Namespaces don’t isolate the network on their own: by default a pod in one namespace can still reach a pod in any other. Blocking that takes a NetworkPolicy. And some objects belong to the whole cluster rather than to a namespace: nodes, persistent volumes, and namespaces themselves. kubectl needs -n payments, or a default namespace set in its context, to see anything outside default, and forgetting that flag was the most common reason I thought something had disappeared.
Helm: how the YAML actually gets there
In practice I almost never wrote those objects by hand. A real service needs a Deployment, a Service, a ConfigMap, maybe an Ingress and an autoscaler, and each environment needs slightly different versions of all of them: more replicas in production, a different hostname in staging. Copying YAML files per environment goes wrong fast. Helm is the tool we used instead, and it’s often called the package manager for Kubernetes.
A chart is a folder of templates plus default values:
web/
Chart.yaml name and version of the chart
values.yaml defaults: image tag, replicas, resources
templates/
deployment.yaml Kubernetes YAML with placeholders
service.yaml
ingress.yaml
The templates are ordinary Kubernetes objects with placeholders in Go’s template syntax:
apiVersion: apps/v1
kind: Deployment
metadata:
name: {{ .Release.Name }}
spec:
replicas: {{ .Values.replicas }}
template:
spec:
containers:
- name: web
image: "{{ .Values.image.repository }}:{{ .Values.image.tag }}"
Each environment then only needs a small file with the values that differ:
# values-prod.yaml
replicas: 6
image:
tag: "2.1"
Running helm upgrade --install web ./web -f values-prod.yaml -n payments merges the defaults with that file, renders the templates into plain Kubernetes YAML, and sends it to the API server, the same way kubectl apply would. From there everything in the sections above happens as before. Helm only decides what the desired state is; the controllers still do the work.
The installed copy of a chart is a release, with a name and a namespace. Helm stores each install and upgrade as a numbered revision, so helm history web lists them and helm rollback web 2 puts back what revision 2 rendered. Because of that, releases, not individual objects, became the thing I thought in: “which revision of web is running in production” is a Helm question.
A few things took me a while with Helm:
helm templaterenders a chart locally without installing it. When a chart misbehaved, reading the rendered YAML was almost always faster than reading the templates.- Helm doesn’t replace understanding the objects. When a pod crashes, the answer is still in
kubectl describeand the logs, not in Helm. - Values files are where environments differ, so they’re where most mistakes are: a wrong indent, or a key the template never reads, fails silently.
- Charts can depend on other charts, which is how a single command installs a database or a monitoring stack from a public chart repository.
Networking, the part I understood last
Kubernetes makes two promises about the network, and hands the job of keeping them to a plugin:
- Every pod gets its own IP address.
- Any pod can reach any other pod at that address, on any node, without address translation.
A CNI (Container Network Interface) plugin keeps those promises, for example by giving each node a range of addresses and setting up routes between nodes. On GKE it’s configured for you.
Pod addresses aren’t much use on their own, because pods are replaced all the time and every replacement gets a new one. That’s what a Service is for. Creating one gives you a DNS name, such as web.default.svc.cluster.local for a Service called web in the default namespace, and a virtual address, the ClusterIP, which stays the same for as long as the Service exists.
What surprised me is that the ClusterIP doesn’t belong to any machine. No network interface has it. kube-proxy runs on every node, watches Services and their pods, and writes rules into the node’s kernel (iptables, or IPVS) that say: a packet for 10.96.0.12 port 80 should go to one of these pod addresses instead, picked at random. The rewrite happens on the sending node, before the packet leaves it.
The detail that mattered to me later is that the choice happens once per connection. A client that opens a connection and keeps it open, as HTTP/2 and gRPC clients do, sends every request down it to the same pod. With a few long-lived clients, load can end up badly uneven, even though the Service is “balancing” traffic.
Traffic from outside the cluster comes in through a Service of type LoadBalancer, which asks the cloud provider for a load balancer, or through an Ingress, which routes HTTP by host and path to Services.
Where a service mesh comes in
Any discussion of Kubernetes networking eventually mentions Istio, so I read up on what a service mesh adds.
A mesh puts a small proxy, usually Envoy, next to every pod as an extra container, called a sidecar. Rules inside the pod send all its traffic in and out through that proxy, so the application talks plain HTTP and the proxy handles the rest. The mesh’s control plane configures all the proxies.
Because the proxy understands HTTP and gRPC, it can do things kube-proxy can’t:
- balance per request instead of per connection, which fixes the gRPC problem above;
- retry failed requests and enforce timeouts;
- encrypt traffic between pods with mutual TLS, without any change to application code;
- split traffic by percentage, for canary releases;
- report latency and error rates for every call between services.
The price is an extra proxy on every hop, more memory per pod, and one more system to understand when something breaks. I haven’t used a mesh myself yet. Reading about one still helped me place kube-proxy: it works on connections, and a mesh works on requests.
What made it click
- Treat YAML as a description of the end state, not a list of steps.
- Helm writes the YAML; Kubernetes still does the work. When in doubt, render the chart and read what it produced.
- When something is wrong, read the object’s status and events with
kubectl describebefore the logs. They tell you which controller gave up, and why. - Components coordinate through the API server, never directly.
- A Service is a set of rules on every node, not a box in the middle.
- Load is balanced per connection, unless something above kube-proxy does better.
Kubernetes is still big. But it’s one idea applied many times, and seeing that made the rest learnable.
References
- Kubernetes documentation. Controllers. kubernetes.io. https://kubernetes.io/docs/concepts/architecture/controller/
- Kubernetes documentation. Kubernetes Components. kubernetes.io. https://kubernetes.io/docs/concepts/overview/components/
- Kubernetes documentation. Service. kubernetes.io. https://kubernetes.io/docs/concepts/services-networking/service/
- Kubernetes documentation. Namespaces. kubernetes.io. https://kubernetes.io/docs/concepts/overview/working-with-objects/namespaces/
- Kubernetes documentation. Network Policies. kubernetes.io. https://kubernetes.io/docs/concepts/services-networking/network-policies/
- Helm. Charts. helm.sh. https://helm.sh/docs/topics/charts/
- Kubernetes documentation. Ingress. kubernetes.io. https://kubernetes.io/docs/concepts/services-networking/ingress/
- Istio. Architecture. istio.io. https://istio.io/latest/docs/ops/deployment/architecture/ — sidecar proxies and the mesh control plane.