Skip to content

Technology

Kubernetes

We operate production Kubernetes and we measure it. Those two things are the same job: almost every decision about a cluster — what it costs, whether it survives a zone loss, whether a deploy is safe — is a question about numbers nobody is currently collecting.

What we think, and why

The scheduler packs by requests, so requests are the unit that matters.

A pod holds what it booked whether or not it touches it. Utilisation tells you what the application does; requests tell you what the cluster is paying for, and confusing the two produces advice that points at hardware which does not exist.

Some namespaces need more, not less.

Any honest rightsizing exercise finds workloads whose measured peak exceeds what they asked for. Cutting across the board to hit a savings number is how a cost project becomes an incident.

A node you cannot resize is not a rightsizing target.

Where an autoscaler chooses node types from pod requests, advice to downsize the node is not executable — the lever is the requests, and the node follows. Tools that recommend otherwise are describing a cluster nobody runs.

Reliability settings are cheap until the day they are not.

Probes, disruption budgets, topology spread and priority classes cost nothing to set and decide what happens during an upgrade, a spot reclaim or a zone failure. We check them because nobody is thanked for them in advance.

What we do with it

  • Day-2 operation: upgrades, node lifecycle, capacity, incident response
  • Rightsizing based on percentiles over weeks, with the raises as well as the cuts
  • Reliability review: probes, PDBs, spread, priorities, and what a zone loss would actually do
  • Cost per namespace and per workload, computed from one measurement rather than two
All services

Running Kubernetes and want a second opinion?

Book a call