Skip to content

Gambling · managed Kubernetes on AWS · under NDA

Three EKS versions upgraded without a maintenance window

Falling behind on Kubernetes is not a single decision. It is the absence of one, repeated every quarter. By the time a cluster is two versions past its support window the upgrade has stopped being routine and become a project — and the longer it waits, the more of the work is archaeology rather than engineering.

The situation

A current cluster upgrades in an afternoon. A late one does not, for three reasons that compound. The control plane moves one minor version at a time, so three releases means three separate migrations rather than one. Each of those releases removes APIs, and the workloads still calling them have usually been untouched for exactly as long as the cluster has. And the add-ons — CNI, CoreDNS, kube-proxy — carry their own compatibility range against the control plane, so they move in lockstep whether or not anyone planned for it.

What also changes is the cost of being wrong. A managed control plane cannot be rolled back to a previous minor version. There is no undo, which means the entire safety argument has to be made before the first upgrade starts rather than assembled after one fails.

Solution

01Inventory what is really running, not what git says

Every API version actually in use was inventoried against the removal list for each target release — not from the manifests in the repository, which describe an intention, but from what the API server was really being asked for. Manifests drift; live usage does not. That produced a short list of workloads to fix and, more usefully, a much longer list of things that were already fine and did not need touching.

02Fix the disruption budgets before the first drain

A drain waits indefinitely for a workload whose PodDisruptionBudget can never be satisfied — a single replica with minAvailable of one is the classic — and a late cluster always has a few. Those were found and corrected before any node was asked to go away, because discovering them mid-rollout is how an unattended upgrade becomes a night shift.

03Move the control plane alone, then replace the nodes

Control plane first and on its own, one minor version at a time, with the add-ons moved to a compatible range at each step. Then node groups replaced rather than upgraded in place, with surge capacity so the new nodes existed before the old ones were drained.

A maintenance window is what you schedule when you cannot predict the blast radius. Here the blast radius was made small enough not to need one: at no point was available capacity below what the workloads required.

Outcome

3
minor versions, one at a time
0
maintenance windows scheduled
0
rollbacks needed

Three minor versions, back inside the support window, with no scheduled downtime and no rollback needed at any step. None of it was clever — it was the documented path, executed in the documented order, with the checks done beforehand instead of discovered during. The reason late upgrades usually need a window is that those checks get skipped.

Core tech

  • Amazon EKS
  • Kubernetes
  • Karpenter
  • Helm
  • Terraform
  • Prometheus

Method

What makes this routine rather than risky

Nothing here is proprietary. This is the part that transfers to your estate whether or not you ever call us.

Live API usage, not the manifests

Deprecated-API findings come from what the API server is actually serving. The repository says what someone intended at the time; the cluster says what is true now.

One minor version at a time

Skipping is unsupported and the failure mode is not graceful. Three releases is three upgrades, each with its own verification step.

Add-ons move with the control plane

CNI, CoreDNS and kube-proxy each have a supported range per version. Leaving them behind produces networking faults that look like anything except an upgrade problem.

Disruption budgets audited in advance

An unsatisfiable PDB turns an automated rollout into a stuck one. Cheap to check beforehand and expensive to meet at two in the morning.

Replace nodes, do not upgrade them

New nodes on the new version, workloads moved across, old nodes removed. With surge capacity the cluster never dips below what it needs — which is what removes the need for a window.

Rehearse with the same inputs

A rehearsal against different manifests proves nothing. The value is in running the real sequence somewhere the consequences are cheap, which matters most when the real thing has no rollback.

The method on this page is standard practice and checkable. The specific figures are illustrative pending the engagement record.