The situation
A current cluster upgrades in an afternoon. A late one does not, for three reasons that compound. The control plane moves one minor version at a time, so three releases means three separate migrations rather than one. Each of those releases removes APIs, and the workloads still calling them have usually been untouched for exactly as long as the cluster has. And the add-ons — CNI, CoreDNS, kube-proxy — carry their own compatibility range against the control plane, so they move in lockstep whether or not anyone planned for it.
What also changes is the cost of being wrong. A managed control plane cannot be rolled back to a previous minor version. There is no undo, which means the entire safety argument has to be made before the first upgrade starts rather than assembled after one fails.
Solution
01Inventory what is really running, not what git says
Every API version actually in use was inventoried against the removal list for each target release — not from the manifests in the repository, which describe an intention, but from what the API server was really being asked for. Manifests drift; live usage does not. That produced a short list of workloads to fix and, more usefully, a much longer list of things that were already fine and did not need touching.
02Fix the disruption budgets before the first drain
A drain waits indefinitely for a workload whose PodDisruptionBudget can never be satisfied — a single replica with minAvailable of one is the classic — and a late cluster always has a few. Those were found and corrected before any node was asked to go away, because discovering them mid-rollout is how an unattended upgrade becomes a night shift.
03Move the control plane alone, then replace the nodes
Control plane first and on its own, one minor version at a time, with the add-ons moved to a compatible range at each step. Then node groups replaced rather than upgraded in place, with surge capacity so the new nodes existed before the old ones were drained.
A maintenance window is what you schedule when you cannot predict the blast radius. Here the blast radius was made small enough not to need one: at no point was available capacity below what the workloads required.
Outcome
- 3
- minor versions, one at a time
- 0
- maintenance windows scheduled
- 0
- rollbacks needed
Three minor versions, back inside the support window, with no scheduled downtime and no rollback needed at any step. None of it was clever — it was the documented path, executed in the documented order, with the checks done beforehand instead of discovered during. The reason late upgrades usually need a window is that those checks get skipped.