Hand someone a report that says “this node is underutilised, resize it to something smaller” and, on a cluster running Karpenter, you have handed them an instruction they cannot carry out. Delete the node and a replacement appears within a minute. Change its instance type and the next provisioning decision ignores you. The node is an output, not an input.
What you actually control
The autoscaler takes pod requests and a set of constraints and produces capacity. So the levers are the requests, and the constraints — never the result. Cut the requests and consolidation removes the node for you, without anybody filing a ticket. Widen the instance requirements and the scheduler finds cheaper shapes on its own. Narrow them and it cannot, no matter what the bill says.
This is also why node-level and workload-level findings are so often the same money, which we wrote about separately. The node exists because of the requests; pricing both and adding them up counts one saving twice.
The disruption reasons are not interchangeable
Consolidation is not one behaviour. Karpenter disrupts nodes for distinct reasons, and conflating them is how a cost fix becomes an availability incident:
- Empty — the node has no workload pods left. Cheap, safe, and almost never the one causing you trouble.
- Underutilized — the workloads would fit elsewhere, so the node is replaced or removed. This is the frequent one, and the one that evicts pods that were perfectly happy.
- Drifted — the node no longer matches its NodePool definition. Harmless in principle and fleet-wide in practice, because one changed field marks every node at once.
How a one-line field rolls an entire fleet
An AMI selector written as an alias pinned to latest resolves to whatever the vendor published most recently. The day a new EKS-optimised image appears, every node in the cluster is drifted simultaneously and the autoscaler begins replacing all of them. Nothing in your repository changed. Nobody deployed. The schedule was set by someone else's release process.
On the incident we published, that was the trigger, and on its own it would have been a slow afternoon of node replacements. What made it expensive was what it compounded with: consolidation reclaiming underutilised nodes every ten minutes, and an autoscaler with no scaling behaviour configured oscillating between four and twenty-four replicas. Three mechanisms, each of which evicts pods, and a workload that read 1.4 GB over the network every time a pod started.
The resulting bill arrived under a storage line item for a filesystem holding 8.57 GiB. The cause was two layers away in node provisioning.
What to check on a cluster you have just met
- 01Are any AMI or image selectors tracking a floating tag? If so, fleet replacement is on the vendor's schedule, not yours.
- 02Is anything else already mutating pod requests — a VPA, a vendor admission webhook? If so, your measurements describe its decisions rather than the team's.
- 03Do the workloads most sensitive to restarts have a disruption budget that can actually be satisfied? A single replica with minAvailable of one blocks a drain forever.
- 04Does any workload pay a per-start cost — a large volume mount, a slow warm-up, a cold cache? Multiply it by the real eviction rate before deciding churn is harmless.
- 05Are the instance requirements wide enough for the scheduler to find a cheaper shape, or has someone pinned a family and then asked why the bill is high?
The general version
Advice is only advice if its reader can execute it. On an autoscaled cluster that rules out most node-level recommendations, and it rules them out whether or not the arithmetic behind them was correct. The useful form is always the same: name the input the team controls, say what the autoscaler will do in response, and price the result once.
The engagements behind this
- A storage bill grew 20× on a filesystem holding 8 GB
E-commerce · production EKS on AWS · under NDA
- Making a savings number smaller
E-commerce · two EKS clusters on AWS · under NDA