Kubernetes management
Day-2 operation of EKS, GKE and self-managed clusters: upgrades, node lifecycle, autoscaling, capacity and the boring reliability work that decides whether a platform is trusted.
Day one was the easy part
A cluster is straightforward to create and expensive to keep. Version support windows expire on someone else's schedule, node pools drift away from what the workloads actually need, autoscalers make decisions nobody reviews, and the person who set it up has moved on. Nothing is broken, exactly — it is just that no one can say what would happen if a zone went away.
What we do
How the work goes
Inventory
Two weeks of measurement before any change: what runs, what it requests, what it actually uses, and where the cluster has no headroom left.
Stabilise
The findings that carry risk go first — version end-of-life, single points of failure, workloads with no disruption budget, nodes that cannot survive a zone loss.
Operate
A standing arrangement: upgrades, capacity reviews and an on-call path, with a monthly written summary of what changed and what it cost.
What you get
- A written cluster inventory: versions, node groups, workloads, and the gaps with severities
- An upgrade plan with dates and rollback points
- Monthly report of changes, incidents and capacity trend
A good fit when
- Teams running production Kubernetes without a dedicated platform person
- Clusters that grew organically and now nobody wants to touch
- Companies whose auditor has started asking about upgrade policy
Not us, and we will say so
- Clusters we are not allowed to observe — we do not operate what we cannot measure
- Teams looking for someone to take the blame rather than the work
Tell us what you are running.
A first call is a conversation. If this is not the service you need, we will point you at the one that is — or tell you that you do not need us yet.
Book a call