Skip to content

E-commerce · production EKS on AWS · under NDA

A storage bill grew 20× on a filesystem holding 8 GB

Daily spend in a production AWS account started climbing: +10% one day, +5% the next. The timing made no sense. That same week the team had shut down a dozen database instances and cleaned up thousands of stale backup recovery points — the bill should have been falling. The question was not “how do we cut costs”. It was “what changed, when nobody changed anything?”

The situation

Amazon EFS went from $6/day to $138/day in three days while every other service moved by single-digit dollars. Reads rose from 150 GB to 2,314 GB per day against flat writes of about 14 GB — on a filesystem holding 8.57 GiB. Something was reading the same few gigabytes roughly three hundred times a day.

EFS gives you no help identifying who. The bill has no per-client breakdown, access points are not billed separately, VPC Flow Logs were not enabled so the network path could not be reconstructed after the fact, and the ClientConnections metric read a constant 1.0.

Solution

01Narrow by billing, not by intuition

Daily cost grouped by service isolated the line item immediately. Grouping that service by usage type narrowed it to a single charge — ETDataAccess-Bytes, the per-gigabyte fee for reading and writing an EFS volume in Elastic throughput mode. One query, before any theory.

02Triangulate the consumer from three independent sources

With no per-client data in the bill, the answer came from Kubernetes state metrics. Counting distinct pods that existed per day, per deployment, produced an unambiguous suspect: one brand-portal deployment went from 15 pods per day to 1,671 over the same three days, while every other workload grew by a factor of two or three. At eight running replicas, that is a pod created roughly every fifty seconds.

Then the arithmetic closed the case: 2,314 GB ÷ 1,671 pods ≈ 1.4 GB per pod start — exactly the size of the Next.js build output, which was mounted from EFS rather than baked into the container image. Every new pod re-read its entire bundle over the network, and the meter ran.

03Find the trigger, which was not in the storage layer at all

Three mechanisms compounded. The trigger sat in the node provisioning config: amiSelectorTerms with alias al2023@latest. AWS published a new EKS-optimised AMI, @latest immediately resolved to it, and all 54 nodes were marked as drifted — so the autoscaler began rolling the entire fleet. On top of that, consolidation was reclaiming underutilised nodes every 10 minutes, and the portal's autoscaler had no scaling behaviour configured, so it oscillated freely between 4 and 24 replicas.

Each of those mechanisms evicts pods. Each eviction cost 1.4 GB of reads.

04Fix to the constraint, not to the ideal

Rather than rebuild the application's storage layout under time pressure, we gated the most frequent disruption source: a budget blocking Underutilized-reason consolidation for 21 hours a day, leaving a three-hour window aligned to the portal's own quietest hour.

That window was chosen from measured data. The portal's load profile bottoms out at a different hour than the databases on the same cluster, and copying the database pool's night window would have been the worse choice.

Outcome

$138
peak daily spend, on an 8.57 GiB filesystem
270×
read volume against the filesystem's own size
6
lines added to a node pool definition
Time to isolate the service
One query against billing data
Time to identify the workload
One metrics query, confirmed by arithmetic
Pod lifetime, before → after
4–8 minutes → 10–125 minutes
Verification
Scheduler reporting DisruptionBlocked within minutes of rollout

The change shipped through the existing GitOps pipeline. Its effect was visible in cluster events immediately and in pod lifetimes within the hour.

We also handed over the two structural fixes the incident exposed but that were out of scope for a same-day response: pin the AMI instead of tracking @latest, so fleet-wide replacements become a planned activity rather than a surprise; and move the application bundle into the container image, so pod churn stops carrying a per-gigabyte price tag at all.

Core tech

  • AWS Cost Explorer
  • CloudWatch
  • Amazon EFS
  • Amazon EKS
  • Karpenter
  • Horizontal Pod Autoscaler
  • kube-state-metrics
  • VictoriaMetrics
  • Argo CD

Method

What transfers to other environments

Nothing here is proprietary. This is the part that transfers to your estate whether or not you ever call us.

@latest is a subscription to unplanned work

It behaves perfectly until an upstream release turns it into a fleet-wide rolling replacement, on the vendor's schedule rather than yours.

Costs rise in services you did not touch

The billing line said EFS. The cause was in node provisioning, two layers away. Following the bill to the service and stopping there would have found nothing to fix.

When the service will not say who, triangulate

Billing gives you the what, infrastructure metrics give you the when, workload metrics give you the who — and dividing one by the other is often the proof that ties them together.

Pick the maintenance window from the workload's own profile

The portal's quietest hour was not the database pool's. Reusing the existing window would have gated consolidation at exactly the wrong time.