The situation
Amazon EFS went from $6/day to $138/day in three days while every other service moved by single-digit dollars. Reads rose from 150 GB to 2,314 GB per day against flat writes of about 14 GB — on a filesystem holding 8.57 GiB. Something was reading the same few gigabytes roughly three hundred times a day.
EFS gives you no help identifying who. The bill has no per-client breakdown, access points are not billed separately, VPC Flow Logs were not enabled so the network path could not be reconstructed after the fact, and the ClientConnections metric read a constant 1.0.
Solution
01Narrow by billing, not by intuition
Daily cost grouped by service isolated the line item immediately. Grouping that service by usage type narrowed it to a single charge — ETDataAccess-Bytes, the per-gigabyte fee for reading and writing an EFS volume in Elastic throughput mode. One query, before any theory.
02Triangulate the consumer from three independent sources
With no per-client data in the bill, the answer came from Kubernetes state metrics. Counting distinct pods that existed per day, per deployment, produced an unambiguous suspect: one brand-portal deployment went from 15 pods per day to 1,671 over the same three days, while every other workload grew by a factor of two or three. At eight running replicas, that is a pod created roughly every fifty seconds.
Then the arithmetic closed the case: 2,314 GB ÷ 1,671 pods ≈ 1.4 GB per pod start — exactly the size of the Next.js build output, which was mounted from EFS rather than baked into the container image. Every new pod re-read its entire bundle over the network, and the meter ran.
03Find the trigger, which was not in the storage layer at all
Three mechanisms compounded. The trigger sat in the node provisioning config: amiSelectorTerms with alias al2023@latest. AWS published a new EKS-optimised AMI, @latest immediately resolved to it, and all 54 nodes were marked as drifted — so the autoscaler began rolling the entire fleet. On top of that, consolidation was reclaiming underutilised nodes every 10 minutes, and the portal's autoscaler had no scaling behaviour configured, so it oscillated freely between 4 and 24 replicas.
Each of those mechanisms evicts pods. Each eviction cost 1.4 GB of reads.
04Fix to the constraint, not to the ideal
Rather than rebuild the application's storage layout under time pressure, we gated the most frequent disruption source: a budget blocking Underutilized-reason consolidation for 21 hours a day, leaving a three-hour window aligned to the portal's own quietest hour.
That window was chosen from measured data. The portal's load profile bottoms out at a different hour than the databases on the same cluster, and copying the database pool's night window would have been the worse choice.
Outcome
- $138
- peak daily spend, on an 8.57 GiB filesystem
- 270×
- read volume against the filesystem's own size
- 6
- lines added to a node pool definition
- Time to isolate the service
- One query against billing data
- Time to identify the workload
- One metrics query, confirmed by arithmetic
- Pod lifetime, before → after
- 4–8 minutes → 10–125 minutes
- Verification
- Scheduler reporting DisruptionBlocked within minutes of rollout
The change shipped through the existing GitOps pipeline. Its effect was visible in cluster events immediately and in pod lifetimes within the hour.
We also handed over the two structural fixes the incident exposed but that were out of scope for a same-day response: pin the AMI instead of tracking @latest, so fleet-wide replacements become a planned activity rather than a surprise; and move the application bundle into the container image, so pod churn stops carrying a per-gigabyte price tag at all.