We ask for a read-only role in the billing account and, optionally, a metrics agent in the cluster. We never ask for administrator credentials. That is a constraint we chose, so it is fair to ask what it costs — what a reviewer with no write access genuinely cannot find out.
What the bill will tell you
Cost Explorer answers the question “what changed, and when” better than anything else available, and it answers it in one query. Daily granularity grouped by service isolates a spike to a line item; grouping that service by usage type usually narrows it to a single charge. On the EFS incident we wrote up, that took the investigation from “spend is rising” to “ETDataAccess-Bytes, the per-gigabyte read fee” without touching the cluster at all.
- Daily cost by service, by usage type, by linked account — all read-only, all immediate.
- Resource-level cost where the service reports it, which is not everywhere.
- Cost allocation tags, but only the ones activated in the billing console. A tag that exists on the resource and was never activated does not appear, and nothing warns you.
- A Cost and Usage Report, if one is already being delivered. Enabling one is a write, and the first file can take up to 24 hours — which is why we ask about it before the engagement rather than during.
What the bill will not tell you
The gap that matters most is attribution inside a shared service. A managed filesystem bills you for reads; it does not bill you per client. EFS has no per-consumer breakdown, access points are not billed separately, and the ClientConnections metric can sit at a constant 1.0 while thousands of pods come and go. If VPC Flow Logs were not already enabled, the network path cannot be reconstructed after the fact — and enabling them is both a write and a change that only helps from that moment forward.
What closes the gap
Workload metrics, which is why we ask for the agent. On that same incident the answer came from counting distinct pods per day per deployment: one workload went from 15 pods a day to 1,671 while everything else grew by a factor of two or three. Dividing the day's read volume by the pod count gave 1.4 GB per pod start — exactly the size of a build output being mounted from the filesystem instead of baked into the image.
None of those three sources could have answered it alone. Billing gave the what, infrastructure metrics gave the when, workload metrics gave the who, and a division tied them together. Triangulation is not a workaround for lacking write access; it is what you do when a managed service will not tell you who consumed it, regardless of your permissions.
The things read-only genuinely cannot do
- Turn on telemetry that was not already on. Flow logs, detailed monitoring and CUR exports all only help from the moment they exist.
- See inside an account that is not in the Organization, or a cluster with no agent. If part of the estate is dark, it is listed as dark.
- Observe a workload that did not run during the measurement window. Fourteen days of a quiet fortnight describes a quiet fortnight.
- Prove a negative. “No findings here” means no findings with the data available, and we say which of the two it is.
Each of those becomes a line in the gap list that ships with the report: what we could not evaluate, and what data would close it. A reviewer who hides that list is not more thorough than one who publishes it — they have just moved the uncertainty somewhere the reader cannot see it.
The engagements behind this
- A storage bill grew 20× on a filesystem holding 8 GB
E-commerce · production EKS on AWS · under NDA
- Making a savings number smaller
E-commerce · two EKS clusters on AWS · under NDA