Skip to content

· 6 min read

Why we count a saving once

Cost tools sum every finding they produce. When two detectors reach the same dollars from different directions, the total is wrong — and it is wrong in the direction that flatters whoever produced it.

A cost tool finds things. It finds an idle node, and prices it. It finds a workload requesting four times the CPU it uses, and prices that too. Then it adds them up, because adding up is what a total does, and hands you a number with a currency symbol in front of it.

On a cluster with an autoscaler, those two findings are frequently the same money. Not similar amounts — the same dollars, reached from two ends. The node exists because the requests asked for it. Cut the requests and the node goes away on its own. Claim both and you have claimed the saving twice.

The arithmetic of a number that cannot hold

On the estate behind our FinOps case study, detectors summed to a headline figure. Reconciling the overlaps between them removed roughly three quarters of it. Thirty-three separate idle-node findings turned out to describe capacity that the workload-level findings had already accounted for, and the honest version of those thirty-three was one finding with a range: the measured request overhead at the low end, the node-level estimate at the high end, and a sentence explaining what the gap between them represents.

Three ways the same dollar gets counted twice

  1. 01Node and workload. An autoscaled node's cost and the over-request of the pods on it are one quantity, observed from the infrastructure side and the application side.
  2. 02Resource and usage type. A finding priced from instance hours and a finding priced from a billing usage type can cover the same spend, because the bill partitions it one way and the inventory partitions it another.
  3. 03Commitment and consumption. Cutting usage you have already paid for through a Reserved Instance or a Savings Plan saves nothing this year. The saving is real only once the commitment expires, and a tool that does not know your coverage cannot tell the difference.

What we do instead

Every finding is attributed to exactly one pool of money, by exact match rather than by pattern. A finding that cannot be attributed is listed as unattributed rather than quietly folded into a total — the list of things we could not place is part of the deliverable, not an embarrassment to be hidden.

Where two detectors reach the same spend, they are folded into one finding and the overlap is written down. Where a recommendation cannot be executed by the person reading it — resizing a node an autoscaler owns is the recurring example — it is reclassified as evidence rather than counted as an independent saving, and rewritten as the action that does exist.

And every figure traces to a line on the bill for a named month, so the reader can check it without us. That constraint does more work than any of the others: a number that has to reconcile to a bill line cannot quietly describe a four-month total under a monthly label, which is a mistake we have found in our own detectors and fixed at the source.

Why the smaller number is worth more

A large savings figure is easy to produce and expensive to defend. The moment someone senior asks an engineer to check it, every double-count becomes a reason to distrust the whole exercise — including the findings that were real. The figure that survives that meeting is the only one that gets acted on, and it is usually a good deal smaller than the one the tool printed.

The engagements behind this