What if we treated cost as another observability signal?
Most organizations already have pretty good visibility into their cloud spend.
AWS, Azure and GCP provide detailed billing data. FinOps teams have dashboards showing spend by account, service, business unit and sometimes application.
But there is a difference between knowing what you spent and understanding why your cost changed.
For Kubernetes, knowing that an AWS account spent $80,000 on compute last month doesn't tell an engineer very much.
You need to be able to move through the infrastructure:
Cloud Account → Cluster → Namespace → Workload → Pod → Container
And then connect that infrastructure back to the organization:
Team → Application → Product → Cost Center
Now imagine Kubernetes spend increases 25% over two weeks.
Knowing that the increase came from three clusters is useful.
Knowing that most of it came from two workloads is much more useful.
Knowing that one of those workloads doubled its CPU requests after a deployment gives an engineer somewhere to start investigating.
This is where I think cost starts moving beyond reporting and becomes another form of operational telemetry.
Applications aren't static.
We deploy new versions. Traffic patterns change. Resource requirements change. Someone changes a request or limit. A workload starts scaling differently. Performance improves or degrades.
So why do we treat cost as something we look at separately, often days or weeks later?
I think cost metrics should sit alongside the other signals we already use to understand an application.
Latency went down. What happened to cost?
We increased CPU requests. Did performance actually improve?
Traffic increased 20%. Did infrastructure cost increase 20%, 5% or 50%?
A deployment improved p95 latency by 15%, but doubled the cost of running the service. Was that the tradeoff we intended to make?
That's the feedback loop I'm interested in.
Not optimizing cost in isolation. And definitely not reducing cost at the expense of performance or reliability.
It's understanding cost, performance and reliability together as the system changes.
If observability is supposed to help us understand the behavior of our systems, cost should be part of that picture.
#CostManagement #Kubernetes #Observability