

Something strange has happened in the observability market. The tools that were supposed to help teams reduce the cost of downtime have themselves become one of the largest line items in the infrastructure budget. I have spoken to teams paying six and seven figures annually for their observability stack, more than the estimated business impact of the downtime they are trying to prevent. When the medicine costs more than the disease, something has gone fundamentally wrong with the treatment plan.
The path here is easy to trace. It starts with good intentions. A team instruments their services with metrics, traces, and logs. The data is useful. They instrument more. They add custom dashboards. They enable distributed tracing across every service boundary. They configure high-resolution metrics because “what if we need the granularity during an incident?” Each decision is individually reasonable. Collectively, they produce a telemetry firehose that costs a fortune to ingest, store, and query, and most of which is never looked at by anyone.
The observability vendors are not incentivised to help you solve this. Their pricing models are based on data volume. More data means more revenue. So the product encourages you to instrument everything, retain everything, and correlate everything. The marketing says “full observability.” The invoice says you are paying per gigabyte for logs that nobody reads, traces that nobody queries, and metrics that feed dashboards nobody watches.
The fix is to treat observability like any other engineering investment: with a cost-benefit analysis. Start by identifying what you actually look at during incidents. Not what you could theoretically look at. What you actually look at. For most teams, that is a surprisingly small subset of the data they collect. Then work backwards from there. What telemetry directly supports incident detection, diagnosis, and resolution? Keep that at full resolution. Everything else can be sampled, aggregated, or retained at lower resolution with shorter TTLs.
OpenTelemetry helps here by decoupling instrumentation from vendors, but it does not solve the underlying problem of over-collection. That requires engineering discipline and, frankly, a willingness to accept that “less data, better curated” is a more mature observability strategy than “collect everything and hope someone finds it useful.” Your observability stack should make incidents cheaper to resolve, not more expensive to prevent than they are to endure.