Skip to main content

> FINOPS // CHAPTER 08

Chapter 8: Observability Cost Governance, Log Sampling & Metric Explosion

Preventing metric label explosion, in-flight OpenTelemetry Collector log filtering, and CloudWatch retention limits.

Canonical FinOps Manual #08|TinyCTO Cloud Bill Bible

Chapter 8: Observability Cost Governance, Log Sampling & Metric Explosion

Preventing metric label explosion, in-flight OpenTelemetry Collector log filtering, and CloudWatch retention limits.

#1. Executive Summary & Problem Statement

Observability bills (Datadog, AWS CloudWatch, Splunk, New Relic) have become notorious for exceeding the underlying infrastructure costs they monitor. It is not uncommon for a company spending `math:50,000/month on AWS compute to spend `70,000/month on Datadog or CloudWatch log ingestion and custom metrics.

The primary culprits are:

  1. Unsampled debug/info logs emitted in high-throughput production services.
  2. High-cardinality metric labels (injecting user IDs, UUIDs, or timestamps into Prometheus/Datadog metric tags).
  3. Ingesting repetitive healthcheck logs from load balancers.

#2. The Custom Metric Explosion Trap

In telemetry platforms, metric cost is calculated as:

Cost∝∏Cardinality of Dimensions\text{Cost} \propto \prod \text{Cardinality of Dimensions}

If a metric http_requests_total has:

  • endpoint: 20 values
  • status_code: 5 values
  • user_id: 1,000,000 unique users

That single metric creates 20×5×1,000,000=100,000,00020 \times 5 \times 1,000,000 = 100,000,000 unique time series, immediately triggering six-figure monthly observability charges.

Metric Cardinality Golden Rule

NEVER put high-cardinality values (User IDs, Session IDs, Order IDs, IP Addresses, UUIDs) into metric dimensions. Put them in structured distributed traces or log payloads instead.


#3. OpenTelemetry Collector In-Flight Log Filtering

Deploy an OpenTelemetry Collector daemonset to filter, sample, and drop unnecessary logs at the edge before they leave the cluster:

# OpenTelemetry Collector Pipeline Configuration
processors:
  filter/drop_healthchecks:
    error_mode: ignore
    logs:
      log_record:
        - 'attributes["http.target"] == "/healthz"'
        - 'attributes["http.target"] == "/readyz"'
        - 'attributes["user_agent"] == "ELB-HealthChecker/2.0"'

  probabilistic_sampler:
    sampling_percentage: 10.0 # Keep only 10% of HTTP 200 INFO logs; retain 100% of 4xx/5xx errors

service:
  pipelines:
    logs:
      receivers: [otlp]
      processors: [filter/drop_healthchecks, probabilistic_sampler]
      exporters: [datadog, cloudwatch]

#4. CloudWatch Retention Rules

By default, AWS CloudWatch log groups retain log streams Never Expire (forever). Enforce a 30-day retention ceiling across all account log groups via AWS CLI or Terraform:

# Set 30-day retention on all CloudWatch log groups
aws logs describe-log-groups --query "logGroups[*].logGroupName" --output text | tr '\t' '\n' | while read group; do
  aws logs put-retention-policy --log-group-name "$group" --retention-in-days 30
done
AI Summary — Chapter 08: Chapter 8: Observability Cost Governance, Log Sampling & Metric Explosion
AEO / GEO / Perplexity Indexable

Preventing metric label explosion, in-flight OpenTelemetry Collector log filtering, and CloudWatch retention limits.

Chapter ScopeChapter 08 canonical FinOps principles and unit cost guardrails.
Core ConceptsMetric Cardinality • OpenTelemetry Filtering • Probabilistic Sampling • CloudWatch Retention
Maturity LevelWALK (Intermediate)
Agent GuardrailEnforce FOCUS 1.0 mandatory tagging schema and automated anomaly gate remediation.