Chapter 8: Observability Cost Governance, Log Sampling & Metric Explosion
Preventing metric label explosion, in-flight OpenTelemetry Collector log filtering, and CloudWatch retention limits.
#1. Executive Summary & Problem Statement
Observability bills (Datadog, AWS CloudWatch, Splunk, New Relic) have become notorious for exceeding the underlying infrastructure costs they monitor. It is not uncommon for a company spending `math:50,000/month on AWS compute to spend `70,000/month on Datadog or CloudWatch log ingestion and custom metrics.
The primary culprits are:
- Unsampled debug/info logs emitted in high-throughput production services.
- High-cardinality metric labels (injecting user IDs, UUIDs, or timestamps into Prometheus/Datadog metric tags).
- Ingesting repetitive healthcheck logs from load balancers.
#2. The Custom Metric Explosion Trap
In telemetry platforms, metric cost is calculated as:
If a metric http_requests_total has:
endpoint: 20 valuesstatus_code: 5 valuesuser_id: 1,000,000 unique users
That single metric creates unique time series, immediately triggering six-figure monthly observability charges.
Metric Cardinality Golden Rule
NEVER put high-cardinality values (User IDs, Session IDs, Order IDs, IP Addresses, UUIDs) into metric dimensions. Put them in structured distributed traces or log payloads instead.
#3. OpenTelemetry Collector In-Flight Log Filtering
Deploy an OpenTelemetry Collector daemonset to filter, sample, and drop unnecessary logs at the edge before they leave the cluster:
# OpenTelemetry Collector Pipeline Configuration
processors:
filter/drop_healthchecks:
error_mode: ignore
logs:
log_record:
- 'attributes["http.target"] == "/healthz"'
- 'attributes["http.target"] == "/readyz"'
- 'attributes["user_agent"] == "ELB-HealthChecker/2.0"'
probabilistic_sampler:
sampling_percentage: 10.0 # Keep only 10% of HTTP 200 INFO logs; retain 100% of 4xx/5xx errors
service:
pipelines:
logs:
receivers: [otlp]
processors: [filter/drop_healthchecks, probabilistic_sampler]
exporters: [datadog, cloudwatch]
#4. CloudWatch Retention Rules
By default, AWS CloudWatch log groups retain log streams Never Expire (forever). Enforce a 30-day retention ceiling across all account log groups via AWS CLI or Terraform:
# Set 30-day retention on all CloudWatch log groups
aws logs describe-log-groups --query "logGroups[*].logGroupName" --output text | tr '\t' '\n' | while read group; do
aws logs put-retention-policy --log-group-name "$group" --retention-in-days 30
done
