---
title: "Chapter 8: Observability Cost Governance, Log Sampling & Metric Explosion — Cloud Economics | TinyCTO"
description: "Preventing metric label explosion, in-flight OpenTelemetry Collector log filtering, and CloudWatch retention limits."
image: "https://tinycto.tv/assets/cloud-economics/cloud_economics_manuals_og.jpg"
canonicalUrl: "https://tinycto.tv/cloud-economics/manuals/08-observability-cost-control"
locale: "en"
---

# Chapter 8: Observability Cost Governance, Log Sampling & Metric Explosion

## 1. Executive Summary & Problem Statement
Observability bills (Datadog, AWS CloudWatch, Splunk, New Relic) have become notorious for exceeding the underlying infrastructure costs they monitor. It is not uncommon for a company spending \$50,000/month on AWS compute to spend \$70,000/month on Datadog or CloudWatch log ingestion and custom metrics.

The primary culprits are:
1. **Unsampled debug/info logs** emitted in high-throughput production services.
2. **High-cardinality metric labels** (injecting user IDs, UUIDs, or timestamps into Prometheus/Datadog metric tags).
3. **Ingesting repetitive healthcheck logs** from load balancers.

---

## 2. The Custom Metric Explosion Trap
In telemetry platforms, metric cost is calculated as:
$$\text{Cost} \propto \prod \text{Cardinality of Dimensions}$$

If a metric `http_requests_total` has:
- `endpoint`: 20 values
- `status_code`: 5 values
- `user_id`: 1,000,000 unique users

That single metric creates $20 \times 5 \times 1,000,000 = 100,000,000$ unique time series, immediately triggering six-figure monthly observability charges.

### Metric Cardinality Golden Rule
**NEVER put high-cardinality values (User IDs, Session IDs, Order IDs, IP Addresses, UUIDs) into metric dimensions.** Put them in structured distributed traces or log payloads instead.

---

## 3. OpenTelemetry Collector In-Flight Log Filtering
Deploy an OpenTelemetry Collector daemonset to filter, sample, and drop unnecessary logs at the edge before they leave the cluster:

```yaml
# OpenTelemetry Collector Pipeline Configuration
processors:
  filter/drop_healthchecks:
    error_mode: ignore
    logs:
      log_record:
        - 'attributes["http.target"] == "/healthz"'
        - 'attributes["http.target"] == "/readyz"'
        - 'attributes["user_agent"] == "ELB-HealthChecker/2.0"'

  probabilistic_sampler:
    sampling_percentage: 10.0 # Keep only 10% of HTTP 200 INFO logs; retain 100% of 4xx/5xx errors

service:
  pipelines:
    logs:
      receivers: [otlp]
      processors: [filter/drop_healthchecks, probabilistic_sampler]
      exporters: [datadog, cloudwatch]
```

---

## 4. CloudWatch Retention Rules
By default, AWS CloudWatch log groups retain log streams **Never Expire** (forever).
Enforce a 30-day retention ceiling across all account log groups via AWS CLI or Terraform:

```bash
# Set 30-day retention on all CloudWatch log groups
aws logs describe-log-groups --query "logGroups[*].logGroupName" --output text | tr '\t' '\n' | while read group; do
  aws logs put-retention-policy --log-group-name "$group" --retention-in-days 30
done
```

```json
{
  "@context": "https://schema.org",
  "@type": "TechArticle",
  "headline": "Chapter 8: Observability Cost Governance, Log Sampling & Metric Explosion",
  "description": "Preventing metric label explosion, in-flight OpenTelemetry Collector log filtering, and CloudWatch retention limits.",
  "url": "https://tinycto.tv/cloud-economics/manuals/08-observability-cost-control",
  "inLanguage": "en-US",
  "author": {
    "@type": "Organization",
    "name": "TinyCTO.tv",
    "url": "https://tinycto.tv"
  },
  "publisher": {
    "@type": "Organization",
    "name": "TinyCTO.tv",
    "url": "https://tinycto.tv"
  }
}
```
