> STREAMING
High-Throughput Streaming: Kafka, Apache Flink & ClickHouse
Battle-tested real-time analytics and financial transaction event stream utilizing self-hosted Kafka with NVMe storage, Flink CEP, and ClickHouse columnar storage.
Mathematical Breakeven Inflection Curve
Inflection champion above 500 million events/month. Self-hosted ClickHouse is 80% cheaper than Snowflake or Databricks for real-time ingest.
3 Maturity Tiers & Infrastructure Specifications
Component stack and cost steps from prototype to hyper-scale enterprise
1. Prototype / Early Stage10M - 50M events/mo
$350 - $1,100 / mo
$0.022 / 1M events
Stack Components:
- 3-node Kafka Cluster (EC2 i3en.xlarge)
- Flink Standalone on K8s
- Single-node ClickHouse NVMe
Cost Allocation:Topic-level Cost Allocation
Autoscaling:Static Dedicated Compute
2. Scaled Production100M - 1B events/mo
$1,800 - $6,500 / mo
$0.0065 / 1M events
Stack Components:
- 5-node Kafka Cluster (Graviton is4gen NVMe)
- Flink on Kubernetes (Spot instances)
- ClickHouse Cluster (3-node HA with S3 cold tier)
Cost Allocation:FOCUS Data Pipeline Tagging per data stream
Autoscaling:Flink Reactive Mode + Kafka Partition Scaling
3. High-Throughput Enterprise1B - 20B events/mo
$6,500 - $26,000 / mo
$0.0013 / 1M events
Stack Components:
- Kafka on Bare-Metal NVMe / AWS Graviton
- Flink Cluster with RocksDB state backend
- Distributed ClickHouse Cluster on S3 Object Storage
Cost Allocation:Granular Table & Stream Chargeback
Autoscaling:Automated Lag-Based Consumer Autoscaling
Key Cost Drivers
- •NVMe instance storage IOPS & capacity
- •Cross-AZ replication between Kafka brokers
- •ClickHouse memory & CPU compute
Waste Vulnerabilities
- •Retaining raw uncompressed Kafka topic logs for > 7 days on expensive NVMe
- •Redundant data replication across 3 AZs for non-critical logs
Mitigation Playbooks
- •Enable Kafka Tiered Storage offloading to S3 after 12 hours (70% storage savings)
- •Use ZSTD compression on all Kafka topics
