> STREAMING
Serverless Data Lakehouse on Apache Iceberg & Athena
Zero-cluster data lakehouse architecture using Apache Iceberg open table format, AWS S3 object storage, and serverless query engines (Athena/DuckDB).
Mathematical Breakeven Inflection Curve
Cost-optimal for batch analytics and ad-hoc BI queries (< 5,000 queries/day). Beyond 20,000 queries/day, dedicated ClickHouse or StarRocks is 60% cheaper.
3 Maturity Tiers & Infrastructure Specifications
Component stack and cost steps from prototype to hyper-scale enterprise
1. Prototype / Early Stage100 GB - 1 TB stored data
$25 - $150 / mo
$0.005 / TB scanned
Stack Components:
- S3 Standard Storage
- Apache Iceberg (Parquet formatted)
- AWS Athena (Serverless SQL)
Cost Allocation:Athena Workgroup Tagging
Autoscaling:Native Athena Serverless Capacity
2. Scaled Production5 TB - 50 TB stored data
$250 - $1,800 / mo
$0.0035 / TB scanned (Partition pruned)
Stack Components:
- S3 Intelligent-Tiering
- Iceberg with compaction lambdas
- Athena Workgroups with per-query limits
Cost Allocation:Cost Allocation per Athena Workgroup & Team
Autoscaling:Serverless Auto-Partitioning
3. High-Throughput Enterprise100 TB - 1 PB stored data
$1,800 - $9,500 / mo
$0.002 / TB scanned
Stack Components:
- S3 Express One Zone (hot tier) + Glacier (cold)
- Apache Iceberg with DuckDB cache
- Athena Provisioned Capacity
Cost Allocation:Enterprise Data Mesh Cost Allocation
Autoscaling:Athena Reserved DPU Pools
Key Cost Drivers
- •Athena Data Scanned ($5.00/TB scanned)
- •S3 GET/PUT API request calls
- •S3 Storage volume
Waste Vulnerabilities
- •Running SELECT * queries without partition filtering scanning full petabytes
- •Small file problem causing excessive S3 API call costs
Mitigation Playbooks
- •Enforce Iceberg automatic file compaction
- •Set Athena workgroup data scan limits (e.g. max 10GB per query)
- •Convert all tables to Parquet with Snappy/ZSTD
