Skip to main content

> STREAMING

Serverless Data Lakehouse on Apache Iceberg & Athena

Zero-cluster data lakehouse architecture using Apache Iceberg open table format, AWS S3 object storage, and serverless query engines (Athena/DuckDB).

Mathematical Breakeven Inflection Curve

Cost-optimal for batch analytics and ad-hoc BI queries (< 5,000 queries/day). Beyond 20,000 queries/day, dedicated ClickHouse or StarRocks is 60% cheaper.

3 Maturity Tiers & Infrastructure Specifications

Component stack and cost steps from prototype to hyper-scale enterprise

1. Prototype / Early Stage100 GB - 1 TB stored data
$25 - $150 / mo
$0.005 / TB scanned
Stack Components:
  • S3 Standard Storage
  • Apache Iceberg (Parquet formatted)
  • AWS Athena (Serverless SQL)
Cost Allocation:Athena Workgroup Tagging
Autoscaling:Native Athena Serverless Capacity
2. Scaled Production5 TB - 50 TB stored data
$250 - $1,800 / mo
$0.0035 / TB scanned (Partition pruned)
Stack Components:
  • S3 Intelligent-Tiering
  • Iceberg with compaction lambdas
  • Athena Workgroups with per-query limits
Cost Allocation:Cost Allocation per Athena Workgroup & Team
Autoscaling:Serverless Auto-Partitioning
3. High-Throughput Enterprise100 TB - 1 PB stored data
$1,800 - $9,500 / mo
$0.002 / TB scanned
Stack Components:
  • S3 Express One Zone (hot tier) + Glacier (cold)
  • Apache Iceberg with DuckDB cache
  • Athena Provisioned Capacity
Cost Allocation:Enterprise Data Mesh Cost Allocation
Autoscaling:Athena Reserved DPU Pools

Key Cost Drivers

  • •Athena Data Scanned ($5.00/TB scanned)
  • •S3 GET/PUT API request calls
  • •S3 Storage volume

Waste Vulnerabilities

  • •Running SELECT * queries without partition filtering scanning full petabytes
  • •Small file problem causing excessive S3 API call costs

Mitigation Playbooks

  • •Enforce Iceberg automatic file compaction
  • •Set Athena workgroup data scan limits (e.g. max 10GB per query)
  • •Convert all tables to Parquet with Snappy/ZSTD