Skip to main content

> ARCHITECTURE CATALOG // 18 MODELS

18 Cloud Architectures & Breakeven Curves

From serverless to bare-metal colocation across 54 maturity tiers with exact unit cost step-functions.

SERVERLESSarch-fin-01

Serverless Event-Driven Web & REST API

Zero-maintenance serverless request tier using API Gateway, Lambda, and DynamoDB On-Demand for bursty traffic with zero idle compute cost.

Breakeven:Cost-optimal below 8 million requests/month (~$120/mo). Above 25 million requests/month, Managed Containers (ECS Fargate) become 40% cheaper.
Prototype:$5 - $45 / mo
Production:$80 - $950 / mo
High-Scale:$950 - $3,800 / mo
SERVERLESSarch-fin-02

Serverless Asynchronous Event & Queue Processing

Decoupled event pipeline using Amazon EventBridge, SQS FIFO queues, and batched Lambda consumers with dead-letter queue resilience.

Breakeven:Cost-optimal up to 50M events/month. Beyond 100M events/month, self-hosted Kafka or managed RabbitMQ on EC2 lowers cost by 65%.
Prototype:$10 - $60 / mo
Production:$60 - $550 / mo
High-Scale:$550 - $2,400 / mo
SERVERLESSarch-fin-03

Edge Serverless Full-Stack Web & Dynamic Routing

Ultra-low-latency edge application platform running on Cloudflare Workers, Pages, and edge KV/D1 SQL databases with zero public egress fees.

Breakeven:Cost-effective across all scales for read-heavy global traffic due to Cloudflare Bandwidth Alliance ($0.00 egress).
Prototype:$5 - $25 / mo
Production:$25 - $220 / mo
High-Scale:$220 - $1,800 / mo
CONTAINERIZEDarch-fin-04

Managed Container Microservices on AWS ECS Fargate

Serverless container execution on AWS ECS Fargate eliminating EC2 cluster management overhead while providing predictable per-second resource billing.

Breakeven:Cost-optimal for workloads with 10 - 80 tasks. Above 100 steady-state tasks, EKS with Karpenter and Spot instances reduces cost by 45%.
Prototype:$85 - $260 / mo
Production:$450 - $2,200 / mo
High-Scale:$2,200 - $9,500 / mo
CONTAINERIZEDarch-fin-05

Google Cloud Run Managed Microservices Platform

Highly elastic managed container platform scaling to zero when idle, billing exclusively during active request processing.

Breakeven:Extremely economical for irregular or daytime-only traffic. For 24/7 continuous high-utilization (>70% CPU), GKE Autopilot becomes 30% cheaper.
Prototype:$15 - $90 / mo
Production:$180 - $1,400 / mo
High-Scale:$1,400 - $6,500 / mo
CONTAINERIZEDarch-fin-06

Azure Container Apps (ACA) Microservices Mesh

Fully managed serverless container environment built on top of Kubernetes and KEDA, providing Dapr microservice capabilities with consumption billing.

Breakeven:Cost-effective up to 60 replicas. Beyond 60 continuous replicas, Azure Kubernetes Service (AKS) with spot node pools is 35% cheaper.
Prototype:$40 - $180 / mo
Production:$250 - $1,600 / mo
High-Scale:$1,600 - $7,200 / mo
KUBERNETESarch-fin-07

High-Density Multi-Tenant AWS EKS with Karpenter & Graviton Spot

State-of-the-art enterprise Kubernetes platform combining Karpenter just-in-time autoscaling, ARM64 Graviton instances, and aggressive Spot node bin-packing.

Breakeven:Inflection champion above 50 vCPUs. Consistently 50% - 65% cheaper than ECS Fargate or standard managed compute across steady-state workloads.
Prototype:$240 - $650 / mo
Production:$850 - $3,800 / mo
High-Scale:$3,800 - $24,000 / mo
KUBERNETESarch-fin-08

Google Kubernetes Engine (GKE) Autopilot Workload Architecture

Hands-off Google Kubernetes Engine cluster where Google manages node infrastructure, billing only for requested pod resources with automated bin-packing.

Breakeven:Cost-optimal for teams without dedicated Kubernetes SRE teams. For teams with SRE capacity, standard GKE with Spot GCE instances is 25% cheaper.
Prototype:$120 - $350 / mo
Production:$600 - $2,900 / mo
High-Scale:$2,900 - $16,000 / mo
KUBERNETESarch-fin-09

Cilium eBPF Hybrid Service Mesh & Private Cloud Topology

High-performance container networking using Cilium eBPF without kube-proxy, eliminating iptables overhead and routing traffic locally to avoid AZ transit tolls.

Breakeven:Reduces network spend by 40% - 70% in high-volume microservice topologies (> 50 TB/mo internal east-west traffic).
Prototype:$280 - $750 / mo
Production:$1,200 - $4,800 / mo
High-Scale:$4,800 - $28,000 / mo
STREAMINGarch-fin-10

High-Throughput Streaming: Kafka, Apache Flink & ClickHouse

Battle-tested real-time analytics and financial transaction event stream utilizing self-hosted Kafka with NVMe storage, Flink CEP, and ClickHouse columnar storage.

Breakeven:Inflection champion above 500 million events/month. Self-hosted ClickHouse is 80% cheaper than Snowflake or Databricks for real-time ingest.
Prototype:$350 - $1,100 / mo
Production:$1,800 - $6,500 / mo
High-Scale:$6,500 - $26,000 / mo
STREAMINGarch-fin-11

Serverless Data Lakehouse on Apache Iceberg & Athena

Zero-cluster data lakehouse architecture using Apache Iceberg open table format, AWS S3 object storage, and serverless query engines (Athena/DuckDB).

Breakeven:Cost-optimal for batch analytics and ad-hoc BI queries (< 5,000 queries/day). Beyond 20,000 queries/day, dedicated ClickHouse or StarRocks is 60% cheaper.
Prototype:$25 - $150 / mo
Production:$250 - $1,800 / mo
High-Scale:$1,800 - $9,500 / mo
STREAMINGarch-fin-12

Change Data Capture (CDC) Pipeline: Debezium to Cloud Data Warehouses

Continuous operational database replication pipeline reading transaction logs (WAL) via Debezium to replicate Postgres/MySQL data into analytical stores.

Breakeven:Cost-effective alternative to costly SaaS ETL tools (Fivetran/Stitch) once table rows exceed 10 million/day (saving $3,000 - $15,000/mo).
Prototype:$60 - $220 / mo
Production:$350 - $1,400 / mo
High-Scale:$1,400 - $5,800 / mo
GENAI_GPUarch-fin-13

Cost-Optimized LLM Inference Cluster: vLLM & Spot GPUs

High-efficiency production LLM serving cluster utilizing vLLM PagedAttention, Ray cluster autoscaling, and spot GPU instances with automated failover.

Breakeven:Inflection champion above 20 million tokens/day. Self-hosting 70B models on 4x A100/H100 Spot instances is 60% cheaper than OpenAI/Anthropic API rates.
Prototype:$450 - $1,800 / mo
Production:$2,800 - $11,500 / mo
High-Scale:$11,500 - $65,000 / mo
GENAI_GPUarch-fin-14

High-Throughput Vector Embedding Pipeline: GPU & CPU AVX-512

Asynchronous text embedding architecture combining GPU batch inference for high-volume ingest with CPU AVX-512 fallback for low-latency point queries.

Breakeven:Generating > 500 million embeddings/month self-hosted costs ~$800/mo vs $10,000+/mo on commercial embedding APIs (92% savings).
Prototype:$65 - $220 / mo
Production:$450 - $1,800 / mo
High-Scale:$1,800 - $7,500 / mo
GENAI_GPUarch-fin-15

Ultra-Low Latency Speculative Decoding Architecture

Enterprise multi-tier inference architecture deploying a fast, lightweight draft model alongside a massive target model to achieve 2.5x inference speedup with zero quality degradation.

Breakeven:Justified for high-concurrency conversational agents where latency reduction directly increases user retention and GPU saturation efficiency.
Prototype:$1,200 - $3,500 / mo
Production:$6,500 - $22,000 / mo
High-Scale:$22,000 - $95,000 / mo
HYBRID_BARE_METALarch-fin-16

Hetzner Bare-Metal Workhorses with Cloudflare Edge Ingress

High-performance hybrid architecture running heavy compute and databases on Hetzner dedicated bare-metal servers shielded by Cloudflare CDN edge security.

Breakeven:Inflection champion for sustained high-compute (> 128 vCPUs) or high-egress (> 20 TB/mo). Bare-metal Hetzner is 75% - 85% cheaper than equivalent AWS EC2/Egress.
Prototype:$85 - $190 / mo
Production:$350 - $1,100 / mo
High-Scale:$1,800 - $8,500 / mo
HYBRID_BARE_METALarch-fin-17

Colocation High-Density NVMe Storage & Ceph Cluster

Private petabyte-scale storage cluster running Ceph distributed object and block storage on enterprise Supermicro hardware in carrier-neutral colocation facilities.

Breakeven:Inflection champion above 500 Terabytes. Blended cost is $0.0035/GB-month vs $0.023/GB-month on AWS S3 Standard (85% recurring savings).
Prototype:$650 - $1,800 / mo
Production:$2,500 - $9,500 / mo
High-Scale:$9,500 - $38,000 / mo
HYBRID_BARE_METALarch-fin-18

Bare-Metal Database Repatriation: Enterprise PostgreSQL & Patroni

High-IOPS bare-metal PostgreSQL deployment with NVMe RAID-10 storage, Patroni automated failover, and pgBouncer connection pooling bypassing AWS RDS hypervisor overhead.

Breakeven:Inflection champion for databases requiring > 15,000 sustained write IOPS. Eliminates $5,000 - $25,000/mo in AWS RDS Provisioned IOPS and instance fees.
Prototype:$180 - $450 / mo
Production:$850 - $2,800 / mo
High-Scale:$2,800 - $9,500 / mo
AI Summary & Agent Operating Digest
AEO / GEO / Perplexity Indexable

The TinyCTO Cloud Economics & FinOps Architecture Canon ('The Cloud Bill Bible') establishes an operational engineering discipline for cloud financial governance. Built upon the FinOps Foundation FOCUS 1.0 open specification, the canon details 18 production cloud architectures across 6 archetypes, 24 cloud waste typologies with detection queries and CLI remediation playbooks, a deterministic sizing wizard engine with itemized Bill-of-Materials calculations, 10 deep bilingual engineering manuals, and a 26-tool comparison matrix.

FOCUS 1.0 Ontology & AttributionStandardized 12-column billing dataset mapping BilledCost, EffectiveCost, and ChargeSubcategory across AWS, GCP, Azure, and Bare Metal with mandatory 5-key tagging.
18 Architectures & Breakeven Curves54 maturity configurations across Serverless, Containers, Karpenter Kubernetes, Streaming, GenAI GPUs, and Bare-Metal Colocation with mathematical breakeven inflection models.
24 Cloud Waste TypologiesExhaustive waste detection queries and automated remediation commands covering Compute, Storage, Networking, Database, AI/ML, and Observability cost leaks.
Deterministic Sizing WizardInteractive calculation engine modeling throughput, storage, and egress to produce itemized BOMs, realistic savings projections, and waste vulnerability disclosures.

Cloud Architecture & Breakeven Curves FAQs

Why is Serverless Event-Driven API (Lambda + DynamoDB) cost-optimal below 8 million requests/month but disadvantageous above 25 million?

Serverless architectures charge strictly for execution duration (GB-seconds) and invocation counts, meaning zero idle cost when traffic is low or bursty. However, at sustained high throughput (> 25M requests/mo), the per-million invocation premium exceeds the cost of continuously running rightsized container instances (ECS Fargate or Kubernetes with Karpenter), where CPU cores can process thousands of concurrent requests at a flat amortized hourly rate.

How does the High-Density EKS architecture achieve an 82% bin-packing efficiency?

By combining Karpenter Just-in-Time scheduling with Graviton ARM64 node pools and Vertical Pod Autoscaler (VPA) recommendation engines. Pod resource requests are right-sized based on p95 historical utilization rather than developer estimates, and Karpenter selects instances whose memory-to-CPU ratios match aggregate pending pod shapes, eliminating stranded memory capacity.

What is the breakeven point between managed streaming (Amazon MSK / Confluent) and self-hosted ClickHouse / Kafka?

Managed streaming services charge hefty operational management premiums and high cross-AZ data transfer markups. Below 50M events/month, managed services save valuable platform engineering time. However, beyond 500M events/month, self-hosting Kafka with Tiered Storage (S3) and ClickHouse on NVMe instances reduces streaming and analytical storage costs by > 80%, saving tens of thousands of dollars per month.

How does Speculative Decoding reduce LLM inference costs by 2.2x?

Large target LLMs (e.g. Llama-3.1-70B) are memory-bandwidth bound during autoregressive token generation. Speculative decoding pairs a lightweight draft model (Llama-3.2-1B) running on cheap compute to propose 4-5 tokens per step. The large target model verifies all candidate tokens in a single parallel forward pass, achieving a 2x-2.5x speedup in throughput and slashing GPU duty cycle duration without any loss in generation accuracy.

How does the Hybrid Bare-Metal Egress Shield prevent cloud bandwidth lock-in?

By locating heavy storage, video streaming, or data replication workloads on dedicated bare-metal servers (e.g. Hetzner) connected via low-latency unmetered transit to Cloudflare or AWS CloudFront edge caches. Steady-state traffic consumes $0 unmetered bare-metal bandwidth, while public cloud compute is utilized only for elastic, stateless burst capacity.

Why is Debezium CDC preferred over recurring ETL batch queries for database synchronization?

Periodic batch ETL jobs run massive `SELECT *` table scans that saturate primary database CPU, spike read IOPS, and require oversized database instance tiers. Debezium Change Data Capture reads the database Write-Ahead Log (WAL) directly at the storage engine level with near-zero CPU overhead, streaming record inserts and updates continuously without degrading operational transaction performance.