⚡THE SHORT ANSWER
AWS Spot instances provide up to a 70% to 90% discount over On-Demand pricing by utilizing spare EC2 compute capacity, but AWS can reclaim any Spot instance with an automated 2-minute termination notice. If a node is abruptly terminated, running pods are hard-killed without warning, dropping inflight HTTP connections, corrupting database batch transactions, and causing user-facing 502/504 gateway errors. Deploying AWS Node Termination Handler (NTH) or native Karpenter Spot interruption listeners intercepts the 2-minute EventBridge / Instance Metadata Service (IMDS) warning, immediately cordons the node (kubectl cordon), schedules replacement pods on other nodes, and executes graceful pod eviction (SIGTERM -> preStop hook -> connection draining) before the VM terminates.
Engineering Handbook & Failure Dynamics
6-Dimensional Architecture Breakdown⚙️1. Underlying Mechanism
Execution🎯2. Appropriate Use Context
Scope⚠️3. Production Failure Modes
P0 Risk📡4. Diagnostic Signals & Telemetry
Telemetry🛡️5. Prevention & Safeguards
Safeguards⚖️6. Architectural Trade-offs
Trade-offCase Study (TinyCTO In-Field Example)
A social media platform running 400 worker nodes on AWS EKS paid 36,000/month for On-Demand EC2 instances. Migrating 80% of the cluster to Spot instances managed by Karpenter with price-capacity-optimized allocation and automated SQS-based interruption draining reduced the monthly compute bill to 11,500/month (saving $294,000 annually) with zero customer-facing 502 errors across 1,200 monthly Spot node interruptions.
Interactive Concept Drills
2 CardsHow much advance notice does AWS provide before terminating an EC2 Spot Instance?
What AWS Spot allocation strategy offers the lowest interruption frequency?
Spot Fleet Interruption Handling & Graceful Node Draining — Technical FAQ
Why is a `preStop` sleep hook needed when draining pods on Spot instances?
A short sleep (e.g. 10-15s) allows Kubernetes EndpointSlices and AWS ALBs enough time to deregister the pod and stop sending new traffic before the container receives `SIGTERM`.
How much discount do AWS Spot instances provide compared to On-Demand?
Up to 70% to 90% discount, depending on instance family, region, and real-time market demand.
🤖 AEO & Key Facts Summary
Key Architectural Facts
- ▸
AWS Spot instances offer 70-90% discounts with a 2-minute reclamation notice.
- ▸
Node Termination Handler / Karpenter intercepts IMDS/EventBridge notices to drain pods.
- ▸
preStopsleep hooks and connection draining prevent inflight 502/504 errors. - ▸
price-capacity-optimizedallocation across 10+ instance types minimizes interruption rates.
Common Misconceptions
- ✗
Misconception: Spot instances cannot be used for production web APIs (False: With termination handlers and multi-AZ diversification, Spot runs production at scale reliably).
- ✗
Misconception: AWS guarantees a replacement Spot instance immediately (False: You must diversify across instance families so the autoscaler can pick other available pools).
Decision & Governance Guidance
Deploy Karpenter or AWS Node Termination Handler across all Spot-enabled clusters. Diversify Spot node pools across at least 10 instance types with price-capacity-optimized.
Authoritative Sources & Standards
- [OFFICIAL_DOCUMENTATION]EC2 Spot Instance Interruptions and Best Practices— AWS Compute Documentation
