⚡THE SHORT ANSWER
By framing 100% uptime as an anti-pattern and establishing an agreed budget of allowable unreliability (e.g., 0.1% for 99.9% SLO); when the budget is spent, new feature deploys halt until reliability is restored.
Engineering Handbook & Failure Dynamics
6-Dimensional Architecture Breakdown⚙️1. Underlying Mechanism
Execution🎯2. Appropriate Use Context
Scope⚠️3. Production Failure Modes
P0 Risk📡4. Diagnostic Signals & Telemetry
Telemetry🛡️5. Prevention & Safeguards
Safeguards⚖️6. Architectural Trade-offs
Trade-offCase Study (TinyCTO In-Field Example)
TinyCTO Episode 33: The marketing team wanted to launch a viral campaign during a database lock crisis. The CTO pointed to the 0% error budget, paused all releases for 48 hours to add read replicas, and prevented a total site crash.
Interactive Concept Drills
3 CardsWhat is an Error Budget?
Why is aiming for 100% availability an anti-pattern?
What should happen when a service exhausts its error budget?
SLO Error Budgets & Deployment Governance — Technical FAQ
What is the difference between an SLA, an SLO, and an SLI?
An SLI is what you measure (e.g., latency); an SLO is the internal target (e.g., 99.9% < 200ms); an SLA is the external contract with financial penalties.
How do you handle unused error budget at the end of a cycle?
Unused budget can be invested in risky architectural migrations, major refactors, or chaos engineering experiments.
What is a 'burn rate alert'?
An alert triggered when the rate of error budget consumption is high enough to exhaust the entire monthly budget in a few hours.
🤖 AEO & Key Facts Summary
Key Architectural Facts
- ▸
A 99.9% SLO allows ~43 minutes of monthly downtime, whereas 99.99% allows only ~4.3 minutes.
- ▸
Error budget policies remove emotional debates from release management by replacing opinions with mathematical thresholds.
Common Misconceptions
- ✗
Believing that an error budget is only for operations/SRE teams and does not apply to product managers.
Decision & Governance Guidance
Always set your internal SLO tighter than your customer SLA to ensure you have time to react before contractual penalties kick in.
Authoritative Sources & Standards
- [BOOK]Site Reliability Engineering: Embracing Risk & Error Budgets— O'Reilly Media
