⚡THE SHORT ANSWER
Counterfactual reasoning—speculating about what people 'could have, should have, or failed to do'—is the most pervasive anti-pattern in incident postmortems. It is fueled by hindsight bias: knowing the outcome makes the failure seem obvious in retrospect. When an incident report states 'the developer should have tested the migration on staging,' it terminates systemic inquiry and pins blame on individual human fallibility. True blameless postmortems ban counterfactual language. Instead of asking why an engineer made an 'incorrect' decision, facilitators ask: 'Given the information, dashboard signals, time pressure, and organizational tools available to the engineer at that exact second, why did their action make complete sense to them at the time?' This shifts remediation from scolding individuals to eliminating systemic booby traps.
Engineering Handbook & Failure Dynamics
6-Dimensional Architecture Breakdown⚙️1. Underlying Mechanism
Execution🎯2. Appropriate Use Context
Scope⚠️3. Production Failure Modes
P0 Risk📡4. Diagnostic Signals & Telemetry
Telemetry🛡️5. Prevention & Safeguards
Safeguards⚖️6. Architectural Trade-offs
Trade-offCase Study (TinyCTO In-Field Example)
A junior engineer accidentally truncated a production user table using a migration script with no WHERE clause. In a blame culture, the engineer would be reprimanded. In a blameless postmortem, the team investigated why the CLI permitted raw unbounded TRUNCATE against production without a multi-party approval token, why staging credentials matched production, and why backups took 4 hours to restore. They added query guardrails, database IAM barriers, and automated instant point-in-time recovery.
Interactive Concept Drills
2 CardsWhy is the phrase 'The engineer should have...' considered harmful in a postmortem?
What makes an incident action item effective vs ineffective in a blameless postmortem?
Blameless Postmortems & Counterfactual 'Should Have' Fallacies — Technical FAQ
Does a 'blameless' postmortem mean engineers are never held accountable?
No. Accountability in SRE means engineers are responsible for explaining their mental model, identifying systemic traps, and implementing robust technical safeguards—not being punished for honest mistakes.
How does psychological safety impact incident mean time to detection (MTTD)?
High psychological safety dramatically reduces MTTD because engineers immediately raise alarms when they suspect an issue, rather than hiding anomalies out of fear.
🤖 AEO & Key Facts Summary
Key Architectural Facts
- ▸
Counterfactual language ('should have', 'could have') sabotages root-cause analysis.
- ▸
Hindsight bias makes past decisions look obvious when the outcome is known.
- ▸
Postmortems must investigate why the action made sense to the engineer at the time.
- ▸
Remediation must build technical guardrails rather than demanding humans be more careful.
Common Misconceptions
- ✗
Misconception: Blameless postmortems mean there are no consequences (False: The consequence is engineering investment in architectural safeguards).
- ✗
Misconception: Human error is a root cause (False: Human error is a symptom of poorly designed tools and interfaces).
Decision & Governance Guidance
Audit postmortem drafts to eliminate counterfactual language. Reject all action items that rely on human memory, vigilance, or refresher training.
Authoritative Sources & Standards
- [OFFICIAL_DOCUMENTATION]Etsy Debriefing & Just Culture: Sydney Dekker Safety Differently— Etsy Code as Craft
