The JUDGE Framework for Human Decisions in Automated Systems
Use JUDGE: Jurisdiction, Uncertainty, Downside, Grounding, and Escalation. The framework converts vague calls for oversight into five testable decision controls. Human judgment is valuable precisely where a system cannot reduce the decision to a stable rule without losing context, authority, values, or accountability. It should not become an ornamental approval step or an excuse to leave unsafe automation unbounded. The control must be designed around the actual decision and its consequences.
01.A Predictable TinyCTO Incident
The agent had high confidence, the operator had no authority, and the customer had no recovery path. Every participant believed someone else owned the decision. The failure is not that a human disappeared from the interface. The failure is that intent, evidence, authority, reversibility, and accountability stopped travelling together. A polished workflow can therefore remain procedurally correct while becoming operationally wrong.
02.The Governing Principle
JUDGE should run before consequential automation is approved and again whenever the system encounters a novel condition. It is a design and runtime checklist, not an ethics poster. This distinction matters because automation changes the economics of decisions. It can repeat a useful action at enormous scale, but it can also repeat an invalid assumption faster than an organization can notice. Good judgment does not compete with automation; it defines the safe operating envelope in which autonomy is earned.
03.What Good Implementation Looks Like
- Jurisdiction: name the authorized decision owner and override owner. - Uncertainty: expose what the system does not know or cannot verify. - Downside: model credible harm, blast radius, and reversibility. - Grounding: show evidence, provenance, constraints, and competing signals. - Escalation: define stop conditions, response time, handoff, and review. These controls must be visible at runtime. A policy document that cannot stop, narrow, explain, or reverse system behavior is not an operational safeguard. Teams should test the path under realistic time pressure, incomplete evidence, unavailable reviewers, and partial failure.
04.Common Failure Modes & Anti-Patterns
- The framework is completed once and never tested at runtime. - Confidence is treated as evidence. - Escalation names a team but not an accountable individual. - Downside analysis ignores people who do not operate the system. The recurring anti-pattern is responsibility without agency: a person is named accountable after the system has hidden evidence, removed time, narrowed options, or completed the action. That is not meaningful human oversight. It is liability routing.
05.Practical Review Framework
1. Who owns the objective and who may override the system? 2. What evidence, uncertainty, provenance, and alternatives are visible? 3. What is the worst credible consequence, and is the full outcome reversible? 4. When must the system pause or escalate? 5. How will the decision and its outcome improve policy, evaluation, and design?
JUDGE BEFORE AUTONOMY > The system had five nines of availability and zero owners of the decision.

