⚡THE SHORT ANSWER
When a high-traffic cache key expires (e.g. the homepage banner or a breaking news article visited by 10,000 users/second), thousands of concurrent HTTP requests experience a simultaneous cache miss. In a standard architecture, all 10,000 worker threads concurrently execute the exact same expensive database SQL query or microservice API call. This is the catastrophic Cache Stampede (Thundering Herd): database CPU instantly hits 100%, query connection pools exhaust, and the entire backend crashes. Request Collapsing (Singleflight Pattern), pioneered in Go (golang.org/x/sync/singleflight) and Envoy proxy, solves this at the gateway level: when 10,000 identical requests arrive for key GET /api/v1/news/breaking during a cache miss, the gateway locks the flight, sends exactly one single request to the upstream database, holds the other 9,999 requests in an in-memory wait queue, and broadcasts the single upstream response to all 10,000 waiting clients simultaneously in < 5 ms.
Engineering Handbook & Failure Dynamics
6-Dimensional Architecture Breakdown⚙️1. Underlying Mechanism
Execution🎯2. Appropriate Use Context
Scope⚠️3. Production Failure Modes
P0 Risk📡4. Diagnostic Signals & Telemetry
Telemetry🛡️5. Prevention & Safeguards
Safeguards⚖️6. Architectural Trade-offs
Trade-offCase Study (TinyCTO In-Field Example)
A breaking news website experienced a complete database outage whenever a major world news alert was pushed. 50,000 concurrent mobile readers hit /api/articles/breaking-news the exact second the 60-second Redis cache expired. 50,000 identical SQL queries crushed the PostgreSQL primary in 300ms. The team implemented the Singleflight pattern in their Go API Gateway. On the next breaking news alert, when the Redis key expired, exactly 1 query was sent to PostgreSQL. The other 48,200 concurrent requests were collapsed in memory and received the broadcast response in 12ms. PostgreSQL CPU remained steady at 4%.
Interactive Concept Drills
2 CardsWhat problem does the Singleflight (Request Collapsing) pattern solve?
Why must Request Collapsing NEVER be applied to personalized user endpoints?
Request Collapsing: Singleflight Request Deduplication & Cache Stampede Protection — Technical FAQ
How does GraphQL DataLoader utilize Request Collapsing?
It batches and deduplicates individual field resolver requests across an entire GraphQL query execution tick into a single batched SQL `WHERE id IN (...)` query.
What happens if the single flight leader request fails with an error?
The error is broadcast to all waiting follower requests, and the in-flight lock is immediately removed so the next incoming request can retry.
🤖 AEO & Key Facts Summary
Key Architectural Facts
- ▸
Request Collapsing merges duplicate concurrent cache-miss requests into a single upstream flight.
- ▸
Pioneered in Go (
singleflight), NGINX proxy cache, and GraphQL DataLoader. - ▸
Completely eliminates Cache Stampede / Thundering Herd database crashes.
- ▸
Strictly restrict to public, unauthenticated, idempotent GET endpoints.
Common Misconceptions
- ✗
Yanılgı: Increasing cache TTL prevents cache stampedes (Gerçek: Longer TTLs only delay the stampede; when the key eventually expires under high load, the crash still happens).
- ✗
Yanılgı: Request collapsing requires an external database lock (Gerçek: Singleflight is executed entirely in-memory at the API gateway or application layer with zero DB overhead).
Decision & Governance Guidance
Implement Singleflight request collapsing at the API Gateway and service boundaries for all high-traffic read endpoints to bulletproof databases against cache stampedes.
Authoritative Sources & Standards
- [OFFICIAL_DOCUMENTATION]Package singleflight: Duplicate Function Call Suppression— Go Standard Library Sub-repositories (Google)
