Canonical Engineering Manual #05|TinyCTO RAG Bible
Context Engineering & Prompt Memory
Mitigating lost-in-the-middle attention degradation, prompt packing, prefix KV caching in vLLM, and episodic memory.
Canon Certified 10 min
#1. Overcoming the Lost-in-the-Middle Phenomenon
Empirical studies demonstrate that transformer attention mechanisms disproportionately prioritize tokens at the extreme beginning and end of long prompt contexts. Passages situated in the middle of long contexts experience severe retrieval degradation.
Context Budgeting & Layout Optimization
- Rank-Aware Context Ordering:
- Place the highest-scoring candidate passages at the very beginning and very end of the context window.
- Relegate supplementary background context to the center.
- Token Headroom Invariants:
- Context assembly must never exceed 70% of the total model context window, guaranteeing at least 30% headroom for complex reasoning and structured output generation.
#2. Dynamic Memory Architectures
- Prefix KV Caching: Reusing precomputed key-value attention tensors across requests for identical system prompts and base retrieval passages, cutting time-to-first-token (TTFT) by up to 75%.
- Episodic & Entity Memory: Extracting key user preferences and relational entities into dynamic knowledge graphs to maintain continuity across multi-turn sessions.
