---
title: "Manual 05: Context Engineering & Prompt Memory — RAG Canon"
description: "Technical implementation manual for Context Engineering & Prompt Memory in the RAG Canon."
image: "https://tinycto.tv/assets/rag-canon/rag_manuals_og.jpg"
canonicalUrl: "https://tinycto.tv/rag-canon/manuals/05-context-engineering-memory"
locale: "en"
---

# Engineering Manual 05: Context Engineering & Prompt Memory

## 1. Overcoming the Lost-in-the-Middle Phenomenon
Empirical studies demonstrate that transformer attention mechanisms disproportionately prioritize tokens at the extreme beginning and end of long prompt contexts. Passages situated in the middle of long contexts experience severe retrieval degradation.

### Context Budgeting & Layout Optimization
1. **Rank-Aware Context Ordering:**
   - Place the highest-scoring candidate passages at the very beginning and very end of the context window.
   - Relegate supplementary background context to the center.
2. **Token Headroom Invariants:**
   - Context assembly must never exceed 70% of the total model context window, guaranteeing at least 30% headroom for complex reasoning and structured output generation.

---

## 2. Dynamic Memory Architectures
- **Prefix KV Caching:** Reusing precomputed key-value attention tensors across requests for identical system prompts and base retrieval passages, cutting time-to-first-token (TTFT) by up to 75%.
- **Episodic & Entity Memory:** Extracting key user preferences and relational entities into dynamic knowledge graphs to maintain continuity across multi-turn sessions.
