---
title: "Manual 01: Executive Architecture Overview — RAG Canon"
description: "Technical implementation manual for Executive Architecture Overview in the RAG Canon."
image: "https://tinycto.tv/assets/rag-canon/rag_manuals_og.jpg"
canonicalUrl: "https://tinycto.tv/rag-canon/manuals/01-executive-overview"
locale: "en"
---

# Engineering Manual 01: Executive Architecture Overview

## 1. Problem Space & Architectural Foundations
Modern enterprise knowledge systems require reliable, factual, and low-latency access to institutional documents without the hallucinations inherent to static parametric language models. Retrieval-Augmented Generation (RAG) decouples parametric reasoning from non-parametric factual storage, allowing organizations to ground LLM generations in verifiable, continuously refreshed source documents.

### The Decision Frontier: RAG vs Fine-Tuning vs Long Context
1. **Retrieval-Augmented Generation (RAG):**
   - *Best for:* Dynamic knowledge bases, rapid daily updates, strict verifiable citations, and multi-tenant access control.
   - *Economics:* Low upfront cost, predictable token operational costs, sub-second indexing of new files.
2. **Parameter-Efficient Fine-Tuning (PEFT / LoRA):**
   - *Best for:* Adapting specialized linguistic style, jargon, syntax, or tone. Not suitable for fast-changing factual knowledge.
   - *Economics:* Moderate GPU compute upfront; low runtime token cost.
3. **Long-Context In-Memory LLMs:**
   - *Best for:* Single-session deep synthesis across few documents (<200k tokens).
   - *Economics:* Prohibitive latency and linear token costs across millions of recurring user requests.

---

## 2. Core Architectural Lifecycle Stages
A resilient production RAG pipeline operates across 6 discrete lifecycle stages:
1. **Document Ingestion & Parsing:** Extracting clean text and preserving layout from heterogeneous files (PDFs, Office documents, HTML, scans).
2. **Chunking & Embedding Generation:** Segmenting text into semantic units and encoding into dense, sparse, or late-interaction vector spaces.
3. **Indexing & Hybrid Search:** Managing multi-tenant vector storage, inverted lexical indices, and approximate nearest neighbor (ANN) graphs.
4. **Context Engineering & Reranking:** Filtering candidate passages via deep cross-encoders and packing prompts to avoid lost-in-the-middle degradation.
5. **Generation & Verification:** Synthesizing answers with constrained logit masking and automated citation verification.
6. **Guardrails & Audit Logging:** Intercepting adversarial prompt injections, redacting PII, and logging immutable compliance trails.
