Canonical Engineering Manual #01|TinyCTO RAG Bible
Executive Architecture Overview
Architectural foundations, the RAG vs Fine-Tuning vs Long-Context decision frontier, and lifecycle stages.
Canon Certified 8 min
#1. Problem Space & Architectural Foundations
Modern enterprise knowledge systems require reliable, factual, and low-latency access to institutional documents without the hallucinations inherent to static parametric language models. Retrieval-Augmented Generation (RAG) decouples parametric reasoning from non-parametric factual storage, allowing organizations to ground LLM generations in verifiable, continuously refreshed source documents.
The Decision Frontier: RAG vs Fine-Tuning vs Long Context
- Retrieval-Augmented Generation (RAG):
- Best for: Dynamic knowledge bases, rapid daily updates, strict verifiable citations, and multi-tenant access control.
- Economics: Low upfront cost, predictable token operational costs, sub-second indexing of new files.
- Parameter-Efficient Fine-Tuning (PEFT / LoRA):
- Best for: Adapting specialized linguistic style, jargon, syntax, or tone. Not suitable for fast-changing factual knowledge.
- Economics: Moderate GPU compute upfront; low runtime token cost.
- Long-Context In-Memory LLMs:
- Best for: Single-session deep synthesis across few documents (<200k tokens).
- Economics: Prohibitive latency and linear token costs across millions of recurring user requests.
#2. Core Architectural Lifecycle Stages
A resilient production RAG pipeline operates across 6 discrete lifecycle stages:
- Document Ingestion & Parsing: Extracting clean text and preserving layout from heterogeneous files (PDFs, Office documents, HTML, scans).
- Chunking & Embedding Generation: Segmenting text into semantic units and encoding into dense, sparse, or late-interaction vector spaces.
- Indexing & Hybrid Search: Managing multi-tenant vector storage, inverted lexical indices, and approximate nearest neighbor (ANN) graphs.
- Context Engineering & Reranking: Filtering candidate passages via deep cross-encoders and packing prompts to avoid lost-in-the-middle degradation.
- Generation & Verification: Synthesizing answers with constrained logit masking and automated citation verification.
- Guardrails & Audit Logging: Intercepting adversarial prompt injections, redacting PII, and logging immutable compliance trails.
