Skip to main content

> RAG_MANUAL_01

Manual 01: Executive Architecture Overview

Architectural foundations, the RAG vs Fine-Tuning vs Long-Context decision frontier, and lifecycle stages.

Canonical Engineering Manual #01|TinyCTO RAG Bible

Executive Architecture Overview

Architectural foundations, the RAG vs Fine-Tuning vs Long-Context decision frontier, and lifecycle stages.

#1. Problem Space & Architectural Foundations

Modern enterprise knowledge systems require reliable, factual, and low-latency access to institutional documents without the hallucinations inherent to static parametric language models. Retrieval-Augmented Generation (RAG) decouples parametric reasoning from non-parametric factual storage, allowing organizations to ground LLM generations in verifiable, continuously refreshed source documents.

The Decision Frontier: RAG vs Fine-Tuning vs Long Context

  1. Retrieval-Augmented Generation (RAG):
    • Best for: Dynamic knowledge bases, rapid daily updates, strict verifiable citations, and multi-tenant access control.
    • Economics: Low upfront cost, predictable token operational costs, sub-second indexing of new files.
  2. Parameter-Efficient Fine-Tuning (PEFT / LoRA):
    • Best for: Adapting specialized linguistic style, jargon, syntax, or tone. Not suitable for fast-changing factual knowledge.
    • Economics: Moderate GPU compute upfront; low runtime token cost.
  3. Long-Context In-Memory LLMs:
    • Best for: Single-session deep synthesis across few documents (<200k tokens).
    • Economics: Prohibitive latency and linear token costs across millions of recurring user requests.

#2. Core Architectural Lifecycle Stages

A resilient production RAG pipeline operates across 6 discrete lifecycle stages:

  1. Document Ingestion & Parsing: Extracting clean text and preserving layout from heterogeneous files (PDFs, Office documents, HTML, scans).
  2. Chunking & Embedding Generation: Segmenting text into semantic units and encoding into dense, sparse, or late-interaction vector spaces.
  3. Indexing & Hybrid Search: Managing multi-tenant vector storage, inverted lexical indices, and approximate nearest neighbor (ANN) graphs.
  4. Context Engineering & Reranking: Filtering candidate passages via deep cross-encoders and packing prompts to avoid lost-in-the-middle degradation.
  5. Generation & Verification: Synthesizing answers with constrained logit masking and automated citation verification.
  6. Guardrails & Audit Logging: Intercepting adversarial prompt injections, redacting PII, and logging immutable compliance trails.