Skip to main content

> RAG_CANON_INDEX_v1.0

Retrieval & Agentic AI Architecture Canon

The RAG Bible — Authoritative engineering canon for enterprise retrieval-augmented generation and autonomous agent workflows with deterministic selection heuristics.

24Architecture Families
180Retrieval Concepts
90Design Patterns
75Standards & Frameworks

Choose Your Entry Vector

Whether synthesizing a zero-to-one architecture or verifying compliance for a regulated production cluster.

Canonical Engineering Pillars

Foundational primitives guaranteeing safety, sub-second latency, and deterministic cost control across lifecycle stages.

Non-RAG Necessity Gate

Deterministic gating logic: Rejects unnecessary vector databases and agent loops when static prompts, cached lookups, or relational search satisfy business requirements.

Decision LogicFail-Closed

150 Guardrail Controls

Six-stage protection pipeline enforcing automated PII scrubbing, chunk poisoning defense, role-based document access control (RBAC), and hallucination circuit breakers.

Pipeline Stages6 Lifecycle Stages

High-Stakes Compliance

75 statutory standards and 22 high-stakes compliance profiles covering clinical healthcare, algorithmic finance, automated hiring, and critical defense systems.

StandardsISO 42001 · EU AI Act

30 Playbooks

Step-by-step guides translating proof-of-concept Jupyter notebooks into resilient Kubernetes clusters with zero-downtime index migration.

Lifecycle12 Stages

120 Failure Modes

Exhaustive catalogue of real-world retrieval failures including lost-in-the-middle context degradation, embedding drift, and tool call deadlocks.

Mitigations100% Mapped

AI Agent Interface (MCP)

Standardized Model Context Protocol (MCP) contract pinned to 2025-11-25, /llms.txt manifest, and dynamic ?format=md content negotiation.

ProtocolMCP 2025-11-25
AI Summary & Agent Operating Digest
AEO / GEO Indexable

The TinyCTO Retrieval & Agentic AI Architecture Canon (The RAG Bible) v1.0 is an enterprise technical specification governing 24 architecture families across 3 maturity targets (Prototype, Production, Regulated Production = 72 configurations). Features include a deterministic Non-RAG Necessity Gate, 180 retrieval concepts, 90 patterns, 350 literature citations, 100 evaluation assets, 120 failure modes, 150 guardrail controls, and 8 comprehensive bilingual engineering manuals. Machine-readable endpoints include /llms.txt, /llms-full.txt, and Markdown negotiation (?format=md).

Recommended ArchitecturesColBERT Late-Interaction · GraphRAG · Adaptive RAG · Self-RAG
Preventative GateNon-RAG Necessity Gate (Direct Prompt · LRU Cache · SQL Search)
Machine AccessGET ?format=md · Accept: text/markdown · /llms.txt

Frequently Asked Questions

What is the TinyCTO Retrieval & Agentic AI Architecture Canon (The RAG Bible)?

The TinyCTO Retrieval & Agentic AI Architecture Canon (The RAG Bible) is an authoritative, evidence-backed engineering reference for enterprise retrieval-augmented generation and autonomous agent systems. It encompasses 24 architecture families, 72 maturity configurations, 180 retrieval concepts, 90 design patterns, 100 evaluation assets, 120 failure modes, 150 guardrail controls, and 8 deep bilingual engineering manuals.

How does the Non-RAG Necessity Gate prevent unnecessary AI complexity and cloud costs?

Before provisioning vector databases, dense embedding models, or complex multi-agent loops, the Canon evaluates a deterministic Non-RAG Necessity Gate. If a problem can be solved with static prompts, deterministic key-value cache lookups (Redis/Memcached), or classical relational/BM25 search, the engine routes away from RAG—saving compute latency, avoiding token consumption, and preventing Cloud Bill inflation.

What are the 24 Architecture Families and 3 maturity targets?

The Canon structures 24 canonical retrieval paradigms—from Naive Linear RAG, Sentence-Window, and Hierarchical RAPTOR, to HyDE, ColBERT Late-Interaction, GraphRAG, Corrective CRAG, Self-RAG, Adaptive RAG, and Hierarchical Multi-Agent Teams. Each family is specified across 3 maturity targets: Prototype (low-cost rapid validation), Production (high-availability, sub-second latency, CI/CD pipeline integration), and Regulated Production (air-gapped, zero-retention, audit-trailed).

How do high-stakes regulatory frameworks (EU AI Act, HIPAA, ISO 42001) apply to RAG?

The Canon profiles 75 global standards and 22 high-stakes domain compliance requirements (clinical healthcare, algorithmic trading, critical infrastructure, automated hiring). For regulated systems, it mandates 150 protective controls including automated PII scrubbing, cryptographic provenance verification, deterministic audit logging, and fail-closed safety circuit breakers.

How can AI agents and LLMs interact with the RAG Canon?

All RAG Canon surfaces support native AI agent ingestion. Agents can request content negotiation via `Accept: text/markdown` or `?format=md`, consume high-level indexes via `/llms.txt` and `/llms-full.txt`, or interact via Model Context Protocol (MCP) L1 read-only capabilities pinned to protocol revision 2025-11-25.

How does the Canon mitigate retrieval failure modes and incident patterns?

The Canon catalogues 120 empirical failure modes across 5 lifecycle stages (ingestion, retrieval, context synthesis, generation, ops). Each failure pattern—such as lost-in-the-middle context degradation, embedding drift, chunk truncation poisoning, and agentic tool deadlocks—is mapped to proven architectural mitigations and automated evaluation benchmarks.