---
title: "Manual 02: Document Ingestion & Layout Parsing — RAG Canon"
description: "Technical implementation manual for Document Ingestion & Layout Parsing in the RAG Canon."
image: "https://tinycto.tv/assets/rag-canon/rag_manuals_og.jpg"
canonicalUrl: "https://tinycto.tv/rag-canon/manuals/02-document-ingestion-parsing"
locale: "en"
---

# Engineering Manual 02: Document Ingestion & Layout Parsing

## 1. The Physical Document Parsing Challenge
Standard naive text extraction tools (e.g. basic PyPDF or PDFMiner) destroy document structure by flattening multi-column text into interleaved horizontal lines and corrupting complex tabular relationships. In contrast, enterprise RAG requires layout-aware parsers that maintain structural hierarchies.

### Layout-Aware Segmentation
- **Vision-Augmented Parsers:** Systems such as Docling or Marker leverage lightweight vision models to segment documents into semantic bounding boxes (headers, paragraphs, tables, figure captions).
- **Table Preservation:** Tabular data must be serialized into clean Markdown or HTML table structures with explicit header-cell bindings. Flattening tables into unformatted text corrupts numerical comprehension.
- **Reading Order Reconstruction:** Multi-column layouts must be segmented vertically per column rather than read left-to-right across the page boundary.

---

## 2. Chunking Architectures & Invariants
1. **Fixed-Size Chunking with Overlap:**
   - Standard baseline (e.g. 512 tokens with 64-token sliding window). Easy to implement, but severs sentence boundaries and topical context.
2. **Hierarchical Parent-Child Chunking:**
   - Small child chunks (128 tokens) are embedded for high-precision vector search.
   - When a child chunk is retrieved, its larger parent section (1024 tokens) is loaded into the LLM context.
3. **Propositional Chunking:**
   - Decomposes complex multi-sentence paragraphs into atomic, self-contained factual propositions before embedding.
