Canonical Engineering Manual #02|TinyCTO RAG Bible
Document Ingestion & Layout Parsing
Layout-aware visual document parsing, OCR error tolerance, tabular structure recovery, and chunking invariants.
Canon Certified 12 min
#1. The Physical Document Parsing Challenge
Standard naive text extraction tools (e.g. basic PyPDF or PDFMiner) destroy document structure by flattening multi-column text into interleaved horizontal lines and corrupting complex tabular relationships. In contrast, enterprise RAG requires layout-aware parsers that maintain structural hierarchies.
Layout-Aware Segmentation
- Vision-Augmented Parsers: Systems such as Docling or Marker leverage lightweight vision models to segment documents into semantic bounding boxes (headers, paragraphs, tables, figure captions).
- Table Preservation: Tabular data must be serialized into clean Markdown or HTML table structures with explicit header-cell bindings. Flattening tables into unformatted text corrupts numerical comprehension.
- Reading Order Reconstruction: Multi-column layouts must be segmented vertically per column rather than read left-to-right across the page boundary.
#2. Chunking Architectures & Invariants
- Fixed-Size Chunking with Overlap:
- Standard baseline (e.g. 512 tokens with 64-token sliding window). Easy to implement, but severs sentence boundaries and topical context.
- Hierarchical Parent-Child Chunking:
- Small child chunks (128 tokens) are embedded for high-precision vector search.
- When a child chunk is retrieved, its larger parent section (1024 tokens) is loaded into the LLM context.
- Propositional Chunking:
- Decomposes complex multi-sentence paragraphs into atomic, self-contained factual propositions before embedding.
