Skip to main content

> RAG_MANUAL_02

Manual 02: Document Ingestion & Layout Parsing

Layout-aware visual document parsing, OCR error tolerance, tabular structure recovery, and chunking invariants.

Canonical Engineering Manual #02|TinyCTO RAG Bible

Document Ingestion & Layout Parsing

Layout-aware visual document parsing, OCR error tolerance, tabular structure recovery, and chunking invariants.

#1. The Physical Document Parsing Challenge

Standard naive text extraction tools (e.g. basic PyPDF or PDFMiner) destroy document structure by flattening multi-column text into interleaved horizontal lines and corrupting complex tabular relationships. In contrast, enterprise RAG requires layout-aware parsers that maintain structural hierarchies.

Layout-Aware Segmentation

  • Vision-Augmented Parsers: Systems such as Docling or Marker leverage lightweight vision models to segment documents into semantic bounding boxes (headers, paragraphs, tables, figure captions).
  • Table Preservation: Tabular data must be serialized into clean Markdown or HTML table structures with explicit header-cell bindings. Flattening tables into unformatted text corrupts numerical comprehension.
  • Reading Order Reconstruction: Multi-column layouts must be segmented vertically per column rather than read left-to-right across the page boundary.

#2. Chunking Architectures & Invariants

  1. Fixed-Size Chunking with Overlap:
    • Standard baseline (e.g. 512 tokens with 64-token sliding window). Easy to implement, but severs sentence boundaries and topical context.
  2. Hierarchical Parent-Child Chunking:
    • Small child chunks (128 tokens) are embedded for high-precision vector search.
    • When a child chunk is retrieved, its larger parent section (1024 tokens) is loaded into the LLM context.
  3. Propositional Chunking:
    • Decomposes complex multi-sentence paragraphs into atomic, self-contained factual propositions before embedding.