> ML_RECIPE // ENTERPRISE-TECHNICAL-DOCUMENTATION-RAG_v1.0
Enterprise Technical Documentation RAG
Empower engineers and customer support agents to query internal runbooks and architecture docs with cited, verifiable answers.
Business Outcome
Empower engineers and customer support agents to query internal runbooks and architecture docs with cited, verifiable answers.
Zero hallucinated links or invented commands permitted on technical procedures.
Heuristic Baseline
Full-text keyword search (Elasticsearch / PostgreSQL BM25) returning top matching document snippets.
BM25 keyword search achieves Context Precision 0.58 and frequently misses semantic paraphrases.
Phase 1: Prototype Path
Chunk Markdown/PDF documentation into 500-token passages. Compute embeddings via sentence-transformers and query an in-memory vector index.
Phase 2: Production Path
Continuous doc sync via webhook pipeline into pgvector. Serve embeddings via ONNX Runtime on CPU and LLM generation via containerized vLLM with citation validation.
Compute & Placement Topologies
Pretrained off-the-shelf models; zero custom model training required
Embedding generation on CPU server; LLM synthesis via vLLM server GPU or secure hosted API
3-Plan Placement Alternatives
In-memory vector index with hosted API LLM completions (OpenAI/Anthropic) using strict chunk citations.
Single GPU server (e.g. 1x NVIDIA L4 or RTX 4090) running vLLM with quantized Mistral/Llama for internal private inference.
pgvector storage with read-replicas + CPU ONNX Runtime embedding microservice + containerized vLLM on dedicated GPU cluster with Ragas evaluation and citation verification gate.
Recommended Libraries & Tools
Governance, Safeguards & Risks
- Enforce strict chunk attribution: every generation must display clickable links to source documents.
- Sanitize input prompts against prompt injection attacks.
- Ensure sensitive intranet documents are restricted by user role authorization (RBAC).
