Skip to main content

> ML_RECIPE // ENTERPRISE-TECHNICAL-DOCUMENTATION-RAG_v1.0

Enterprise Technical Documentation RAG

Empower engineers and customer support agents to query internal runbooks and architecture docs with cited, verifiable answers.

semantic search ragsaas ecommerceApache-2.0moderate-cloud
Back to All Recipes

Business Outcome

Empower engineers and customer support agents to query internal runbooks and architecture docs with cited, verifiable answers.

Acceptance Criteria:

Zero hallucinated links or invented commands permitted on technical procedures.

Heuristic Baseline

Full-text keyword search (Elasticsearch / PostgreSQL BM25) returning top matching document snippets.

Baseline Evaluation:

BM25 keyword search achieves Context Precision 0.58 and frequently misses semantic paraphrases.

Phase 1: Prototype Path

Chunk Markdown/PDF documentation into 500-token passages. Compute embeddings via sentence-transformers and query an in-memory vector index.

Hardware: Developer workstation with 16GB RAM

Phase 2: Production Path

Continuous doc sync via webhook pipeline into pgvector. Serve embeddings via ONNX Runtime on CPU and LLM generation via containerized vLLM with citation validation.

Hardware: Server with 1x NVIDIA A10G/L4 (24GB VRAM) for vLLM, or CPU host with external private API

Compute & Placement Topologies

Training Placement

Pretrained off-the-shelf models; zero custom model training required

Inference Placement

Embedding generation on CPU server; LLM synthesis via vLLM server GPU or secure hosted API

3-Plan Placement Alternatives

Plan A: Simplest Viable

In-memory vector index with hosted API LLM completions (OpenAI/Anthropic) using strict chunk citations.

Plan B: Hardware-Fitted

Single GPU server (e.g. 1x NVIDIA L4 or RTX 4090) running vLLM with quantized Mistral/Llama for internal private inference.

Plan C: Production-Ready

pgvector storage with read-replicas + CPU ONNX Runtime embedding microservice + containerized vLLM on dedicated GPU cluster with Ragas evaluation and citation verification gate.

Recommended Libraries & Tools

★ PRIMARY TOOLTransformersHugging Face
View
vLLMvLLM Project / UC Berkeley
View
Transformers.jsHugging Face
View
Sentence TransformersHugging Face / UKP Lab
View

Governance, Safeguards & Risks

Governance Safeguards:
  • Enforce strict chunk attribution: every generation must display clickable links to source documents.
  • Sanitize input prompts against prompt injection attacks.
  • Ensure sensitive intranet documents are restricted by user role authorization (RBAC).