Skip to main content

> RC_ARCH_FAM_20

Long-Context In-Memory RAG with KV Caching

Leverages large LLM context windows with prefix KV cache reuse for low-latency multi-document reasoning. Configured for prototype maturity target.

Back to All Architectures

Maturity Tier Configurations (3 Tiers)

PROTOTYPE

Leverages large LLM context windows with prefix KV cache reuse for low-latency multi-document reasoning. Configured for prototype maturity target.

Latencyp95: 1200ms
CostLow ($0.005/query)
Stack Components
cand-rag-comp-001cand-rag-comp-021cand-rag-store-001
PRODUCTION

Leverages large LLM context windows with prefix KV cache reuse for low-latency multi-document reasoning. Configured for production maturity target.

Latencyp95: 350ms
CostMedium ($0.015/query)
Stack Components
cand-rag-comp-002cand-rag-comp-022cand-rag-comp-081cand-rag-store-011
REGULATED PRODUCTION

Leverages large LLM context windows with prefix KV cache reuse for low-latency multi-document reasoning. Configured for regulated production maturity target.

Latencyp95: 550ms
CostHigh ($0.045/query)
Stack Components
cand-rag-comp-003cand-rag-comp-023cand-rag-comp-161cand-rag-comp-241cand-rag-store-021

Canonical Data Flow & Pipeline

Client Query -> API Gateway -> Query Preprocessing -> Vector Search -> Reranking / Guardrail -> Generator -> Client Stream