> RC_ARCH_FAM_20
Long-Context In-Memory RAG with KV Caching
Leverages large LLM context windows with prefix KV cache reuse for low-latency multi-document reasoning. Configured for prototype maturity target.
Maturity Tier Configurations (3 Tiers)
PROTOTYPE
Leverages large LLM context windows with prefix KV cache reuse for low-latency multi-document reasoning. Configured for prototype maturity target.
Latencyp95: 1200ms
CostLow ($0.005/query)
Stack Components
cand-rag-comp-001cand-rag-comp-021cand-rag-store-001
PRODUCTION
Leverages large LLM context windows with prefix KV cache reuse for low-latency multi-document reasoning. Configured for production maturity target.
Latencyp95: 350ms
CostMedium ($0.015/query)
Stack Components
cand-rag-comp-002cand-rag-comp-022cand-rag-comp-081cand-rag-store-011
REGULATED PRODUCTION
Leverages large LLM context windows with prefix KV cache reuse for low-latency multi-document reasoning. Configured for regulated production maturity target.
Latencyp95: 550ms
CostHigh ($0.045/query)
Stack Components
cand-rag-comp-003cand-rag-comp-023cand-rag-comp-161cand-rag-comp-241cand-rag-store-021
Canonical Data Flow & Pipeline
Client Query -> API Gateway -> Query Preprocessing -> Vector Search -> Reranking / Guardrail -> Generator -> Client Stream
