> ML_LITERATURE // ZHENG-2024-SGLANG-EFFICIENT-EXECUTION-STRUCTURED-LANGUAGE-MODEL-PROGRAMS_v1.0
SGLang: Efficient Execution of Structured Language Model Programs
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Jeff Huang, Chuyue Sun, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Hao Zhang · arXiv preprint (2024)
systems2024industry-standardartifactsAvailable
Principal Contribution
Introduced RadixAttention for multi-turn hierarchical KV-cache sharing and a structured decoding runtime, drastically accelerating multi-call agent workflows and JSON generation.
Operational Relevance
Serves as qualified reference for implementing task-text-generation in production systems.
Assumptions
- Underlying computational topology and mathematical bounds adhere to established convexity/smoothness guarantees
Limitations
- Hardware runtime speedups, privacy budgets, and convergence depend on hyperparameters and network communication limits
Connected Algorithms, Architectures & Tools
Related Algorithms:
Related Architectures:
Implementing Libraries:
