Skip to main content

> ML_LITERATURE // ZHENG-2024-SGLANG-EFFICIENT-EXECUTION-STRUCTURED-LANGUAGE-MODEL-PROGRAMS_v1.0

SGLang: Efficient Execution of Structured Language Model Programs

Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Jeff Huang, Chuyue Sun, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Hao Zhang · arXiv preprint (2024)

systems2024industry-standardartifactsAvailable

Principal Contribution

Introduced RadixAttention for multi-turn hierarchical KV-cache sharing and a structured decoding runtime, drastically accelerating multi-call agent workflows and JSON generation.

Operational Relevance

Serves as qualified reference for implementing task-text-generation in production systems.

Assumptions

  • Underlying computational topology and mathematical bounds adhere to established convexity/smoothness guarantees

Limitations

  • Hardware runtime speedups, privacy budgets, and convergence depend on hyperparameters and network communication limits

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: