> ML_LITERATURE // XIAO-2023-SMOOTHQUANT-ACCURATE-EFFICIENT-POST-TRAINING-QUANTIZATION_v1.0
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, Song Han · International Conference on Machine Learning (ICML) (2023)
algorithm2023industry-standardthirdPartyReproduced
Principal Contribution
Designed a mathematically equivalent per-channel scaling transformation smoothing activation outliers onto weights, enabling true W8A8 INT8 matrix multiplication on Tensor Cores.
Operational Relevance
Directly guides architectural decisions, alignment strategy, and serving infrastructure for task-text-generation.
Assumptions
- Empirical distribution regularity holds and target domain adheres to pretraining linguistic/visual support
Limitations
- Resource scaling, inference memory requirements, and alignment robustness vary with model size and hardware topology
Connected Algorithms, Architectures & Tools
Related Algorithms:
Related Architectures:
Implementing Libraries:
