Skip to main content

> ML_LITERATURE // PENG-2023-RWKV-REINVENTING-RNNS-TRANSFORMER-ERA_v1.0

RWKV: Reinventing RNNs for the Transformer Era

Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Huanqi Shen, Rui-Jie Zhu, Matteo Bartolo, Hao-Chen Song, Zhen-Yu Zhang, Evelina Xu, Médéric Boquien, Fei Chen, Gyu-Wan Kim, Haowen Hou, He-Ming Li, Jan Kocoń, Jiaming Kong, Linhai Ren, Liying Zheng, Marius Valdenegro-Toro, Nektarios Tsoutsos, Roberto Perez, Shitao Tang, Soufiane Guessous, Tomasz Szandala, Xuan-Shen Chen, Yong-Shuai Hou, Yuxuan Song, Zhen-Yu Chen, Zhi-Hao Lu, Zhi-Ming Zhao · Conference on Empirical Methods in Natural Language Processing (EMNLP) (2023)

seminal-architecture2023industry-standardthirdPartyReproduced

Principal Contribution

Designed an architecture combining the parallelized training advantages of transformers with the constant-time, constant-memory inference efficiency of RNNs (Receptance Weighted Key Value).

Operational Relevance

Directly guides architectural decisions, alignment strategy, and serving infrastructure for task-text-generation.

Assumptions

  • Empirical distribution regularity holds and target domain adheres to pretraining linguistic/visual support

Limitations

  • Resource scaling, inference memory requirements, and alignment robustness vary with model size and hardware topology

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: