Kimi Linear: An Expressive, Efficient Attention Architecture

Other: machine learning, computer science, artificial intelligence(arxiv.org)view on HackerNews
Kimi Linearattention architecturemachine learningartificial intelligencecomputer sciencehybrid linear attentionfair comparisonsreinforcement learningscaling regimespretrainingdecoding throughputKV cache usage

Author: ronfriedhaber

Date: 7/28/2026

Article Summary:
The paper introduces Kimi Linear, a hybrid linear attention architecture that outperforms full attention in various scenarios, including short-context, long-context, and reinforcement learning (RL) scaling regimes.