Kimi K3 Efficiency Surges 2.5x with 2.8T Parameters and Rewritten Architecture
Moonshot AI releases Kimi K3 featuring 2.8T parameters and 2.5x scaling efficiency. The model matches Fable 5 in benchmarks via architectural upgrades and specialized post-training.
Woofun AI reports that Moonshot AI has published the technical report for Kimi K3, a model with 2.8 trillion parameters where each token activates 104 billion. The new architecture achieves 2.5 times higher scaling efficiency than K2, requiring only 40% of the computational volume for equivalent validation loss.
The model introduces KDA for long sequences, Attention Residuals to prevent information dilution, and a redesigned MoE with 896 routing experts activating 16 per token. Post-training involves merging nine specialized experts trained on general, Agent, and code tasks. K3 performance now approaches or surpasses Fable 5 and GPT-5.6 Sol in code, search, and tool call benchmarks.
Comments
No comments yet.