Bullish
Ant Group Releases 7.9B Ling-3.0-tiny with 90 Token/s Local Speed
16:45
Ant Group opens up Ling-3.0-tiny weights under MIT license. The 7.9B MoE model activates only 1.3B parameters per token, enabling high-speed local inference on consumer hardware.
Woofun AI reports that Ant Group has released the open-source weights for Ling-3.0-tiny, providing BF16, FP8, and INT4 versions under the MIT license. The 7.9 billion parameter model utilizes a mixture-of-experts architecture with 128 routing experts, activating only 1.3 billion parameters per token alongside one shared expert. It employs a hybrid attention mechanism combining Kimi Delta Attention and MLA in a 3:1 layer ratio. Performance benchmarks indicate the FP8 variant delivers 86–90 tokens per second on an M4 Pro MacBook and 100–105 tokens per second on a DGX Spark.
WOOFUN AI
Impact Assessment · Quick Read
The release of a highly efficient 7.9B MoE model lowers the barrier for local AI deployment on consumer-grade hardware. By achieving near-90 token/s on an M4 Pro, this model could accelerate the adoption of edge-based Agent applications. The MIT licensing further encourages rapid integration into third-party developer ecosystems, potentially increasing competition in the lightweight LLM sector.
Generated by WOOFUN AI · For reference only, not investment advice
Comments
No comments yet.