Bullish

Ant Group Releases 7.9B Ling-3.0-tiny with 90 Token/s Local Speed

16:45

Ant Group opens up Ling-3.0-tiny weights under MIT license. The 7.9B MoE model activates only 1.3B parameters per token, enabling high-speed local inference on consumer hardware.

Woofun AI reports that Ant Group has released the open-source weights for Ling-3.0-tiny, providing BF16, FP8, and INT4 versions under the MIT license. The 7.9 billion parameter model utilizes a mixture-of-experts architecture with 128 routing experts, activating only 1.3 billion parameters per token alongside one shared expert. It employs a hybrid attention mechanism combining Kimi Delta Attention and MLA in a 3:1 layer ratio. Performance benchmarks indicate the FP8 variant delivers 86–90 tokens per second on an M4 Pro MacBook and 100–105 tokens per second on a DGX Spark.

WOOFUN AI

Impact Assessment · Quick Read

The release of a highly efficient 7.9B MoE model lowers the barrier for local AI deployment on consumer-grade hardware. By achieving near-90 token/s on an M4 Pro, this model could accelerate the adoption of edge-based Agent applications. The MIT licensing further encourages rapid integration into third-party developer ecosystems, potentially increasing competition in the lightweight LLM sector.
Generated by WOOFUN AI · For reference only, not investment advice

Comments

Me
Replying to @User
0/800

No comments yet.

Notifications

Sign in to view messages
View all messagesManage subscriptions