Ant L3.0-flash Open Sources 124B MoE Model with 128GB FP8 Variant
inclusionAI releases Ling-3.0-flash under MIT license, offering a 128GB FP8 option. The 124B parameter model targets Agent tasks, matching larger predecessors in benchmarks.
Woofun AI reports that inclusionAI has officially open-sourced the Ling-3.0-flash model weights under the MIT license, deploying both BF16 and FP8 variants on Hugging Face and ModelScope for self-deployment via SGLang or vLLM.
The 124-billion parameter Mixture-of-Experts model activates only 5.1 billion parameters per generation and supports a 256,000-token context. While the BF16 version requires approximately 255GB, the FP8 quantized version reduces this to 128GB with a maximum benchmark deviation of 1.57 points. Official evaluations indicate performance parity or superiority over the trillion-parameter Ring-2.6-1T on most metrics, specifically targeting programming, search, and tool invocation tasks.
Comments
No comments yet.