Ant Group's Bailing Releases New Native Hybrid Reasoning Model Ling-3.0-Flash

On July 24, 2026, Ant Group's Bailing officially released the new generation native hybrid reasoning model Ling-3.0-Flash. The model has a total of 124 billion parameters, with only 5.1 billion activated per computation. While significantly reducing model size and computational cost, Ling-3.0-Flash matches or even surpasses industry-leading models that are 2 to 3 times its parameter count in core metrics such as basic reasoning, instruction following, and long text processing, demonstrating higher intelligence-to-efficiency ratio and cost-effectiveness for deployment. The model is now available on OpenRouter, free for one week, and will be open-sourced thereafter.

Ant Group's Bailing Releases New Native Hybrid Reasoning Model Ling-3.0-Flash image

A key highlight of this version is its deep optimization for real-world agent applications. The model expands the training environment to over 10,000 real interactive scenarios and improves self-correction and long-term planning mechanisms for complex tasks. Whether in code writing, complex task decomposition, or in-depth multi-source research, Ling-3.0-Flash exhibits stronger autonomous end-to-end delivery capabilities, addressing the common issues of deviation or interruption in large tasks faced by traditional models.

The achievement of high intelligence with a small model size stems from a redesigned underlying computational architecture. Ling-3.0-Flash abandons the approach of simply stacking parameters and instead adopts a hybrid attention mechanism from the pre-training stage, alternating KDA linear attention layers and MLA layers in a 5:1 ratio. This design achieves a better balance between long-context efficiency and model capability.

Ant Group's Bailing Releases New Native Hybrid Reasoning Model Ling-3.0-Flash image

Furthermore, Ling-3.0-Flash upgrades the previous Lightning Attention to KDA (Kimi Delta Attention), introducing fine-grained diagonal gating in the state update of the Delta Rule. This allows the model to more accurately retain key information when processing long documents and code bases. Additionally, the model further compresses the expert activation ratio per token from 1/32 in the previous generation to 1/64, resulting in higher efficiency leverage.

To enable faster and more stable AI agent operation, Bailing has also developed supporting engineering and collaboration architectures. In terms of response speed, cluster-level hierarchical caching is introduced to avoid redundant computation in long conversations and multi-turn interactions, reducing the first-token latency for long inputs by 60% to 80% or more. For task stability, an upgraded multi-agent collaboration architecture allows different agents to work in parallel and cross-validate, effectively reducing the risk of misjudgment by a single model, providing support for high-frequency online services and real business deployment.

 

Share this article