Ant Group unveils Ling-3.0-Flash AI, achieving top-tier performance at a reduced scale.

Ant Group has unveiled Ling-3.0-Flash, a new AI model that punches well above its weight. The model carries 124 billion total parameters but activates only 5.1 billion per token — a tiny fraction of what rival models use, according to National Post.
Despite that lean footprint, Ling-3.0-Flash matches or beats models two to three times its size on key benchmarks. The model is now available free through August 3, 2026, via OpenRouter and Vercel AI Gateway, according to The Province.
Most large AI models activate a huge share of their parameters for every task. Ling-3.0-Flash flips that script. It activates just 5.1 billion of its 124 billion parameters per token. That ratio — called the expert activation ratio — has been cut from 1-in-32 to 1-in-64 compared to earlier designs, according to Fort McMurray Today.
The result is a much higher "efficiency leverage" — meaning the model gets more done with less compute. Ant Group says Ling-3.0-Flash outperforms competitors on foundational reasoning, instruction following, and long-context processing, according to Northern News.
Ant Group built Ling-3.0-Flash from scratch using a "native hybrid-linear attention" design. It blends two types of layers: KDA (Kimi Delta Attention) and MLA layers at a 5-to-1 ratio. KDA handles long-term efficiency, while MLA layers keep strong memory of context, according to National Post.
This hybrid approach lets the model handle very long pieces of text without slowing down. It is designed specifically for production-grade AI agent workflows — real business tasks, not just lab tests. The goal is a "superior balance" between speed and accuracy, according to Cochrane Times Post.
Ant Group is making Ling-3.0-Flash easy to try. The model is live on OpenRouter and Vercel AI Gateway right now. Developers can use it for free through August 3, 2026, according to Ontario Farmer.
That free window is a direct play for developer mindshare. By removing cost as a barrier, Ant Group is betting teams will build with Ling-3.0-Flash and stick with it. It follows a growing trend of Chinese AI labs offering aggressive pricing to win global users, according to Woodstock Sentinel Review.
Ling-3.0-Flash arrives as AI labs race to build smarter models without ballooning costs. Smaller active parameter counts mean cheaper inference — the cost to run a model on real queries. A model that beats larger rivals while activating only 5.1 billion parameters is a meaningful leap, according to Goderich Signal Star.
Ant Group positions Ling-3.0-Flash as a "high-speed execution node" for AI agents — software that completes multi-step tasks autonomously. If the benchmark claims hold up in real-world use, the model could pressure rivals like OpenAI and Google to justify the cost of their far larger systems, according to National Post.
Publishers
8
Articles
8
Reach
8