skip to content
The Weighted Average

Wire

Ant opens 560B Ling-3.1-flash

Ant Group’s InclusionAI launched Ling-3.1-flash with about 560 billion total parameters and 25 billion active per token, positioning it for agent tasks, search, and office software. TechNode’s report says the model is designed for a 1 million-token context, but the two-week free trial currently caps requests at 256,000 tokens; InclusionAI plans to enable the larger window and release the model as open source after the trial. Builders should treat the sparse-activation/context combination as a model to benchmark when it opens, not assume the 1M limit or open weights exist today; the archive’s local-model size analysis shows why memory, not headline parameter count, decides edge feasibility.