Xiaomi’s MiMo-V2.6 Tackles Agents, Cyber And 3D Worlds In Fully Open-Source Release
In Brief
Xiaomi open-sources MiMo-V2.6-Pro and Flash, scaling RL across 750K trajectories in six days to claim the top open-source spot on Artificial Analysis, with unchanged API pricing and broad agentic, coding and cyber capabilities.

Xiaomi has released and open-sourced the MiMo-V2.6 series, its latest generation of natively omnimodal AI models, positioning the release as a key step in scaling reinforcement learning on verifiable, complex tasks. The lineup comprises two models — MiMo-V2.6-Pro, described as the company’s most capable model to date, and the more cost-efficient MiMo-V2.6-Flash — alongside a Pro-UltraSpeed variant offering output speeds up to 20 times faster.
On the Artificial Analysis Intelligence Index (v4.3, September 2026), MiMo-V2.6-Pro scored 46.32, surpassing Kimi K3 and Qwen3.8 Max to become the highest-rated open-source model. Notably, the new series retains the API pricing of its V2.5 predecessor, pushing the intelligence-versus-cost Pareto frontier outward at unchanged cost.
Benchmark results place the models competitively across disciplines. On DeepSWE v1.1, a long-horizon software engineering test, MiMo-V2.6-Pro scored 71.9, trailing DeepSeek V4.1 Flash (74.2), Claude Opus 5 and GPT 6 Astra (74.0 each), while MiMo-V2.6-Flash reached 67.9 — a dramatic improvement over MiMo-V2.5-Pro’s 19.0. The series also posted strong results in general agentic workflows, leading Automation Bench v1.0.6 among frontier models except DeepSeek V4.1 Flash, and scoring 94.0 and 95.1 on the cyber-focused CyberGym benchmark.
Reinforcement Learning at Scale, Streamed in Public
The technical foundation of the release is a scaled RL training run that Xiaomi streamed live. In under six days, each model completed 30 RL steps over roughly 750,000 trajectories, at costs of approximately $0.85 million (Flash) and $2.62 million (Pro). Average pass rates on training tasks rose by 25% and 12% in relative terms, and held-out benchmark gains were substantial: DeepSWE v1.1 scores improved by roughly 17 points for Flash (48.8 to 65.68) and 14 points for Pro (58.4 to 72.57), with RL proving sample-efficient and generalizing beyond the training distribution.
Scaling was pursued along three axes: larger batches on a fully asynchronous architecture (1,568 samples per update, up to 1 million context length, 3.5–3.7 billion tokens per step); a multi-task suite spanning coding, general agents, visual and cyber domains; and increased grader compute using relative comparisons for more precise reward signals. The team also froze the router to suppress training drift and built a defense against reward hacking combining reward design, adversarial evaluation, anomaly detection and cross-verification.
Beyond benchmarks, Xiaomi highlights applied capabilities under what it calls “Vibe World” — extending natural-language programming from software to interactive 3D worlds, game development, Blender-based 3D modeling, closed-loop robotic arm control, frontend and presentation design, and end-to-end video and music production. In research settings, MiMo-V2.6-Pro contributed to computational screening of novel MOF materials for PFAS adsorption and assisted in formalizing the Li–Yorke theorem in Lean 4, producing over 6,000 lines of kernel-verified code without Lean-specific post-training.
The models are available through the MiMo API Platform, AI Studio, MiMo Desktop and OpenRouter, with pricing unchanged from V2.5. Xiaomi is open-sourcing the full technical report, training environments and RL code for reproduction and further research.
Disclaimer
In line with the Trust Project guidelines, please note that the information provided on this page is not intended to be and should not be interpreted as legal, tax, investment, financial, or any other form of advice. It is important to only invest what you can afford to lose and to seek independent financial advice if you have any doubts. For further information, we suggest referring to the terms and conditions as well as the help and support pages provided by the issuer or advertiser. MetaversePost is committed to accurate, unbiased reporting, but market conditions are subject to change without notice.
About The Author
Alisa, a dedicated journalist at the MPost, specializes in crypto, AI, investments, and the expansive realm of Web3. With a keen eye for emerging trends and technologies, she delivers comprehensive coverage to inform and engage readers in the ever-evolving landscape of digital finance.
More articles
Alisa, a dedicated journalist at the MPost, specializes in crypto, AI, investments, and the expansive realm of Web3. With a keen eye for emerging trends and technologies, she delivers comprehensive coverage to inform and engage readers in the ever-evolving landscape of digital finance.



