Dev Radar
Support
LiveUpdated 2026-09-22 08:21 UTC

“MiMo-V2.6 Scaling Reinforcement Learning Towards Self-Improvement”

“MiMo-V2.6 Scaling Reinforcement Learning Towards Self-Improvement” Xiaomi MiMo just completed their largest RL…

This is a dev post classified by Jev as AI dev tools (a launch), kept by the Dev Radar because it carries real work, not commentary.

“MiMo-V2.6 Scaling Reinforcement Learning Towards Self-Improvement” Xiaomi MiMo just completed their largest RL scaling run so far, openly. And it came along with a technical report, outlining how they've scaled RL across compute, environments, and grading. tl;dr The run costs $2.6M for Pro and $0.9M for Flash, with ~44% spent on rollouts, 41-44% on training, and 14% on grading. Each RL step uses 1,568 prompts x 16 rollouts, producing ~25K trajectories and 2.7-3.7B tokens, with contexts up to 1M tokens. New techniques include groupwise agentic grading, which compares successful trajecto

Posted by alphaXiv (56.1k followers) 3 h ago · 49 likes · 2.5k views · view the original post on X. Kept by the Dev Radar as AI dev tools. Tools mentioned: alphaXiv.

More dev work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 20.7k posts from 4.9k X accounts over the last 21 days, 2.4k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-22 08:21 UTC. Full method.