Dev Radar
Support
LiveUpdated 2026-09-23 15:57 UTC

@XiaomiMiMo is already working on next MiMo-V3 architecture!

This is awesome! @XiaomiMiMo is already working on next MiMo-V3 architecture! So there might not be anymore v2 models,…

This is a dev post classified by Jev as Other (a launch), kept by the Dev Radar because it carries real work, not commentary.

This is awesome! @XiaomiMiMo is already working on next MiMo-V3 architecture! So there might not be anymore v2 models, MiMo-v2.6 might be the last one. Their motivation is interesting. Their paper HySparse2: Hybrid Sparse Attention with Two-Level KV Sharing has it very clear. Reading through paper, what i learned is this will result in much improved - prompt processing speed - TTFT - prefill energy usage - GPU requirements for prefill servers Cause prompt doesn't necessarily need to travel through the entire network. I really like these kind of research going on to making inference effic

Posted by AJ (7.7k followers) 1 h ago · 8 likes · 627 views · view the original post on X. Kept by the Dev Radar as Other.

More dev work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 22.6k posts from 5k X accounts over the last 21 days, 2.6k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-23 15:57 UTC. Full method.