Dev Radar
Support
LiveUpdated 2026-09-19 21:10 UTC

My RTX 4090 just hit 181 tok/s on Qwen 3.8 27B model. O Monday I thought 140 was the…

My RTX 4090 just hit 181 tok/s on Qwen 3.8 27B model. O Monday I thought 140 was the ceiling. I was wrong. @mr_r0b0t…

This is a dev post classified by Jev as AI dev tools (a tool drop), kept by the Dev Radar because it carries real work, not commentary.

My RTX 4090 just hit 181 tok/s on Qwen 3.8 27B model. O Monday I thought 140 was the ceiling. I was wrong. @mr_r0b0t built a different engine. I wanted to push the card again to see if I can squeeze more out of it. The headline number moved 30%, the like-for-like number moved 4%, and the gap between them is the actual picture. Here is my full breakdown of whether the recipe has legs. A DIFFERENT ENGINE Two pieces. A better quant and a better drafter. The quant is EXL3 at 4.0 bits per weight, 15.4 GiB for the whole 27B model. The drafter is DFlash2, a block-diffusion model that proposes u

Posted by Yume_X (2.1k followers) 2 days ago · 80 likes · 5.9k views · view the original post on X. Kept by the Dev Radar as AI dev tools.

More dev work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 14.4k posts from 4.7k X accounts over the last 21 days, 1.6k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 21:10 UTC. Full method.