My RTX 4090 just hit 181 tok/s on Qwen 3.8 27B model. O Monday I thought 140 was the…
This is a dev post classified by Jev as AI dev tools (a tool drop), kept by the Dev Radar because it carries real work, not commentary.
My RTX 4090 just hit 181 tok/s on Qwen 3.8 27B model. O Monday I thought 140 was the ceiling. I was wrong. @mr_r0b0t built a different engine. I wanted to push the card again to see if I can squeeze more out of it. The headline number moved 30%, the like-for-like number moved 4%, and the gap between them is the actual picture. Here is my full breakdown of whether the recipe has legs. A DIFFERENT ENGINE Two pieces. A better quant and a better drafter. The quant is EXL3 at 4.0 bits per weight, 15.4 GiB for the whole 27B model. The drafter is DFlash2, a block-diffusion model that proposes u
Posted by Yume_X (2.1k followers) 2 days ago · 80 likes · 5.9k views · view the original post on X. Kept by the Dev Radar as AI dev tools.
More dev work like this
- I think we’re looking for this — @vaibcode
- When a hard coding decision needs a second opinion, don’t settle for one model. — @DanKornas
- Banger paper from MIT and Sakana AI. — @dair_ai
- Nautilo agents use memory like memory competition champs. They never forget and they… — @Dan_Jeffries1
- AI agents need guardrails before they reach your tools. — @DanKornas
- 20-30 agents are only useful when you stop babysitting every one of them. — @catmanyau
- Okay, everyone wants us to give the unbiased facts. — @Teknium
- A single RTX 3090 hit 381 tok/s on Qwen3.8-27B. — @0x0SojalSec
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 14.4k posts from 4.7k X accounts over the last 21 days, 1.6k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 21:10 UTC. Full method.