A single RTX 3090 hit 381 tok/s on Qwen3.8-27B.
This is a dev post classified by Jev as AI dev tools (a tool drop), kept by the Dev Radar because it carries real work, not commentary.
A single RTX 3090 hit 381 tok/s on Qwen3.8-27B. DFlash2 + lookup: 138 Optimized MTP: 114 context lookup so the model verifies 16 tokens at once when the answer already lives in the prompt. Longer verification & prefix cache: 381 tok/s The 381 number only appears when the model is reproducing or extracting from a long prompt. That’s the point. Prefix cache also drops TTFT from 22s to 0.56s on the second question against the same document. Best for RAG, document Q&A, and coding agents. Regular chat still lands around 133 tok/s. 25K-token context. https://x.com/TeksEdge/status/209059599011
Posted by Md Ismail Šojal 🕷️ (55.6k followers) 49 min ago · 6 likes · 621 views · view the original post on X. Kept by the Dev Radar as AI dev tools.
More dev work like this
- AI-written code gets risky when nobody checks the diffs. — @DanKornas
- The new "bakeoff" capability in Compound Engineering is so damn useful. Even on smaller… — @trevin
- Tried the new @zcode_ai update with dynamic workflows. — @Foshowithit
- Hermes Agent can now stream its THINKING into other AI apps. — @tonysimons_
- Hallmark has been installed 50k+ times! — @nutlope
- Your AI coding agents need a workflow, not more babysitting. — @DanKornas
- Register by October 7 and save >> https://bit.ly/4vkntyD — @linuxfoundation
- Nimble is able to process images! — @madiator
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 14.4k posts from 4.7k X accounts over the last 21 days, 1.6k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 19:58 UTC. Full method.