Dev Radar
Support
LiveUpdated 2026-09-22 18:46 UTC

Just launched my usual Qwen 3.8 27B configs and, for a brief moment, couldn't understand…

Just launched my usual Qwen 3.8 27B configs and, for a brief moment, couldn't understand why I had so little VRAM in…Just launched my usual Qwen 3.8 27B configs and, for a brief moment, couldn't understand why I had so little VRAM in…

This is a dev post classified by Jev as AI dev tools (a launch), kept by the Dev Radar because it carries real work, not commentary.

Just launched my usual Qwen 3.8 27B configs and, for a brief moment, couldn't understand why I had so little VRAM in use, like multiple gigabytes less than usual. This was, of course, due to the latest change in BeeLlama v0.4.7 Preview, where independent draft batch sizing drastically reduces VRAM wasted on buffers, with no performance impact. Here's a breakdown of the savings for MTP and DFlash 2 on 2×3090, using Qwen 3.8 27B UD-Q6_K_M with a 256k KVarN6 KV cache.

Posted by Ivan Neustroev (273 followers) 1 h ago · 2 likes · 47 views · view the original post on X. Kept by the Dev Radar as AI dev tools.

More dev work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 21.4k posts from 5k X accounts over the last 21 days, 2.4k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-22 18:46 UTC. Full method.