Dev Radar
Support
LiveUpdated 2026-09-21 06:12 UTC

My First Model On @huggingface Qwen 3.8 27B Unleashed HIT 105,666 Downloads 🔥

My First Model On @huggingface Qwen 3.8 27B Unleashed HIT 105,666 Downloads 🔥 Qwen3.8-27B Unleashed UD-Q3_K_XL hits…

This is a dev post classified by Jev as AI dev tools (a launch), kept by the Dev Radar because it carries real work, not commentary.

My First Model On @huggingface Qwen 3.8 27B Unleashed HIT 105,666 Downloads 🔥 Qwen3.8-27B Unleashed UD-Q3_K_XL hits near 100 tk/s @ 250k ctx on 1 x 4090. - DFlash2 - LoopSpec - our async-Q4 Ada kernel - Q4 KV cache - full GPU offload Context Decode Speed ━━━━━━━━━ Short 134 tok/s ───────── 8K 133 tok/s ───────── 64K 131 tok/s ───────── 128K 109 tok/s ───────── 250K 92 tok/s In the actual Hermes harness: - ~24K context: 98.5 tok/s - ~250K context: 80 tok/s - Completed the full 250K three-tool workflow correctly - Preserved the full conversation history - No retrieval or compression

Posted by Eric ⚡️ Building... (10k followers) 3 days ago · 25 likes · 1.5k views · view the original post on X. Kept by the Dev Radar as AI dev tools.

More dev work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 18.9k posts from 4.9k X accounts over the last 21 days, 2.1k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-21 06:12 UTC. Full method.