Dev Radar
Support
LiveUpdated 2026-09-19 18:39 UTC

Qwen3.8-Flash-Next EXL3 just got a BIG one-Spark update. 🔥

Qwen3.8-Flash-Next EXL3 just got a BIG one-Spark update. 🔥 58.8 tok/s single-stream through native ExLlamaV3.🚀 157.6…

This is a dev post classified by Jev as Hosting & infra (a tool drop), kept by the Dev Radar because it carries real work, not commentary.

Qwen3.8-Flash-Next EXL3 just got a BIG one-Spark update. 🔥 58.8 tok/s single-stream through native ExLlamaV3.🚀 157.6 tok/s aggregate through vLLM across 8 streams.🤯 FULL 262,144 context on ONE DGX Spark. 4.05 bpw EXL3 pack. A much better serving envelope. 𝗕𝗘𝗙𝗢𝗥𝗘 → 𝗡𝗢𝗪 Previous public headline: 47.6 tok/s greedy p50 64K configured context MTP k=2 Concurrency not characterized Now: Native ExLlamaV3: 58.8 tok/s single stream vLLM + vllm-exl3: ~50–55 tok/s single stream 155.6 tok/s @ 4 streams 157.6 tok/s @ 8 streams Configured context: 64K → 262,144 KV pool: 416,163 tokens at the default 2

Posted by Cruz (2k followers) 4 days ago · 158 likes · 36.1k views · view the original post on X. Kept by the Dev Radar as Hosting & infra. Tools mentioned: qwen3.8-flash-next-exl3-dgx-spark-recipe, Qwen EXL3 benchmark.

More dev work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 12.2k posts from 4.7k X accounts over the last 21 days, 1.4k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:39 UTC. Full method.