Qwen3.8-Flash-Next EXL3 just got a BIG one-Spark update. 🔥
This is a dev post classified by Jev as Hosting & infra (a tool drop), kept by the Dev Radar because it carries real work, not commentary.
Qwen3.8-Flash-Next EXL3 just got a BIG one-Spark update. 🔥 58.8 tok/s single-stream through native ExLlamaV3.🚀 157.6 tok/s aggregate through vLLM across 8 streams.🤯 FULL 262,144 context on ONE DGX Spark. 4.05 bpw EXL3 pack. A much better serving envelope. 𝗕𝗘𝗙𝗢𝗥𝗘 → 𝗡𝗢𝗪 Previous public headline: 47.6 tok/s greedy p50 64K configured context MTP k=2 Concurrency not characterized Now: Native ExLlamaV3: 58.8 tok/s single stream vLLM + vllm-exl3: ~50–55 tok/s single stream 155.6 tok/s @ 4 streams 157.6 tok/s @ 8 streams Configured context: 64K → 262,144 KV pool: 416,163 tokens at the default 2
Posted by Cruz (2k followers) 4 days ago · 158 likes · 36.1k views · view the original post on X. Kept by the Dev Radar as Hosting & infra. Tools mentioned: qwen3.8-flash-next-exl3-dgx-spark-recipe, Qwen EXL3 benchmark.
More dev work like this
- Last chance: 4 days left to get your ticket to @WeAreDevs! — @Docker
- #MachineLearning with #AmazonSageMaker Cookbook! #BigData #Analytics #DataScience #AI… — @gp_pulipaka
- GSP644: Build a Serverless App with Cloud Run that Creats PDF Files 📄☁️ — @orbitofops
- One cluster. Multiple workloads. 🔥 — @k8sAMD
- i need a usa vpn like 3 times per year so paying for any vpn monthly/annually doesnt… — @thekitze
- The OpenInfra community in East Africa is expanding with the launch of the OpenInfra… — @openinfradev
- Kubernetes Pod Anti-Affinity for better replica distribution — @twtayaan
- Just 3 days to go! The stage is set for #ApsaraConference2026. — @alibaba_cloud
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 12.2k posts from 4.7k X accounts over the last 21 days, 1.4k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:39 UTC. Full method.