Dev Radar
Support
LiveUpdated 2026-09-19 18:49 UTC

qwen3.8-flash-next-exl3-dgx-spark-recipe

qwen3.8-flash-next-exl3-dgx-spark-recipe — Serve Qwen3.8-Flash-Next (turboderp ExLlamaV3 pack) on one NVIDIA DGX Spark with vLLM +…

qwen3.8-flash-next-exl3-dgx-spark-recipe is Serve Qwen3.8-Flash-Next (turboderp ExLlamaV3 pack) on one NVIDIA DGX Spark with vLLM + vllm-exl3 - vcruz305/Qwen3.8-Flash-Next-EXL3-DGX-Spark-recipe. It is ranked #574 on the Dev Radar, in Hosting & infra, first seen 12 days ago and shared in 6 posts (96k views).

Visit github.com

GitHub - vcruz305/Qwen3.8-Flash-Next-EXL3-DGX-Spark-recipe: Serve Qwen3.8-Flash-Next (turboderp ExLlamaV3 pack) on one NVIDIA DGX Spark with vLLM + vllm-exl3 · GitHub Skip to content Navigation Menu Sign in Appearance settings Search / Sign in Sign up Appearance settings You signed in with another tab or window. Reload to refresh your session. You signed out in another tab or window. Reload to refresh your session. You switched accounts on another tab or window. Reload to refresh your session. Dismiss alert {{ message }} vcruz305 / Qwen3.8-Flash-Next-EXL3-DGX-Spark-recipe Public Notifications You must be signed in to change notification…

What people said about qwen3.8-flash-next-exl3-dgx-spark-recipe on X

Qwen3.8-Flash-Next -> 79.95 tok/s. 🔥 ONE DGX Spark. Full 262K cache configured. 🚀 Qwen3.8-Flash-Next EXL3 just got another major update. New measured default: MTP ndt=5 DSpark, dc=0.6 8-bit KV 262,144-token cache At an actual 240K-token prompt: 72.0 tok/s decode ~1,150 tok/s prefill Exact needle retrieval 𝗙𝗣𝟭𝟲…

@ViC305, 1 days ago · 263 likes · see the post

Qwen3.8-Flash-Next EXL3 just got a BIG one-Spark update. 🔥 58.8 tok/s single-stream through native ExLlamaV3.🚀 157.6 tok/s aggregate through vLLM across 8 streams.🤯 FULL 262,144 context on ONE DGX Spark. 4.05 bpw EXL3 pack. A much better serving envelope. 𝗕𝗘𝗙𝗢𝗥𝗘 → 𝗡𝗢𝗪 Previous public headline: 47.6 tok/s…

@ViC305, 4 days ago · 158 likes · see the post

Qwen3.8-Flash-Next now reaches ~43 tok/s after a 122,902-token prompt on ONE DGX Spark. ⚡🚀 MTP k=2 won my draft-depth sweep, with +42.5% mean decode over no draft. The PLE table stays fully on-device. I promised the deeper MTP tests. Here are the results, and now you can explore them in an interactive benchmark page…

@ViC305, 11 days ago · 58 likes · see the post

Alternatives to qwen3.8-flash-next-exl3-dgx-spark-recipe

qwen3.8-flash-next-exl3-dgx-spark-recipe in numbers

FAQ

What is qwen3.8-flash-next-exl3-dgx-spark-recipe?

Serve Qwen3.8-Flash-Next (turboderp ExLlamaV3 pack) on one NVIDIA DGX Spark with vLLM + vllm-exl3 - vcruz305/Qwen3.8-Flash-Next-EXL3-DGX-Spark-recipe It was first shared on X 12 days ago and is ranked #574 on the Dev Radar.

Is qwen3.8-flash-next-exl3-dgx-spark-recipe free?

It is open source.

Who shared qwen3.8-flash-next-exl3-dgx-spark-recipe?

1 account on X, including @ViC305, in 6 posts totalling 96k views.

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 12.3k posts from 4.7k X accounts over the last 21 days, 1.4k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:49 UTC. Full method.