qwen3.8-flash-next-exl3-dgx-spark-recipe
qwen3.8-flash-next-exl3-dgx-spark-recipe is Serve Qwen3.8-Flash-Next (turboderp ExLlamaV3 pack) on one NVIDIA DGX Spark with vLLM + vllm-exl3 - vcruz305/Qwen3.8-Flash-Next-EXL3-DGX-Spark-recipe. It is ranked #574 on the Dev Radar, in Hosting & infra, first seen 12 days ago and shared in 6 posts (96k views).
GitHub - vcruz305/Qwen3.8-Flash-Next-EXL3-DGX-Spark-recipe: Serve Qwen3.8-Flash-Next (turboderp ExLlamaV3 pack) on one NVIDIA DGX Spark with vLLM + vllm-exl3 · GitHub Skip to content Navigation Menu Sign in Appearance settings Search / Sign in Sign up Appearance settings You signed in with another tab or window. Reload to refresh your session. You signed out in another tab or window. Reload to refresh your session. You switched accounts on another tab or window. Reload to refresh your session. Dismiss alert {{ message }} vcruz305 / Qwen3.8-Flash-Next-EXL3-DGX-Spark-recipe Public Notifications You must be signed in to change notification…
What people said about qwen3.8-flash-next-exl3-dgx-spark-recipe on X
Qwen3.8-Flash-Next -> 79.95 tok/s. 🔥 ONE DGX Spark. Full 262K cache configured. 🚀 Qwen3.8-Flash-Next EXL3 just got another major update. New measured default: MTP ndt=5 DSpark, dc=0.6 8-bit KV 262,144-token cache At an actual 240K-token prompt: 72.0 tok/s decode ~1,150 tok/s prefill Exact needle retrieval 𝗙𝗣𝟭𝟲…
— @ViC305, 1 days ago · 263 likes · see the post
Qwen3.8-Flash-Next EXL3 just got a BIG one-Spark update. 🔥 58.8 tok/s single-stream through native ExLlamaV3.🚀 157.6 tok/s aggregate through vLLM across 8 streams.🤯 FULL 262,144 context on ONE DGX Spark. 4.05 bpw EXL3 pack. A much better serving envelope. 𝗕𝗘𝗙𝗢𝗥𝗘 → 𝗡𝗢𝗪 Previous public headline: 47.6 tok/s…
— @ViC305, 4 days ago · 158 likes · see the post
Qwen3.8-Flash-Next now reaches ~43 tok/s after a 122,902-token prompt on ONE DGX Spark. ⚡🚀 MTP k=2 won my draft-depth sweep, with +42.5% mean decode over no draft. The PLE table stays fully on-device. I promised the deeper MTP tests. Here are the results, and now you can explore them in an interactive benchmark page…
— @ViC305, 11 days ago · 58 likes · see the post
Alternatives to qwen3.8-flash-next-exl3-dgx-spark-recipe
- boat by ASCII — ascii box is now boat ( )
- Railway — Bring context from your work and personal accounts into the same conversation. Connect your accounts in the plugin…
- Nebius Token Factory — Start building
- recipes.vllm.ai — DeepSeek's first experimental multimodal V4 model — the V4-Flash MoE backbone plus a 32-layer vision tower, 1M…
- click.alibabacloud.com — Model Studio
- serverkit — ServerKit is a lightweight, modern server control panel for managing web applications, databases, and services on your…
qwen3.8-flash-next-exl3-dgx-spark-recipe in numbers
- Rank on the Dev Radar: #574 of 1355
- Shared in 6 posts by 1 account: @ViC305
- 96k views on those posts
- First seen 12 days ago, last shared 1 days ago
- Pricing seen by Jev: open source
- Market: Hosting & infra
FAQ
What is qwen3.8-flash-next-exl3-dgx-spark-recipe?
Serve Qwen3.8-Flash-Next (turboderp ExLlamaV3 pack) on one NVIDIA DGX Spark with vLLM + vllm-exl3 - vcruz305/Qwen3.8-Flash-Next-EXL3-DGX-Spark-recipe It was first shared on X 12 days ago and is ranked #574 on the Dev Radar.
Is qwen3.8-flash-next-exl3-dgx-spark-recipe free?
It is open source.
Who shared qwen3.8-flash-next-exl3-dgx-spark-recipe?
1 account on X, including @ViC305, in 6 posts totalling 96k views.
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 12.3k posts from 4.7k X accounts over the last 21 days, 1.4k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:49 UTC. Full method.