Dev Radar
Support
LiveUpdated 2026-09-19 18:06 UTC

glm-5.3-flash-exl3-k2-dgx-spark-recipe

glm-5.3-flash-exl3-k2-dgx-spark-recipe — vLLM recipe: GLM-5.3-Flash EXL3 K2 on one DGX Spark GB10. Native MTP k=2. Measured…

glm-5.3-flash-exl3-k2-dgx-spark-recipe is vLLM recipe: GLM-5.3-Flash EXL3 K2 on one DGX Spark GB10. Native MTP k=2. Measured tok/s. - vcruz305/GLM-5.3-Flash-EXL3-K2-DGX-Spark-recipe. It is ranked #1347 on the Dev Radar, in Other, first seen 20 days ago and shared in 4 posts (36.3k views).

Visit github.com

GitHub - vcruz305/GLM-5.3-Flash-EXL3-K2-DGX-Spark-recipe: vLLM recipe: GLM-5.3-Flash EXL3 K2 on one DGX Spark GB10. Native MTP k=2. Measured tok/s. · GitHub Skip to content Navigation Menu Sign in Appearance settings Search / Sign in Sign up Appearance settings You signed in with another tab or window. Reload to refresh your session. You signed out in another tab or window. Reload to refresh your session. You switched accounts on another tab or window. Reload to refresh your session. Dismiss alert {{ message }} vcruz305 / GLM-5.3-Flash-EXL3-K2-DGX-Spark-recipe Public Notifications You must be signed in to change notification settings Fork 2…

What people said about glm-5.3-flash-exl3-k2-dgx-spark-recipe on X

vllm-exl3 v0.3.0 is LIVE with custom native CUDA kernels for 2-bit EXL3 on NVIDIA DGX Spark GB10. GLM-5.3-Flash-EXL3-K2 jumped from 16.9 → 24.6 tok/s average single-stream decode, a +45.6% gain. Coding hit 27.6 tok/s, +85.6%. 🚀 The previous ExLlamaV3-backed path inside vLLM was leaving a lot of GB10 bandwidth on the…

@ViC305, 16 days ago · 128 likes · see the post

GLM-5.3-Flash EXL3 now cold-prefills 258,048 tokens on ONE DGX Spark. 258,048 tokens → HTTP 200 in 427 seconds. Memory stays flat. CUDA graphs stay ON. Native MTP k=2 also completes at 258K in 464 seconds. 🔥 The >163K long-context hang is fixed. 𝗧𝗪𝗢 𝗗𝗜𝗙𝗙𝗘𝗥𝗘𝗡𝗧 𝗣𝗥𝗢𝗕𝗟𝗘𝗠𝗦 What looked like one wall…

@ViC305, 17 days ago · 98 likes · see the post

GLM-5.3-Flash EXL3 K2 packs a 320B / 18B-active model into 91.017 GiB and runs it on ONE DGX Spark. Native MTP at 64K averaged 17.29 tok/s across 4 workloads and reached 20.60 tok/s on structured output. 𝗤𝗨𝗔𝗟𝗜𝗧𝗬-𝗙𝗜𝗥𝗦𝗧 𝗞𝟮 This is not “put the entire model at 2-bit.” I applied EXL3 K2 MCG trellis…

@ViC305, 20 days ago · 61 likes · see the post

Alternatives to glm-5.3-flash-exl3-k2-dgx-spark-recipe

glm-5.3-flash-exl3-k2-dgx-spark-recipe in numbers

FAQ

What is glm-5.3-flash-exl3-k2-dgx-spark-recipe?

vLLM recipe: GLM-5.3-Flash EXL3 K2 on one DGX Spark GB10. Native MTP k=2. Measured tok/s. - vcruz305/GLM-5.3-Flash-EXL3-K2-DGX-Spark-recipe It was first shared on X 20 days ago and is ranked #1347 on the Dev Radar.

Is glm-5.3-flash-exl3-k2-dgx-spark-recipe free?

It is open source.

Who shared glm-5.3-flash-exl3-k2-dgx-spark-recipe?

1 account on X, including @ViC305, in 4 posts totalling 36.3k views.

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 12.2k posts from 4.7k X accounts over the last 21 days, 1.4k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:06 UTC. Full method.