glm-5.3-flash-exl3-k2-dgx-spark-recipe
glm-5.3-flash-exl3-k2-dgx-spark-recipe is vLLM recipe: GLM-5.3-Flash EXL3 K2 on one DGX Spark GB10. Native MTP k=2. Measured tok/s. - vcruz305/GLM-5.3-Flash-EXL3-K2-DGX-Spark-recipe. It is ranked #1347 on the Dev Radar, in Other, first seen 20 days ago and shared in 4 posts (36.3k views).
GitHub - vcruz305/GLM-5.3-Flash-EXL3-K2-DGX-Spark-recipe: vLLM recipe: GLM-5.3-Flash EXL3 K2 on one DGX Spark GB10. Native MTP k=2. Measured tok/s. · GitHub Skip to content Navigation Menu Sign in Appearance settings Search / Sign in Sign up Appearance settings You signed in with another tab or window. Reload to refresh your session. You signed out in another tab or window. Reload to refresh your session. You switched accounts on another tab or window. Reload to refresh your session. Dismiss alert {{ message }} vcruz305 / GLM-5.3-Flash-EXL3-K2-DGX-Spark-recipe Public Notifications You must be signed in to change notification settings Fork 2…
What people said about glm-5.3-flash-exl3-k2-dgx-spark-recipe on X
vllm-exl3 v0.3.0 is LIVE with custom native CUDA kernels for 2-bit EXL3 on NVIDIA DGX Spark GB10. GLM-5.3-Flash-EXL3-K2 jumped from 16.9 → 24.6 tok/s average single-stream decode, a +45.6% gain. Coding hit 27.6 tok/s, +85.6%. 🚀 The previous ExLlamaV3-backed path inside vLLM was leaving a lot of GB10 bandwidth on the…
— @ViC305, 16 days ago · 128 likes · see the post
GLM-5.3-Flash EXL3 now cold-prefills 258,048 tokens on ONE DGX Spark. 258,048 tokens → HTTP 200 in 427 seconds. Memory stays flat. CUDA graphs stay ON. Native MTP k=2 also completes at 258K in 464 seconds. 🔥 The >163K long-context hang is fixed. 𝗧𝗪𝗢 𝗗𝗜𝗙𝗙𝗘𝗥𝗘𝗡𝗧 𝗣𝗥𝗢𝗕𝗟𝗘𝗠𝗦 What looked like one wall…
— @ViC305, 17 days ago · 98 likes · see the post
GLM-5.3-Flash EXL3 K2 packs a 320B / 18B-active model into 91.017 GiB and runs it on ONE DGX Spark. Native MTP at 64K averaged 17.29 tok/s across 4 workloads and reached 20.60 tok/s on structured output. 𝗤𝗨𝗔𝗟𝗜𝗧𝗬-𝗙𝗜𝗥𝗦𝗧 𝗞𝟮 This is not “put the entire model at 2-bit.” I applied EXL3 K2 MCG trellis…
— @ViC305, 20 days ago · 61 likes · see the post
Alternatives to glm-5.3-flash-exl3-k2-dgx-spark-recipe
- Ling-3.0-flash-Fin — Qwen-Drive-1.0-4B is a vision-language foundation model that handles 3D perception, driving VQA, and motion planning…
- bespoke-nimble-9b — We’re on a journey to advance and democratize artificial intelligence through open source and open science.
- Cactus Compute — It runs on mobiles, wearables, smart home devices, small robots and microcontrollers, with prebuilt engines for macOS,…
- orcabonsai-27b-uncensored — Open source
- Spark-X2.5 — 🤖 ModelScope
- grok.com —
glm-5.3-flash-exl3-k2-dgx-spark-recipe in numbers
- Rank on the Dev Radar: #1347 of 1353
- Shared in 4 posts by 1 account: @ViC305
- 36.3k views on those posts
- First seen 20 days ago, last shared 15 days ago
- Pricing seen by Jev: open source
- Market: Other
FAQ
What is glm-5.3-flash-exl3-k2-dgx-spark-recipe?
vLLM recipe: GLM-5.3-Flash EXL3 K2 on one DGX Spark GB10. Native MTP k=2. Measured tok/s. - vcruz305/GLM-5.3-Flash-EXL3-K2-DGX-Spark-recipe It was first shared on X 20 days ago and is ranked #1347 on the Dev Radar.
Is glm-5.3-flash-exl3-k2-dgx-spark-recipe free?
It is open source.
Who shared glm-5.3-flash-exl3-k2-dgx-spark-recipe?
1 account on X, including @ViC305, in 4 posts totalling 36.3k views.
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 12.2k posts from 4.7k X accounts over the last 21 days, 1.4k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:06 UTC. Full method.