Dev Radar
Support
LiveUpdated 2026-09-20 22:16 UTC

glm-5.3-flash-nvfp4-1m-kv-4x-dgx-spark

glm-5.3-flash-nvfp4-1m-kv-4x-dgx-spark — GLM-5.3-Flash (320B MoE) at TP4 across four DGX Sparks, same day as the model drop: 36…

glm-5.3-flash-nvfp4-1m-kv-4x-dgx-spark is GLM-5.3-Flash (320B MoE) at TP4 across four DGX Sparks, same day as the model drop: 36 tok/s, 1.26M-token FP8 KV pool, 262K context, MTP spec decode. First TP4 glm5_next outside B200 hardware. Full.... It is ranked #261 on the Dev Radar, in Other, first seen 6 h ago and shared in 7 posts (13.4k views).

Visit github.com

GitHub - tonyd2wild/GLM-5.3-Flash-NVFP4-1M-KV-4x-DGX-Spark: GLM-5.3-Flash (320B MoE) at TP4 across four DGX Sparks, same day as the model drop: 36 tok/s, 1.26M-token FP8 KV pool, 262K context, MTP spec decode. First TP4 glm5_next outside B200 hardware. Full patched-image recipe + the GB10 memory study. · GitHub Skip to content Navigation Menu Sign in Appearance settings Search / Sign in Sign up Appearance settings You signed in with another tab or window. Reload to refresh your session. You signed out in another tab or window. Reload to refresh your session. You switched accounts on another tab or window. Reload to refresh your session.…

What people said about glm-5.3-flash-nvfp4-1m-kv-4x-dgx-spark on X

🚀 New GLM-5.3-Flash default on 4x DGX Spark (TP4) Found 18 GiB of attention + MLP weights still sitting in bf16, read on every single step. Quantized them to NVFP4. ⚡ Per-step time: 67.6ms → 57ms 🧠 500K ctx, fp8 KV, 3.53M token pool 📈 vs our previous GLM lane: Structure +30% | Math +29% | Prose +27% Counting +21%…

@Tech2Wild, 6 h ago · 26 likes · see the post

🚀 Now running the official NVIDIA quant on 4x DGX Spark (TP4) Rebuilt the NVFP4 attention work on nvidia/GLM-5.3-Flash-NVFP4 instead of LibertAI. Official weights, same speeds. 📊 vs our previous lane: 9 of 9 categories within measurement noise ⚡ Concurrency actually better: +4.8% at C12 → +10.0% at C32 🧠 500K ctx,…

@Tech2Wild, 2 h ago · 10 likes · see the post

Right Now I will swapping the Redhat with the OFFICIAL NVIDIA NVFP4 Checkpoint and see how it holds up. Right now I have not had ANY issues with RedHat and the checkpoint size is smaller so MORE KV Pool.

@Tech2Wild, 6 h ago · 11 likes · see the post

Alternatives to glm-5.3-flash-nvfp4-1m-kv-4x-dgx-spark

glm-5.3-flash-nvfp4-1m-kv-4x-dgx-spark in numbers

FAQ

What is glm-5.3-flash-nvfp4-1m-kv-4x-dgx-spark?

GLM-5.3-Flash (320B MoE) at TP4 across four DGX Sparks, same day as the model drop: 36 tok/s, 1.26M-token FP8 KV pool, 262K context, MTP spec decode. First TP4 glm5_next outside B200 hardware. Full... It was first shared on X 6 h ago and is ranked #261 on the Dev Radar.

Is glm-5.3-flash-nvfp4-1m-kv-4x-dgx-spark free?

It is open source.

Who shared glm-5.3-flash-nvfp4-1m-kv-4x-dgx-spark?

1 account on X, including @Tech2Wild, in 7 posts totalling 13.4k views.

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 17k posts from 4.9k X accounts over the last 21 days, 1.9k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-20 22:16 UTC. Full method.