glm-5.3-flash-nvfp4-1m-kv-4x-dgx-spark
glm-5.3-flash-nvfp4-1m-kv-4x-dgx-spark is GLM-5.3-Flash (320B MoE) at TP4 across four DGX Sparks, same day as the model drop: 36 tok/s, 1.26M-token FP8 KV pool, 262K context, MTP spec decode. First TP4 glm5_next outside B200 hardware. Full.... It is ranked #261 on the Dev Radar, in Other, first seen 6 h ago and shared in 7 posts (13.4k views).
GitHub - tonyd2wild/GLM-5.3-Flash-NVFP4-1M-KV-4x-DGX-Spark: GLM-5.3-Flash (320B MoE) at TP4 across four DGX Sparks, same day as the model drop: 36 tok/s, 1.26M-token FP8 KV pool, 262K context, MTP spec decode. First TP4 glm5_next outside B200 hardware. Full patched-image recipe + the GB10 memory study. · GitHub Skip to content Navigation Menu Sign in Appearance settings Search / Sign in Sign up Appearance settings You signed in with another tab or window. Reload to refresh your session. You signed out in another tab or window. Reload to refresh your session. You switched accounts on another tab or window. Reload to refresh your session.…
What people said about glm-5.3-flash-nvfp4-1m-kv-4x-dgx-spark on X
🚀 New GLM-5.3-Flash default on 4x DGX Spark (TP4) Found 18 GiB of attention + MLP weights still sitting in bf16, read on every single step. Quantized them to NVFP4. ⚡ Per-step time: 67.6ms → 57ms 🧠 500K ctx, fp8 KV, 3.53M token pool 📈 vs our previous GLM lane: Structure +30% | Math +29% | Prose +27% Counting +21%…
— @Tech2Wild, 6 h ago · 26 likes · see the post
🚀 Now running the official NVIDIA quant on 4x DGX Spark (TP4) Rebuilt the NVFP4 attention work on nvidia/GLM-5.3-Flash-NVFP4 instead of LibertAI. Official weights, same speeds. 📊 vs our previous lane: 9 of 9 categories within measurement noise ⚡ Concurrency actually better: +4.8% at C12 → +10.0% at C32 🧠 500K ctx,…
— @Tech2Wild, 2 h ago · 10 likes · see the post
Right Now I will swapping the Redhat with the OFFICIAL NVIDIA NVFP4 Checkpoint and see how it holds up. Right now I have not had ANY issues with RedHat and the checkpoint size is smaller so MORE KV Pool.
— @Tech2Wild, 6 h ago · 11 likes · see the post
Alternatives to glm-5.3-flash-nvfp4-1m-kv-4x-dgx-spark
- Spark-X2.5 — Model Scope
- mimo-v2.6 RL — Streaming the run
- qwen-image-2.1 — Qwen's most powerful open-source image generation model - QwenLM/Qwen-Image-2.1
- Ling-3.0-flash-Fin — Qwen-Drive-1.0-4B is a vision-language foundation model that handles 3D perception, driving VQA, and motion planning…
- mcdma — Mac Studio A — Gateway. The only Studio with a Mellanox ConnectX‑4 on the MikroTik fabric (via a Thunderbolt 5…
- Cactus Compute — It runs on mobiles, wearables, smart home devices, small robots and microcontrollers, with prebuilt engines for macOS,…
glm-5.3-flash-nvfp4-1m-kv-4x-dgx-spark in numbers
- Rank on the Dev Radar: #261 of 1886
- Shared in 7 posts by 1 account: @Tech2Wild
- 13.4k views on those posts
- First seen 6 h ago, last shared 2 h ago
- Pricing seen by Jev: open source
- Market: Other
FAQ
What is glm-5.3-flash-nvfp4-1m-kv-4x-dgx-spark?
GLM-5.3-Flash (320B MoE) at TP4 across four DGX Sparks, same day as the model drop: 36 tok/s, 1.26M-token FP8 KV pool, 262K context, MTP spec decode. First TP4 glm5_next outside B200 hardware. Full... It was first shared on X 6 h ago and is ranked #261 on the Dev Radar.
Is glm-5.3-flash-nvfp4-1m-kv-4x-dgx-spark free?
It is open source.
Who shared glm-5.3-flash-nvfp4-1m-kv-4x-dgx-spark?
1 account on X, including @Tech2Wild, in 7 posts totalling 13.4k views.
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 17k posts from 4.9k X accounts over the last 21 days, 1.9k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-20 22:16 UTC. Full method.