Cooked up a GLM-5.3-Flash-EXL3–2.32bpw for the 128GB crowd, perfect for your @NVIDIAAI…
This is a dev post classified by Jev as Other (a tool drop), kept by the Dev Radar because it carries real work, not commentary.
Cooked up a GLM-5.3-Flash-EXL3–2.32bpw for the 128GB crowd, perfect for your @NVIDIAAI DGX Spark ⚡️ Testing it now, looks like quantized DFlash2-EXL3 is working well! 58.4 tok/s isn’t to be taken as any sort of eval here but it is a very positive signal 🤩
Posted by mr_r0b0t (10.6k followers) 17 h ago · 78 likes · 3.9k views · view the original post on X. Kept by the Dev Radar as Other.
More dev work like this
- Nice: banked codex reset for all of us. — @kimmonismus
- The #Mathematical #Programming - #FPGA. #BigData #Analytics #DataScience #AI… — @gp_pulipaka
- A Constraint Based Approach, ML. #BigData #Analytics #DataScience #AI #MachineLearning… — @gp_pulipaka
- Meet AliceAI-Foundation-80B-A3B-Base: a massive 80B parameter Mixture-of-Experts model… — @HuggingModels
- I am working on a new skill for Antigravity to prepare for Gemini 4, and I decided to… — @IamEmily2050
- There’s no way — @luckeyfaraday
- xAI's Grok Imagine Image 2.0 takes #4 on the Artificial Analysis Text to Image… — @ArtificialAnlys
- NVIDIA just dropped DeepSeek-V4.1-Flash-NVFP4, a quantized version of DeepSeek V4.1 that… — @HuggingModels
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 12.2k posts from 4.7k X accounts over the last 21 days, 1.3k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 17:33 UTC. Full method.