Spent a bit of time tuning NCCL for TP=4 GLM 5.3 Flash on 4x DGX Sparks. Bumped to…
This is a dev post classified by Jev as Hosting & infra (a tutorial), kept by the Dev Radar because it carries real work, not commentary.
Spent a bit of time tuning NCCL for TP=4 GLM 5.3 Flash on 4x DGX Sparks. Bumped to latest vLLM nightly, enabled _both_ ConnectX ports at 200gb. Might have one of the fastest DGX Spark deploys out there too. ~60 tok/s for prose, ~88 tok/s for code and >3000 tok/s for prefill.
Posted by Matt Mastracci (8.7k followers) 1 h ago · 0 likes · 172 views · view the original post on X. Kept by the Dev Radar as Hosting & infra.
More dev work like this
- #RHEL is now supported by Robot Operating System (ROS2) as a tier-1 platform, bringing… — @RedHat
- Agent Runners picks these up along with AI Gateway, so an agent building in Netlify can… — @Netlify
- We got @DeepSeek-V4.1-Flash running in EXL3 on 4× Sparks! — @cfontes
- Introducing the prerelease of Amazon Bedrock AgentCore runtime as a compute provider for… — @temporalio
- every open RL framework handwaves the sandbox layer like it's someone else's problem. — @arjunkocher
- #CloudflareConnect 2026 is bringing together the engineers, architects, and leaders… — @Cloudflare
- Cloud APIs don’t need a custom SDK for every query. — @DanKornas
- Google just open-sourced a Kubernetes-like orchestrator for running AI agents. — @twtayaan
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 22.6k posts from 5k X accounts over the last 21 days, 2.6k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-23 15:21 UTC. Full method.