Qwen3.8-2.4T on vLLM: a Pareto frontier spanning 5K total tokens/s/GPU at high…
This is a dev post classified by Jev as Hosting & infra (a tutorial), kept by the Dev Radar because it carries real work, not commentary.
Qwen3.8-2.4T on vLLM: a Pareto frontier spanning 5K total tokens/s/GPU at high throughput and 180 output tokens/s/user at low latency, across tuned PD configurations on @nvidia GB300 NVL72. Workload: 8K input / 1K output. Drawing on lessons from trial and error, we walk through the tuning decisions step by step: budget KV cache, benchmark prefill and decode separately, then choose topologies and MTP settings for each serving target. Deployment configs are included so you can reproduce the results. Great work from the @NVIDIAAI contributors and the vLLM community! Explore the frontier: https
Posted by vLLM (49.8k followers) 17 h ago · 69 likes · 6k views · view the original post on X. Kept by the Dev Radar as Hosting & infra.
More dev work like this
- At HUAWEI CONNECT, Huawei Deputy Chairman of the Board and Rotating Chairman David Wang… — @Huawei
- Join us on October 7th for an exclusive webinar on Cloudflare OS. We’re discussing the… — @Cloudflare
- We Got Mimo V2.6 Running on 2 x DGX Sparks. Recipes Incoming and it is FAST ! 30 - 90… — @Tech2Wild
- RT @Tech2Wild: We Got Mimo V2.6 Running on 2 x DGX Sparks. Recipes Incoming and it is… — @TechMDAI
- Microsoft Teams is now on Vercel Connect. — @vercel_dev
- a easy trap for infra startups is to try get to the highest ARR the fastest by focusing… — @AniC_dev
- 如果只是想把一个文件发给别人,真的没必要折腾复杂网盘。 — @bkdgiffug
- ☁️🐸 $200 in AWS credits for the 1st 50 teams who go PRO: http://jfrog.com/cloud-pro — @jfrog
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 20.7k posts from 4.9k X accounts over the last 21 days, 2.4k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-22 08:21 UTC. Full method.