Dev Radar
Support
LiveUpdated 2026-09-19 18:39 UTC

MLPerf Inference v6.1 is out, and we have two firsts: the first agentic inference…

MLPerf Inference v6.1 is out, and we have two firsts: the first agentic inference workload on datacenter hardware, and…

This is a dev post classified by Jev as Hosting & infra (a launch), kept by the Dev Radar because it carries real work, not commentary.

MLPerf Inference v6.1 is out, and we have two firsts: the first agentic inference workload on datacenter hardware, and the first MLPerf deployment of a model over a trillion parameters. On 4x NVIDIA Blackwell Ultra GPUs, we posted the leading Offline throughput on GPT-OSS 120B among all Blackwell Ultra GPU submissions, plus an 8.85% throughput gain over v6.0 on identical hardware. That's six months of pure software optimization. On NVIDIA HGX B200, we swapped in Kimi K2.6 for the open-division agentic benchmark, running a 1T+ parameter model where the reference workload expects 27B. Same har

Posted by Lambda (20.9k followers) 3 days ago · 18 likes · 1.2k views · view the original post on X. Kept by the Dev Radar as Hosting & infra.

More dev work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 12.2k posts from 4.7k X accounts over the last 21 days, 1.4k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:39 UTC. Full method.