Great speed-up: Full Kimi K3 — the 2.8T-parameter model running on 16× NVIDIA GB10…
This is a dev post classified by Jev as Hosting & infra (a launch), kept by the Dev Radar because it carries real work, not commentary.
Great speed-up: Full Kimi K3 — the 2.8T-parameter model running on 16× NVIDIA GB10 (Spark) - V5 coming by tomorrow, cca 20% faster and much improved concurrency: (~30t/s C1, 87t/s C8 ) 136t/s peak speed at C8: Coding: game-bench 1 2 4 8 (3000-token runs) C1: 29.81 tok/s C2: 42.00 agg / 21.00 mean stream C4: 58.00 agg / 14.50 mean stream C8: 87.12 agg / 10.89 mean stream Prose: llama-bench coherent corpus modeltestt/speak t/s KIMI-K3pp2048 @ d4000880.35 ± 0.00 KIMI-K3tg2048 @ d400023.59 ± 0.0046.00 ± 0.00 KIMI-K3pp2048 @ d100000812.56 ± 0.00 KIMI-K3tg2048 @ d10000021.13 ± 0.0033.00 ± 0.00 KIMI
Posted by llmwish.open (1.6k followers) 1 days ago · 242 likes · 38.6k views · view the original post on X. Kept by the Dev Radar as Hosting & infra.
More dev work like this
- We’re down to the final day. — @alibaba_cloud
- @maria_rcks — @maria_rcks
- Last night, in a basement in Florida, one conversation ran across two kinds of silicon… — @volatilemarkts
- I have to retract what I posted as the fastest speed of Qwen 3.8 Flash on a single DGX… — @yume_arasaki
- How does a global bank navigate regulations and rapid AI cycles without vendor lock-in?… — @RedHat
- Level up your skills while shaping open source priorities! — @CloudNativeFdn
- 🤙 DeepSeek V4.1 Flash. > 100 t/s reached. max 178 t/s stream decode at b1. — @qubitium
- Azure AI Engineer Roadmap Explained! — @AiswaryaVenkit1
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 17.6k posts from 4.9k X accounts over the last 21 days, 2k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-21 02:20 UTC. Full method.