🤙 DeepSeek V4.1 Flash. > 100 t/s reached. max 178 t/s stream decode at b1.
This is a dev post classified by Jev as Hosting & infra (a tool drop), kept by the Dev Radar because it carries real work, not commentary.
🤙 DeepSeek V4.1 Flash. > 100 t/s reached. max 178 t/s stream decode at b1. Full native, no additional quantization. ~$5.50/hr compute cost. 256k ctx, Max batch 12. No caching. Cold prefill/decode. Nvidia platform, b1/b2 has mtp, mtp has ~45% uplift in decode.
Posted by Qubitium (1.6k followers) 1 h ago · 3 likes · 227 views · view the original post on X. Kept by the Dev Radar as Hosting & infra.
More dev work like this
- Azure AI Engineer Roadmap Explained! — @AiswaryaVenkit1
- 文科生有救了,妈妈再也不怕我的localhost:8888别人访问不了了! — @mylifcc
- Kubernetes Node Failure & Recovery in action 👇 — @twtayaan
- 一个 3.3 万 Star 的 Rust 项目,刚刚发布 1.0。 — @mylifcc
- boat — @AniC_dev
- The next chapter of AI is agentic. — @alibaba_cloud
- 📣 New Inferenceing Engine Alert! — @TeksEdge
- Happy weekend! Let's forget all the troubles in daily life; just go outside to touch… — @alibaba_cloud
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 15.3k posts from 4.8k X accounts over the last 21 days, 1.7k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-20 11:54 UTC. Full method.