Dev Radar
Support
LiveUpdated 2026-09-19 22:52 UTC

Insanity? Who has 8x5090s? 🤯 ... this is one for the record books for Local AI!

Insanity? Who has 8x5090s? 🤯 ... this is one for the record books for Local AI! A new optimization project has…

This is a dev post classified by Jev as Hosting & infra (a tool drop), kept by the Dev Radar because it carries real work, not commentary.

Insanity? Who has 8x5090s? 🤯 ... this is one for the record books for Local AI! A new optimization project has DeepSeek-V4.1-Flash running across 👇 🎮 8× RTX 5090 32GB 👈 👀 🧠 256GB total VRAM 💾 503GiB system RAM 🔗 NO NVLink And the latest results ... ⚡ 74.75 tok/s — single request 🚀 168.49 tok/s — concurrency 8 🔥 217.55 aggregate tok/s — concurrency 64 📚 4,078 tok/s — 8K prefill They didn't shrink the model into some tiny fine-tune. The gains come from the inference stack: ✅ vLLM fixes ✅ CUDA graphs ✅ Expert parallelism ✅ Selective CPU offload ✅ Better movement of MoE experts between RAM + G

Posted by David Hendrickson (11.1k followers) 3 days ago · 29 likes · 5.8k views · view the original post on X. Kept by the Dev Radar as Hosting & infra.

More dev work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 15.1k posts from 4.7k X accounts over the last 21 days, 1.7k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 22:52 UTC. Full method.