📣 New Inferenceing Engine Alert!
This is a dev post classified by Jev as Hosting & infra (a tool drop), kept by the Dev Radar because it carries real work, not commentary.
📣 New Inferenceing Engine Alert! One DGX Spark is now running DeepSeek V4 Flash + Qwen3.8 Flash-Next at 262K context with ~1,000 tps PP. A new engine called Athena recently dropped for Nvidia GB10 systems. On a single 128GB DGX Spark you can get ... DeepSeek V4 Flash ⚡ 8K prefill: 1,126 tps 📚 256K: 948 tps 🚀 Decode @256K: 19.4 tps Qwen3.8 Flash-Next ⚡ 8K prefill: 1,071 tps 📚 256K: 961 tps 🚀 Decode @256K: 32.1 tps Going from 8K → 256K barely rocks on prefill performance. And Athena caches long conversations to disk. 👈👀 A 141,519-token conversation reportedly restores in: ⚡ 2.1 seconds
Posted by David Hendrickson (11.1k followers) 1 h ago · 4 likes · 628 views · view the original post on X. Kept by the Dev Radar as Hosting & infra.
More dev work like this
- The next chapter of AI is agentic. — @alibaba_cloud
- Happy weekend! Let's forget all the troubles in daily life; just go outside to touch… — @alibaba_cloud
- Managing a MikroTik router shouldn’t mean hunting through CLI commands. — @DanKornas
- To handle bursty, mission-critical workloads without overspending on idle cloud… — @OracleDevs
- Live now: our Inference Engineering Track from AI Engineer World's Fair 2026. — @aiDotEngineer
- gm SF — @AniC_dev
- Last chance: 4 days left to get your ticket to @WeAreDevs! — @Docker
- Accelerating LLM on #AWS Platform! #BigData #Analytics #DataScience #AI #MachineLearning… — @gp_pulipaka
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 15.2k posts from 4.8k X accounts over the last 21 days, 1.7k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-20 03:24 UTC. Full method.