Dev Radar
Support
LiveUpdated 2026-09-19 22:52 UTC

Strix Halo and DGX Spark recently got a boost from llama.cpp.

Strix Halo and DGX Spark recently got a boost from llama.cpp. A new llama.cpp PR adds Vulkan support for…

This is a dev post classified by Jev as AI dev tools (a tool drop), kept by the Dev Radar because it carries real work, not commentary.

Strix Halo and DGX Spark recently got a boost from llama.cpp. A new llama.cpp PR adds Vulkan support for Qwen3.8-Flash-Next’s new Hyper-Connection ops — and the same hardware immediately gets more performance. On Ryzen AI Max+ 395 / Radeon 8060S 👇 📥 Prefill 330.9 → 337.0 tok/s (+1.9%) 🚀 Decode 25.55 → 26.17 tok/s (+2.4%) But THIS caught my eye: 🔴 ROCm reference: 21.34 tok/s 🟣 Vulkan after PR: 26.17 tok/s That’s 22.6% faster decode than the ROCm reference in this test. 👀 And DGX Spark gets a bump too: ⚡ Prefill 746.6 → 803.7 tok/s (+7.6%) 🚀 Decode 30.55 → 31.50 tok/s (+3.1%) And Apple

Posted by David Hendrickson (11.1k followers) 2 days ago · 60 likes · 4.4k views · view the original post on X. Kept by the Dev Radar as AI dev tools.

More dev work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 15.1k posts from 4.7k X accounts over the last 21 days, 1.7k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 22:52 UTC. Full method.