Dev Radar
Support
LiveUpdated 2026-09-19 22:52 UTC

Qwen3.8-27B at up to 144 tok/s on an M5 Max MacBook Pro?! You have to try this Splash…

Qwen3.8-27B at up to 144 tok/s on an M5 Max MacBook Pro?! You have to try this Splash Engine! It's a new open-source…

This is a dev post classified by Jev as AI dev tools (a tool drop), kept by the Dev Radar because it carries real work, not commentary.

Qwen3.8-27B at up to 144 tok/s on an M5 Max MacBook Pro?! You have to try this Splash Engine! It's a new open-source inference engine called Splash. Instead of being a universal runtime like llama.cpp or Ollama, Splash optimizes the whole stack around the exact model 👇 ⚙️ model-specific kernels 🧠 hardware-aware memory planning 🚀 DFlash2 speculative decoding 💾 prompt-cache reuse 👥 continuous batching On the same 48GB M5 Pro running Qwen3.8-27B: 🚀 Splash: 74 tok/s ⚡ oMLX: 38 tok/s 🐌 Ollama: 24 tok/s At 32K context: 🚀 Splash: 54 tok/s And with 4 concurrent requests: 🔥 170 aggregate tok/s vs

Posted by David Hendrickson (11.1k followers) 4 h ago · 17 likes · 1.6k views · view the original post on X. Kept by the Dev Radar as AI dev tools. Tools mentioned: LM Studio Blog.

More dev work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 15.1k posts from 4.7k X accounts over the last 21 days, 1.7k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 22:52 UTC. Full method.