Qwen3.8-27B at up to 144 tok/s on an M5 Max MacBook Pro?! You have to try this Splash…
This is a dev post classified by Jev as AI dev tools (a tool drop), kept by the Dev Radar because it carries real work, not commentary.
Qwen3.8-27B at up to 144 tok/s on an M5 Max MacBook Pro?! You have to try this Splash Engine! It's a new open-source inference engine called Splash. Instead of being a universal runtime like llama.cpp or Ollama, Splash optimizes the whole stack around the exact model 👇 ⚙️ model-specific kernels 🧠 hardware-aware memory planning 🚀 DFlash2 speculative decoding 💾 prompt-cache reuse 👥 continuous batching On the same 48GB M5 Pro running Qwen3.8-27B: 🚀 Splash: 74 tok/s ⚡ oMLX: 38 tok/s 🐌 Ollama: 24 tok/s At 32K context: 🚀 Splash: 54 tok/s And with 4 concurrent requests: 🔥 170 aggregate tok/s vs
Posted by David Hendrickson (11.1k followers) 4 h ago · 17 likes · 1.6k views · view the original post on X. Kept by the Dev Radar as AI dev tools. Tools mentioned: LM Studio Blog.
More dev work like this
- Did you already know, that you can interact via CLI with the apple foundation models? — @haukejung
- A new benchmark called JevBench just dropped. — @rohanpaul_ai
- Stop burning turns rewriting vague AI prompts — @DanKornas
- I WOKE UP TO MONEY — @michael_chomsky
- we launched the most comprehensive ai performance engineering repo in the world — @wafer_ai
- Structures a Claude Code session into a game studio with 49 specialized AI agents and 73… — @tom_doerr
- Jev is cool. So is it's OSS companion, Laya. — @BenjDicken
- I think we’re looking for this — @vaibcode
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 15.1k posts from 4.7k X accounts over the last 21 days, 1.7k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 22:52 UTC. Full method.