mlx-serve
mlx-serve is Extend Qwen3.8-Flash-Next (qwen4_exp, head_dim 256, partial_rotary_factor 0.25, rotary_dim 64) from 262144 to 1048576 via HF/vLLM-compatible YaRN. Matches transformers modeling_rope_utils._compute_.... It is ranked #749 on the Dev Radar, in AI dev tools, first seen 18 days ago and shared in 4 posts (54k views).
feat(qwen4_exp): YaRN rope scaling to 1M context by beamivalice · Pull Request #323 · ddalcu/mlx-serve · GitHub Skip to content Navigation Menu Sign in Appearance settings Search / Sign in Sign up Appearance settings You signed in with another tab or window. Reload to refresh your session. You signed out in another tab or window. Reload to refresh your session. You switched accounts on another tab or window. Reload to refresh your session. Dismiss alert {{ message }} ddalcu / mlx-serve Public Uh oh! There was an error while loading. Please reload this page . Notifications You must be signed in to change notification settings Fork 131 Star…
What people said about mlx-serve on X
MLX-Serve 26.9.3 is out. Fastest engine to run local models on a Mac: LLM's, Video, Image, Voice Clone, Music Gen. https://github.com/ddalcu/mlx-serve/releases/tag/v26.9.3 Qwen Flash Next goes over 100 tok/s on M4/M5Max and insane prefill speeds. This release is huge, consider it "beta". Next release will be focused…
— @ddalcu, 4 days ago · 478 likes · see the post
MLX-Serve 26.9.2 is out. Fixes, Polish, Speed. If you have 96-128GB Ram, 👑 Qwen Flash Next is the king of local models for now. Without KV8 Cache + M5Max, you might see it go above 150 Tok/s. https://github.com/ddalcu/mlx-serve/releases/tag/v26.9.2
— @ddalcu, 11 days ago · 137 likes · see the post
MLXServe Bonsai 2 edition pre-release: https://github.com/ddalcu/mlx-serve/releases/tag/v26.9.5-pre-release.1 A lot of people are asking me for this. Still working on performance and correctness, Final release once testing is done, but use this for now, it’s still way faster than anything else, at about 50-70 tok/s…
— @ddalcu, 2 days ago · 110 likes · see the post
Alternatives to mlx-serve
- muse.ai — New connectors are live today. Come build with us.
- classifier.dev — now outperforms jev and is free
- academy.dair.ai — Chat with Paper
- StepFun Open Platform — Step API · Stable · High-Performance · Easy Integration. Leading models and tools to accelerate the deployment of your…
- qwen-image-2.1 — Hugging Face
- Union Alpha — Union Alpha is a multimodal model built for research, coding, and agentic workflows, while delivering frontier-level…
mlx-serve in numbers
- Rank on the Dev Radar: #749 of 1818
- Shared in 4 posts by 1 account: @ddalcu
- 54k views on those posts
- First seen 18 days ago, last shared 2 days ago
- Pricing seen by Jev: open source
- Market: AI dev tools
FAQ
What is mlx-serve?
Extend Qwen3.8-Flash-Next (qwen4_exp, head_dim 256, partial_rotary_factor 0.25, rotary_dim 64) from 262144 to 1048576 via HF/vLLM-compatible YaRN. Matches transformers modeling_rope_utils._compute_... It was first shared on X 18 days ago and is ranked #749 on the Dev Radar.
Is mlx-serve free?
It is open source.
Who shared mlx-serve?
1 account on X, including @ddalcu, in 4 posts totalling 54k views.
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 16.1k posts from 4.9k X accounts over the last 21 days, 1.8k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-20 19:24 UTC. Full method.