Dev Radar
Support
LiveUpdated 2026-09-20 19:24 UTC

mlx-serve

mlx-serve — Extend Qwen3.8-Flash-Next (qwen4_exp, head_dim 256, partial_rotary_factor 0.25,…

mlx-serve is Extend Qwen3.8-Flash-Next (qwen4_exp, head_dim 256, partial_rotary_factor 0.25, rotary_dim 64) from 262144 to 1048576 via HF/vLLM-compatible YaRN. Matches transformers modeling_rope_utils._compute_.... It is ranked #749 on the Dev Radar, in AI dev tools, first seen 18 days ago and shared in 4 posts (54k views).

Visit github.com

feat(qwen4_exp): YaRN rope scaling to 1M context by beamivalice · Pull Request #323 · ddalcu/mlx-serve · GitHub Skip to content Navigation Menu Sign in Appearance settings Search / Sign in Sign up Appearance settings You signed in with another tab or window. Reload to refresh your session. You signed out in another tab or window. Reload to refresh your session. You switched accounts on another tab or window. Reload to refresh your session. Dismiss alert {{ message }} ddalcu / mlx-serve Public Uh oh! There was an error while loading. Please reload this page . Notifications You must be signed in to change notification settings Fork 131 Star…

What people said about mlx-serve on X

MLX-Serve 26.9.3 is out. Fastest engine to run local models on a Mac: LLM's, Video, Image, Voice Clone, Music Gen. https://github.com/ddalcu/mlx-serve/releases/tag/v26.9.3 Qwen Flash Next goes over 100 tok/s on M4/M5Max and insane prefill speeds. This release is huge, consider it "beta". Next release will be focused…

@ddalcu, 4 days ago · 478 likes · see the post

MLX-Serve 26.9.2 is out. Fixes, Polish, Speed. If you have 96-128GB Ram, 👑 Qwen Flash Next is the king of local models for now. Without KV8 Cache + M5Max, you might see it go above 150 Tok/s. https://github.com/ddalcu/mlx-serve/releases/tag/v26.9.2

@ddalcu, 11 days ago · 137 likes · see the post

MLXServe Bonsai 2 edition pre-release: https://github.com/ddalcu/mlx-serve/releases/tag/v26.9.5-pre-release.1 A lot of people are asking me for this. Still working on performance and correctness, Final release once testing is done, but use this for now, it’s still way faster than anything else, at about 50-70 tok/s…

@ddalcu, 2 days ago · 110 likes · see the post

Alternatives to mlx-serve

mlx-serve in numbers

FAQ

What is mlx-serve?

Extend Qwen3.8-Flash-Next (qwen4_exp, head_dim 256, partial_rotary_factor 0.25, rotary_dim 64) from 262144 to 1048576 via HF/vLLM-compatible YaRN. Matches transformers modeling_rope_utils._compute_... It was first shared on X 18 days ago and is ranked #749 on the Dev Radar.

Is mlx-serve free?

It is open source.

Who shared mlx-serve?

1 account on X, including @ddalcu, in 4 posts totalling 54k views.

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 16.1k posts from 4.9k X accounts over the last 21 days, 1.8k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-20 19:24 UTC. Full method.