Dev Radar
Support
LiveUpdated 2026-09-19 19:24 UTC

vllm-exl3

vllm-exl3 — Serve EXL3 (ExLlamaV3 trellis) quantized models on vLLM fork runtimes — any…

vllm-exl3 is Serve EXL3 (ExLlamaV3 trellis) quantized models on vLLM fork runtimes — any architecture, mixed per-layer bitrates, composable with source-format non-routed weights - vcruz305/vllm-exl3. It is ranked #551 on the Dev Radar, in AI dev tools, first seen 16 days ago and shared in 8 posts (70.2k views).

Visit github.com

GitHub - vcruz305/vllm-exl3: Serve EXL3 (ExLlamaV3 trellis) quantized models on vLLM fork runtimes — any architecture, mixed per-layer bitrates, composable with source-format non-routed weights · GitHub Skip to content Navigation Menu Sign in Appearance settings Search / Sign in Sign up Appearance settings You signed in with another tab or window. Reload to refresh your session. You signed out in another tab or window. Reload to refresh your session. You switched accounts on another tab or window. Reload to refresh your session. Dismiss alert {{ message }} vcruz305 / vllm-exl3 Public Notifications You must be signed in to change…

What people said about vllm-exl3 on X

vllm-exl3 v0.3.0 is LIVE with custom native CUDA kernels for 2-bit EXL3 on NVIDIA DGX Spark GB10. GLM-5.3-Flash-EXL3-K2 jumped from 16.9 → 24.6 tok/s average single-stream decode, a +45.6% gain. Coding hit 27.6 tok/s, +85.6%. 🚀 The previous ExLlamaV3-backed path inside vLLM was leaving a lot of GB10 bandwidth on the…

@ViC305, 16 days ago · 128 likes · see the post

Victor’s quietly becoming one of the people to watch for real GB10/Blackwell kernel work not benchmarks, actual in-register Trellis dequant + fused MoE decode shipped as native CUDA. These aren’t just “Quick ships” these are well built and pushing the limits especially with EXL3 🔥 +45.6% decode, 13x prefill, 2.7x…

@Blackwellboy, 16 days ago · 84 likes · see the post

Qwen3.8-Flash-Next now reaches ~43 tok/s after a 122,902-token prompt on ONE DGX Spark. ⚡🚀 MTP k=2 won my draft-depth sweep, with +42.5% mean decode over no draft. The PLE table stays fully on-device. I promised the deeper MTP tests. Here are the results, and now you can explore them in an interactive benchmark page…

@ViC305, 11 days ago · 58 likes · see the post

Alternatives to vllm-exl3

vllm-exl3 in numbers

FAQ

What is vllm-exl3?

Serve EXL3 (ExLlamaV3 trellis) quantized models on vLLM fork runtimes — any architecture, mixed per-layer bitrates, composable with source-format non-routed weights - vcruz305/vllm-exl3 It was first shared on X 16 days ago and is ranked #551 on the Dev Radar.

Is vllm-exl3 free?

It is open source.

Who shared vllm-exl3?

2 accounts on X, including @ViC305, @Blackwellboy, in 8 posts totalling 70.2k views.

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 13.5k posts from 4.7k X accounts over the last 21 days, 1.5k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 19:24 UTC. Full method.