vllm-exl3
vllm-exl3 is Serve EXL3 (ExLlamaV3 trellis) quantized models on vLLM fork runtimes — any architecture, mixed per-layer bitrates, composable with source-format non-routed weights - vcruz305/vllm-exl3. It is ranked #551 on the Dev Radar, in AI dev tools, first seen 16 days ago and shared in 8 posts (70.2k views).
GitHub - vcruz305/vllm-exl3: Serve EXL3 (ExLlamaV3 trellis) quantized models on vLLM fork runtimes — any architecture, mixed per-layer bitrates, composable with source-format non-routed weights · GitHub Skip to content Navigation Menu Sign in Appearance settings Search / Sign in Sign up Appearance settings You signed in with another tab or window. Reload to refresh your session. You signed out in another tab or window. Reload to refresh your session. You switched accounts on another tab or window. Reload to refresh your session. Dismiss alert {{ message }} vcruz305 / vllm-exl3 Public Notifications You must be signed in to change…
What people said about vllm-exl3 on X
vllm-exl3 v0.3.0 is LIVE with custom native CUDA kernels for 2-bit EXL3 on NVIDIA DGX Spark GB10. GLM-5.3-Flash-EXL3-K2 jumped from 16.9 → 24.6 tok/s average single-stream decode, a +45.6% gain. Coding hit 27.6 tok/s, +85.6%. 🚀 The previous ExLlamaV3-backed path inside vLLM was leaving a lot of GB10 bandwidth on the…
— @ViC305, 16 days ago · 128 likes · see the post
Victor’s quietly becoming one of the people to watch for real GB10/Blackwell kernel work not benchmarks, actual in-register Trellis dequant + fused MoE decode shipped as native CUDA. These aren’t just “Quick ships” these are well built and pushing the limits especially with EXL3 🔥 +45.6% decode, 13x prefill, 2.7x…
— @Blackwellboy, 16 days ago · 84 likes · see the post
Qwen3.8-Flash-Next now reaches ~43 tok/s after a 122,902-token prompt on ONE DGX Spark. ⚡🚀 MTP k=2 won my draft-depth sweep, with +42.5% mean decode over no draft. The PLE table stays fully on-device. I promised the deeper MTP tests. Here are the results, and now you can explore them in an interactive benchmark page…
— @ViC305, 11 days ago · 58 likes · see the post
Alternatives to vllm-exl3
- muse.ai — New connectors are live today. Come build with us.
- classifier.dev — now outperforms jev and is free
- mimo-v2.6 RL — Streaming the run
- academy.dair.ai — Chat with Paper
- Union Alpha — Union Alpha is a multimodal model built for research, coding, and agentic workflows, while delivering frontier-level…
- cua — Draft #3943
vllm-exl3 in numbers
- Rank on the Dev Radar: #551 of 1500
- Shared in 8 posts by 2 accounts: @ViC305, @Blackwellboy
- 70.2k views on those posts
- First seen 16 days ago, last shared 10 days ago
- Pricing seen by Jev: open source
- Market: AI dev tools
FAQ
What is vllm-exl3?
Serve EXL3 (ExLlamaV3 trellis) quantized models on vLLM fork runtimes — any architecture, mixed per-layer bitrates, composable with source-format non-routed weights - vcruz305/vllm-exl3 It was first shared on X 16 days ago and is ranked #551 on the Dev Radar.
Is vllm-exl3 free?
It is open source.
Who shared vllm-exl3?
2 accounts on X, including @ViC305, @Blackwellboy, in 8 posts totalling 70.2k views.
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 13.5k posts from 4.7k X accounts over the last 21 days, 1.5k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 19:24 UTC. Full method.