Dev Radar
Support
LiveUpdated 2026-09-19 18:39 UTC

vllm

vllm — A high-throughput and memory-efficient inference and serving engine for LLMs -…

vllm is A high-throughput and memory-efficient inference and serving engine for LLMs - mmastrac/vllm. It is ranked #323 on the Dev Radar, in AI dev tools, first seen 2 days ago and shared in 1 post (3k views).

Visit github.com

What people said about vllm on X

Now you can run a Jev-like on your own hardware using DiffusionGemma from @GoogleDeepMind. Untuned, on a DGX Spark, 0.20s median structured results! Branch is in progress, but tested and working on real hardware. https://github.com/mmastrac/vllm/tree/structured-reads

@mmastrac, 2 days ago · 50 likes · see the post

Alternatives to vllm

vllm in numbers

FAQ

What is vllm?

A high-throughput and memory-efficient inference and serving engine for LLMs - mmastrac/vllm It was first shared on X 2 days ago and is ranked #323 on the Dev Radar.

Is vllm free?

It is open source.

Who shared vllm?

1 account on X, including @mmastrac, in 1 post totalling 3k views.

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 12.2k posts from 4.7k X accounts over the last 21 days, 1.4k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:39 UTC. Full method.