llama.cpp
llama.cpp is Summary Integration checkpoint: Bonsai 2 27B served at its full 262,144-token window with the MTP draft head, on a 12 GB RTX 4070, at 103 tok/s greedy (three-prompt mean, up from 60 without the dra.... It is ranked #1637 on the Dev Radar, in Other, first seen 19 h ago and shared in 1 post (33 views).
cuda: Bonsai 2 27B at the full 262k window with the MTP head on 12 GB Ada: #218 + #215 hybrid PTQ1_0 dispatch, in-place q4_0/q8_0 K/V flash attention (integration checkpoint) by professorpalmer · Pull Request #221 · PrismML-Eng/llama.cpp · GitHub Skip to content Navigation Menu Sign in Appearance settings Search / Sign in Sign up Appearance settings You signed in with another tab or window. Reload to refresh your session. You signed out in another tab or window. Reload to refresh your session. You switched accounts on another tab or window. Reload to refresh your session. Dismiss alert {{ message }} PrismML-Eng / llama.cpp Public forked…
What people said about llama.cpp on X
A 7.8 point jump off a patch is exactly the kind of number my own box can fake with zero code changes. Same 200 token prompt, same flags, warm, four repeats: 18.29 / 19.49 / 17.95 / 18.04 tok/s. Ninety minutes earlier the same prompt measured 55.4. Until a run_order column says what state the engine was in, patch…
— @cozybearlog, 19 h ago · 0 likes · see the post
Alternatives to llama.cpp
- Spark-X2.5 — Model Scope
- mimo-v2.6 RL — Streaming the run
- qwen-image-2.1 — Qwen's most powerful open-source image generation model - QwenLM/Qwen-Image-2.1
- Ling-3.0-flash-Fin — Qwen-Drive-1.0-4B is a vision-language foundation model that handles 3D perception, driving VQA, and motion planning…
- mcdma — Mac Studio A — Gateway. The only Studio with a Mellanox ConnectX‑4 on the MikroTik fabric (via a Thunderbolt 5…
- Cactus Compute — It runs on mobiles, wearables, smart home devices, small robots and microcontrollers, with prebuilt engines for macOS,…
llama.cpp in numbers
- Rank on the Dev Radar: #1637 of 1886
- Shared in 1 post by 1 account: @cozybearlog
- 33 views on those posts
- First seen 19 h ago, last shared 19 h ago
- Pricing seen by Jev: open source
- Market: Other
FAQ
What is llama.cpp?
Summary Integration checkpoint: Bonsai 2 27B served at its full 262,144-token window with the MTP draft head, on a 12 GB RTX 4070, at 103 tok/s greedy (three-prompt mean, up from 60 without the dra... It was first shared on X 19 h ago and is ranked #1637 on the Dev Radar.
Is llama.cpp free?
It is open source.
Who shared llama.cpp?
1 account on X, including @cozybearlog, in 1 post totalling 33 views.
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 17k posts from 4.9k X accounts over the last 21 days, 1.9k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-20 22:16 UTC. Full method.