Transformers has supported loading GGUF files for a few years now, by unquantizing them.
This is a dev post classified by Jev as AI dev tools (a launch), kept by the Dev Radar because it carries real work, not commentary.
Transformers has supported loading GGUF files for a few years now, by unquantizing them. Thanks to @_marcsun, we're now using GGML kernels through the `kernels` library to run at the same performance as llama.cpp Huge kudos to the entire @ggml_org for making these kernels!
Posted by Lysandre (12.5k followers) 1 h ago · 11 likes · 472 views · view the original post on X. Kept by the Dev Radar as AI dev tools.
More dev work like this
- What happens when AI makes developers faster, but the rest of the software organisation… — @johncrickett
- Champion! HKUST Students Forged Dual Workbenches with WorkBuddy — @WorkBuddy_AI
- Similar to MiMo if you wanna generate synthetic coding RL environments from GitHub at… — @adithya_s_k
- Jev has been exploding in popularity recently. — @0x_rody
- Cool~😎 — @YouWareAI
- Great paper from Google and colleagues. — @omarsar0
- Your research agent is only as good as the fetch layer behind it! — @Sumanth_077
- Today's video is a real demo of a multi-agent workflow in Orca! — @tonbistudio
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 21k posts from 5k X accounts over the last 21 days, 2.4k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-22 14:50 UTC. Full method.