PrismML just launched a version of Qwen3.8 27B weights that are compressed down to 5.95…
This is a dev post classified by Jev as AI dev tools (a tool drop), kept by the Dev Radar because it carries real work, not commentary.
Exciting! PrismML just launched a version of Qwen3.8 27B weights that are compressed down to 5.95 GB and that runs on my browser on Macbook Pro with 30 tok/s If it indeed is as lossless as they claim, then this will be a gamechanger for devices with 16 GB of ram! I will run some benchmarks on it and will keep you posted! Go try it out in the Hugging Face space: https://huggingface.co/spaces/webml-community/ternary-bonsai-2-webgpu-kernels
Posted by Onur Solmaz (10k followers) 1 days ago · 161 likes · 29.7k views · view the original post on X. Kept by the Dev Radar as AI dev tools. Tools mentioned: ternary-bonsai-2-webgpu-kernels.
More dev work like this
- I think we’re looking for this — @vaibcode
- When a hard coding decision needs a second opinion, don’t settle for one model. — @DanKornas
- Banger paper from MIT and Sakana AI. — @dair_ai
- Nautilo agents use memory like memory competition champs. They never forget and they… — @Dan_Jeffries1
- AI agents need guardrails before they reach your tools. — @DanKornas
- 20-30 agents are only useful when you stop babysitting every one of them. — @catmanyau
- Okay, everyone wants us to give the unbiased facts. — @Teknium
- A single RTX 3090 hit 381 tok/s on Qwen3.8-27B. — @0x0SojalSec
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 14.4k posts from 4.7k X accounts over the last 21 days, 1.6k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 21:10 UTC. Full method.