A 744B-parameter model is mostly sitting idle at any given token. Colibri takes that…
This is a dev post classified by Jev as Other (a tool drop), kept by the Dev Radar because it carries real work, not commentary.
A 744B-parameter model is mostly sitting idle at any given token. Colibri takes that literally. Instead of forcing every MoE expert into RAM or VRAM, it keeps dense weights resident and streams routed experts through VRAM, RAM and SSD as one memory hierarchy. The engine is pure C, and the hardware mainly changes how fast those experts arrive. It is a useful reminder that sparse inference is partly a data-movement problem. The repo documents the 744B GLM-5.2 path, pure-C engine and storage/RAM/VRAM hierarchy. GitHub Repo: https://github.com/JustVugg/colibri
Posted by Sandhya (1.3k followers) 2 days ago · 2 likes · 76 views · view the original post on X. Kept by the Dev Radar as Other. Tools mentioned: colibri.
More dev work like this
- dear rtx 3060 owners, and every 12gb card behind it. bonsai 2 27b dense went from 26 to… — @sudoingX
- Nice: banked codex reset for all of us. — @kimmonismus
- The #Mathematical #Programming - #FPGA. #BigData #Analytics #DataScience #AI… — @gp_pulipaka
- A Constraint Based Approach, ML. #BigData #Analytics #DataScience #AI #MachineLearning… — @gp_pulipaka
- Meet AliceAI-Foundation-80B-A3B-Base: a massive 80B parameter Mixture-of-Experts model… — @HuggingModels
- I am working on a new skill for Antigravity to prepare for Gemini 4, and I decided to… — @IamEmily2050
- There’s no way — @luckeyfaraday
- Cooked up a GLM-5.3-Flash-EXL3–2.32bpw for the 128GB crowd, perfect for your @NVIDIAAI… — @mr_r0b0t
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 12.2k posts from 4.7k X accounts over the last 21 days, 1.4k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:39 UTC. Full method.