Your cache is everything when it comes to inference, but how do you make sure you're…
This is a dev post classified by Jev as Hosting & infra (a tutorial), kept by the Dev Radar because it carries real work, not commentary.
Your cache is everything when it comes to inference, but how do you make sure you're keeping it around with local models? Typically when you start a server, it's a fresh slate. Then the KV cache grows, evolves, and your cache hits keep growing. But then you shut down the server, swap to a new fancy toy (model), and your cache is destroyed. Nothing we can do there, since caches are model-specific. But let's say then you want to go *back* to deploying the original model. Now *it's* KV cache was also destroyed! @sgl_project supports directly setting up hicache through some configuration a
Posted by Zach Mueller (17.9k followers) 16 h ago · 26 likes · 1.7k views · view the original post on X. Kept by the Dev Radar as Hosting & infra.
More dev work like this
- At HUAWEI CONNECT, Huawei Deputy Chairman of the Board and Rotating Chairman David Wang… — @Huawei
- Join us on October 7th for an exclusive webinar on Cloudflare OS. We’re discussing the… — @Cloudflare
- We Got Mimo V2.6 Running on 2 x DGX Sparks. Recipes Incoming and it is FAST ! 30 - 90… — @Tech2Wild
- RT @Tech2Wild: We Got Mimo V2.6 Running on 2 x DGX Sparks. Recipes Incoming and it is… — @TechMDAI
- Microsoft Teams is now on Vercel Connect. — @vercel_dev
- a easy trap for infra startups is to try get to the highest ARR the fastest by focusing… — @AniC_dev
- 如果只是想把一个文件发给别人,真的没必要折腾复杂网盘。 — @bkdgiffug
- ☁️🐸 $200 in AWS credits for the 1st 50 teams who go PRO: http://jfrog.com/cloud-pro — @jfrog
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 20.7k posts from 4.9k X accounts over the last 21 days, 2.4k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-22 08:21 UTC. Full method.