Dev Radar
Support
LiveUpdated 2026-09-22 08:21 UTC

Your cache is everything when it comes to inference, but how do you make sure you're…

Your cache is everything when it comes to inference, but how do you make sure you're keeping it around with local…

This is a dev post classified by Jev as Hosting & infra (a tutorial), kept by the Dev Radar because it carries real work, not commentary.

Your cache is everything when it comes to inference, but how do you make sure you're keeping it around with local models? Typically when you start a server, it's a fresh slate. Then the KV cache grows, evolves, and your cache hits keep growing. But then you shut down the server, swap to a new fancy toy (model), and your cache is destroyed. Nothing we can do there, since caches are model-specific. But let's say then you want to go *back* to deploying the original model. Now *it's* KV cache was also destroyed! @sgl_project supports directly setting up hicache through some configuration a

Posted by Zach Mueller (17.9k followers) 16 h ago · 26 likes · 1.7k views · view the original post on X. Kept by the Dev Radar as Hosting & infra.

More dev work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 20.7k posts from 4.9k X accounts over the last 21 days, 2.4k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-22 08:21 UTC. Full method.