New on Dedicated Model Inference: canary rollouts.
This is a dev post classified by Jev as Hosting & infra (a launch), kept by the Dev Radar because it carries real work, not commentary.
New on Dedicated Model Inference: canary rollouts. Upgrade the model behind a live endpoint without downtime. Traffic moves from your current deployment to the new checkpoint in gated steps (default 5% → 25% → 50% → 100%). Health checks run before any traffic shifts. After every step, metric gates compare the new model's p95 latency and error rate against the old one. If a gate trips, the rollout pauses at the canary share and waits for you: resume, promote to 100%, or roll back. Three strategies: canary, blue-green, and rolling. Available now via the tg CLI, REST API, and Python SDK. Lea
Posted by Together AI (63.4k followers) 1 h ago · 5 likes · 1.2k views · view the original post on X. Kept by the Dev Radar as Hosting & infra.
More dev work like this
- how do you upgrade an endpoint to a new model while it is serving millions of prod… — @zainhas
- speedmasters, daytonas, tanks — @AniC_dev
- As AI and agentic workloads reshape production infrastructure, how is your team running… — @CloudNativeFdn
- Deploying DiffusionGemma-Jev (djev) just got a lot easier. You can now spin up a Jev… — @googlegemma
- Getting these 8x RTX PRO 6000 bad boys ready for MiMo-V2.6-Pro-RL — @TheAhmadOsman
- Join @AWScloud, @NVIDIA, and LangChain for an in-person session focused on how teams can… — @LangChain
- Arc Mainnet launched with more than infrastructure. — @arc
- Exciting day for Databricks customers! We’re launching three new frontier models on… — @databricks
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 21.7k posts from 5k X accounts over the last 21 days, 2.5k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-23 01:30 UTC. Full method.