Dev Radar
Support
LiveUpdated 2026-09-19 17:33 UTC

Agent regressions are hard to catch when traces are read-only.

Agent regressions are hard to catch when traces are read-only. Kitaru is a replay-based evaluation tool for AI agents…

This is a dev post classified by Jev as Testing & observability (a tool drop), kept by the Dev Radar because it carries real work, not commentary.

Agent regressions are hard to catch when traces are read-only. Kitaru is a replay-based evaluation tool for AI agents that turns recorded or imported production runs into sessions you can test. It helps you see what changed before shipping by replaying those sessions against a new model, prompt, or code change. Key features: • Replay-based evals – re-executes your agent code while answering tool calls from the recording • Trace imports – brings in runs from Langfuse, LangSmith, Braintrust, Logfire, or Arize Phoenix • Baseline and forked replays – compare an unchanged replay with the effect

Posted by Dan Kornas (99.1k followers) 1 days ago · 14 likes · 1.6k views · view the original post on X. Kept by the Dev Radar as Testing & observability.

More dev work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 12.2k posts from 4.7k X accounts over the last 21 days, 1.3k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 17:33 UTC. Full method.