Agent regressions are hard to catch when traces are read-only.
This is a dev post classified by Jev as Testing & observability (a tool drop), kept by the Dev Radar because it carries real work, not commentary.
Agent regressions are hard to catch when traces are read-only. Kitaru is a replay-based evaluation tool for AI agents that turns recorded or imported production runs into sessions you can test. It helps you see what changed before shipping by replaying those sessions against a new model, prompt, or code change. Key features: • Replay-based evals – re-executes your agent code while answering tool calls from the recording • Trace imports – brings in runs from Langfuse, LangSmith, Braintrust, Logfire, or Arize Phoenix • Baseline and forked replays – compare an unchanged replay with the effect
Posted by Dan Kornas (99.1k followers) 1 days ago · 14 likes · 1.6k views · view the original post on X. Kept by the Dev Radar as Testing & observability.
More dev work like this
- What does observability look like when a judgment costs nothing? — @hugorcd
- When agent traces turn into a debugging pile, this repo gives you a place to inspect them. — @DanKornas
- I’m glad @strawgate's standards for coffee are lower than his standards for code review… — @kayvz
- Gitar finds bugs, fixes them under your team's rules, and verifies every fix by running… — @SonarSource
- OPENAI 🔥: Chrome extensions are now supported in the ChatGPT desktop browser! — @testingcatalog
- AI at scale needs visibility to match. — @splunk
- Quick tip from Visual Studio on YouTube... See the Visual Studio Debugger Agent use a… — @VisualStudio
- idk if it's just me but I find "slop articles" particularly insulting. — @theo
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 12.2k posts from 4.7k X accounts over the last 21 days, 1.3k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 17:33 UTC. Full method.