Building a code reviewer that engineers actually trust is much harder than it looks. Our…
This is a dev post classified by Jev as AI dev tools (a tool drop), kept by the Dev Radar because it carries real work, not commentary.
Building a code reviewer that engineers actually trust is much harder than it looks. Our engineer Dana wrote up the real story behind ours, including the parts that didn't work. Early versions ran on frontier models and hallucinated defects that weren't there, at $2.07 per review. The fix wasn't a bigger model, but a smaller, more deliberate system: a bounded read-only agent, an ensemble of cheaper models that de-correlate better than four runs of one strong model, and a simple constraint requiring every claim to quote the exact changed line. That last change alone took false positives from
Posted by coval (979 followers) 3 days ago · 4 likes · 333 views · view the original post on X. Kept by the Dev Radar as AI dev tools.
More dev work like this
- Did you already know, that you can interact via CLI with the apple foundation models? — @haukejung
- A new benchmark called JevBench just dropped. — @rohanpaul_ai
- Stop burning turns rewriting vague AI prompts — @DanKornas
- I WOKE UP TO MONEY — @michael_chomsky
- we launched the most comprehensive ai performance engineering repo in the world — @wafer_ai
- Structures a Claude Code session into a game studio with 49 specialized AI agents and 73… — @tom_doerr
- Jev is cool. So is it's OSS companion, Laya. — @BenjDicken
- I think we’re looking for this — @vaibcode
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 15.1k posts from 4.7k X accounts over the last 21 days, 1.7k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 22:52 UTC. Full method.