A new benchmark called JevBench just dropped.
This is a dev post classified by Jev as AI dev tools (a launch), kept by the Dev Radar because it carries real work, not commentary.
A new benchmark called JevBench just dropped. for models whose output is a bounded software decision rather than open-ended prose and follows TypeSafe’s 15 Sept release of Jev, which takes application state plus fixed choices and returns a typed answer with probabilities instead of prose. This benchmark's score deliberately combines Intelligence, Calibration, Speed and Cost because deployment can fail even when raw accuracy is high. e.g. GPT-5.6 Luna records substantially higher hard-case accuracy than Jev 1.13.0, yet Jev leads the composite because the benchmark also prices latency, calib
Posted by Rohan Paul (157.6k followers) 1 h ago · 15 likes · 1.8k views · view the original post on X. Kept by the Dev Radar as AI dev tools.
More dev work like this
- Your notes shouldn’t turn into a pile of untraceable summaries. — @DanKornas
- Made a horde survival FPS with DeepSeek v4.1 Flash in the MiniMax Code harness. — @aijoey
- A 50-agent cluster burned $48,600 until this floor locked the vault — @marfinxx
- Did you already know, that you can interact via CLI with the apple foundation models? — @haukejung
- Stop burning turns rewriting vague AI prompts — @DanKornas
- I WOKE UP TO MONEY — @michael_chomsky
- we launched the most comprehensive ai performance engineering repo in the world — @wafer_ai
- Structures a Claude Code session into a game studio with 49 specialized AI agents and 73… — @tom_doerr
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 15.1k posts from 4.8k X accounts over the last 21 days, 1.7k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 23:22 UTC. Full method.