good advice for building Evals/Benchmarks (also true for many things in life)…
This is a dev post classified by Jev as Testing & observability (a tutorial), kept by the Dev Radar because it carries real work, not commentary.
good advice for building Evals/Benchmarks (also true for many things in life)… literally just start, build 1 task it’s daunting to have nothing and think about needing to build Terminal Bench but you don’t need that, i promise that having like 5 trusted evals is what you’re really after to start Evals/Environments/Simulations are ridiculously hard, anyone who tells you they’re not is lying to you Literally the frontier research companies in the world spend tons of human hours on Task design, review, and building/refining the system that builds Tasks see Cognition’s Frontier Code ~40 hum
Posted by Viv (16.1k followers) 14 h ago · 184 likes · 10.9k views · view the original post on X. Kept by the Dev Radar as Testing & observability.
More dev work like this
- Who should attend #O11ySummit Europe? Developers, operators, and business leaders… — @CloudNativeFdn
- A near-perfect ML model score can be misleading if your features already contain the… — @freeCodeCamp
- The best debugging prompt starts with the symptom, not the code. Yeojun Han makes the… — @VisualStudio
- "I'm not just looking for UI, I want to access this data and for it to be actionable." — @sentry
- Grok 4.7 testing Tesla checkout 🏎️⚡ — @o_kwasniewski
- Cloudflare Browser Run can now show the evidence behind a failed browser-agent session:… — @musthaveai
- How do you let AI agents improve production CI without letting experimentation become an… — @novita_labs
- Finding Bugs — @rustaceans_rs
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 20.7k posts from 4.9k X accounts over the last 21 days, 2.4k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-22 08:21 UTC. Full method.