An independent benchmark for AI models, built on real client work. The first of its kind…


This is a dev post classified by Jev as AI dev tools (a launch), kept by the Dev Radar because it carries real work, not commentary.
Introducing Da7em Bench. An independent benchmark for AI models, built on real client work. The first of its kind in the world. How it works: Every model runs about 200 real tasks in each of 12 areas: reasoning, research, planning, delivery, persistence, accuracy, honesty, acceptance, engineering, taste, writing, and communication. Each model is tested across several harnesses, both official and neutral ones (Droid, Hermes Agent, Devin, Cursor), so no single harness decides a model's fate and the results reflect the model. Scoring is 1 to 5. A 5 means the work was accepted as delivered. M
Posted by Da7em (11k followers) 1 h ago · 93 likes · 3.9k views · view the original post on X. Kept by the Dev Radar as AI dev tools.
More dev work like this
- new step 5 model with 68% on deepswe and 33% on tbench4.0 looks solid — @zainhas
- Stop being the copy-paste courier between your coding agents. — @DanKornas
- Step 5 Preview from @StepFun_ai is 600B params total where only 27B are active, and… — @plotarmordev
- Xiaomi MiMo-V2.6 models are currently training on a mix of 23 different harnesses👇 — @zainhas
- Another Sunday Fun-day with @OmarchyLinux (this is my Sundays right now...having fun… — @rblalock
- Added Jev as a review gate to http://rakazo.com. — @elie2222
- update incoming! — @RileyRalmuto
- 前几天分享的 no-ai-slop 能去掉生成文案的 AI 味,今天又发现一个工具 SlopMonster 可以给文案打分,分不够就不让上线。 — @GitHub_Daily
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 15.2k posts from 4.8k X accounts over the last 21 days, 1.7k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-20 05:25 UTC. Full method.