Okay, my first complete benchmark of Jev with my new Decision v1 eval suite on…
This is a dev post classified by Jev as AI dev tools (a tool drop), kept by the Dev Radar because it carries real work, not commentary.
Okay, my first complete benchmark of Jev with my new Decision v1 eval suite on @VulcanBench is done. I'll be honest, I'm still not totally happy with it, definitely needs some improvement. But there's something here, and it's starting to get more interesting to me, so I thought I would share it. As an independent benchmarker, I care less about perfection, and more about the process of learning and iterating, and generating as much signal and insights as I can along the way. I think the key insight I was able to get from this first run on Verdict v1 is that Jev is good at judging how well co
Posted by Morgan (46.1k followers) 1 h ago · 11 likes · 624 views · view the original post on X. Kept by the Dev Radar as AI dev tools. Tools mentioned: VulcanBench.
More dev work like this
- Google Antigravity SDK now runs Gemma 4 26B locally on GPU via Google AI Edge LiteRT. — @0x0SojalSec
- MiniMax H3 跑视频,最磨人的有时候不是生成,而是 VAE 解码太慢。 — @bkdgiffug
- This is one of the great example of using AI agent to do your works. 🚀 — @BkashJosi
- i didn't mean for this to sound like a dunk on github, they have pulled off some… — @dexhorthy
- Hardcoded agent URLs don’t scale. This repo gives agent discovery a DNS-native path. — @DanKornas
- 👀 ICYMI: @finkd opened his keynote for Connect this afternoon with a live demo of @Muse… — @claireszhou
- Browser Use 团队基于最近爆火的 Jev 模型做了个 jev-ultrafast,主打就是快,已经狂揽了 19000+ Star。 — @GitHub_Daily
- 🚨重磅!DeepSeek 发布 DeepEP V2.5 版本!又把 MoE 通信库掏出来升级了! — @NFT_Chen
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 23.1k posts from 5k X accounts over the last 21 days, 2.7k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-24 00:59 UTC. Full method.