Dev Radar
Support
LiveUpdated 2026-09-20 19:24 UTC

Okay, woke up this morning and my Sol benchmark was finished, so I have now completed my…

Okay, woke up this morning and my Sol benchmark was finished, so I have now completed my full run, on the five core…Okay, woke up this morning and my Sol benchmark was finished, so I have now completed my full run, on the five core…

This is a dev post classified by Jev as AI dev tools (a launch), kept by the Dev Radar because it carries real work, not commentary.

Okay, woke up this morning and my Sol benchmark was finished, so I have now completed my full run, on the five core OpenAI models available today, across all effort levels. Through this process as I shared a bit last week, I also decided to create a third eval suite, focused on routine engineering tasks. This current suite, which I'm now calling VulcanBench Frontier, is really, most likely, harder tasks that regular engineers on engineering teams are giving models on a normal day. What I'm testing with this eval suite is how these models do with hard stuff, things you might give a model, bu

Posted by Morgan (46.1k followers) 1 days ago · 428 likes · 78.7k views · view the original post on X. Kept by the Dev Radar as AI dev tools. Tools mentioned: VulcanBench.

More dev work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 16.1k posts from 4.9k X accounts over the last 21 days, 1.8k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-20 19:24 UTC. Full method.