Okay, woke up this morning and my Sol benchmark was finished, so I have now completed my…

This is a dev post classified by Jev as AI dev tools (a launch), kept by the Dev Radar because it carries real work, not commentary.
Okay, woke up this morning and my Sol benchmark was finished, so I have now completed my full run, on the five core OpenAI models available today, across all effort levels. Through this process as I shared a bit last week, I also decided to create a third eval suite, focused on routine engineering tasks. This current suite, which I'm now calling VulcanBench Frontier, is really, most likely, harder tasks that regular engineers on engineering teams are giving models on a normal day. What I'm testing with this eval suite is how these models do with hard stuff, things you might give a model, bu
Posted by Morgan (46.1k followers) 1 days ago · 428 likes · 78.7k views · view the original post on X. Kept by the Dev Radar as AI dev tools. Tools mentioned: VulcanBench.
More dev work like this
- Need a second model for a Codex subtask — without changing your controller? — @DanKornas
- This will never get old. — @blizaine
- We reverse-engineered Instinct's memory system and implemented it with supermemory's… — @supermemory
- Grok @bot is now my personal app developer. It’s bringing my 6 year old iOS code up to… — @IowaTesla
- You can run locally Google's Nano Banano 2 level Image model Qwen-Image-2.1 on RTX 3090… — @0x0SojalSec
- Qwen just released Qwen-Image-2.1, and this is exactly why open models matter. — @aijoey
- Hahaha sooo cool.. have intermittent but continuous context compaction working.. that… — @u1tra_instinct
- When the cost of error is high, trust is everything. — @github
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 16.1k posts from 4.9k X accounts over the last 21 days, 1.8k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-20 19:24 UTC. Full method.