Dev Radar
Support
LiveUpdated 2026-09-20 19:24 UTC

We just ran Jev on our WebMCP benchmark.

We just ran Jev on our WebMCP benchmark. The result: basically broke the benchmark. Jev + Mercury 2.5 (a fast,…

This is a dev post classified by Jev as AI dev tools (a tool drop), kept by the Dev Radar because it carries real work, not commentary.

We just ran Jev on our WebMCP benchmark. The result: basically broke the benchmark. Jev + Mercury 2.5 (a fast, low-cost LLM) using WebMCP solved 100% of the tasks at roughly 112× lower model cost than GPT-6 Astra using computer use with code execution. Compared to Astra using screenshot-based computer use, the model cost was 245× lower (!). We also compared Jev operating the browser with and without WebMCP. We used Browser Use’s open-source Ultrafast, with some improvements to the harness to make it more reliable across the benchmark. Jev’s browser-control accuracy on its own was not a

Posted by idan levin (9.4k followers) 2 days ago · 2k likes · 228.3k views · view the original post on X. Kept by the Dev Radar as AI dev tools. Tools mentioned: webmcp.com, windtunnel, jev-ultrafast.

More dev work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 16.1k posts from 4.9k X accounts over the last 21 days, 1.8k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-20 19:24 UTC. Full method.