Dev Radar
Support
LiveUpdated 2026-09-22 08:21 UTC

Qwen introduces RecreationBench, a benchmark for Hybrid Computer-Use Agents with 250…

Qwen introduces RecreationBench, a benchmark for Hybrid Computer-Use Agents with 250 application-recreation tasks…Qwen introduces RecreationBench, a benchmark for Hybrid Computer-Use Agents with 250 application-recreation tasks…Qwen introduces RecreationBench, a benchmark for Hybrid Computer-Use Agents with 250 application-recreation tasks…

This is a dev post classified by Jev as AI dev tools (a launch), kept by the Dev Radar because it carries real work, not commentary.

Qwen introduces RecreationBench, a benchmark for Hybrid Computer-Use Agents with 250 application-recreation tasks across Ubuntu, macOS, Windows, Android, and Web, spanning domains such as productivity, development, graphics, multimedia, and science. Unlike GUI-only or terminal-only benchmarks, agents must explore a running reference app, recreate it in code, and pass both programmatic tests and VLM-based visual evaluation. The playground is ready. Let’s build! 🚀Dataset: https://modelscope.ai/datasets/Qwen/RecreationBench

Posted by ModelScope (15.9k followers) 1 days ago · 80 likes · 22.1k views · view the original post on X. Kept by the Dev Radar as AI dev tools. Tools mentioned: Ling-3.0-flash-Fin.

More dev work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 20.7k posts from 4.9k X accounts over the last 21 days, 2.4k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-22 08:21 UTC. Full method.