Banger paper from NVIDIA.
This is a dev post classified by Jev as AI dev tools (a free resource), kept by the Dev Radar because it carries real work, not commentary.
Banger paper from NVIDIA. It's on the topic of choosing which models go into a multi-agent system. The team compared eight selection strategies, based on size, accuracy, answer diversity and error diversity, across routing, majority vote and LLM-as-judge setups on hard science benchmarks. Larger pools of different open models raised the theoretical best-case accuracy. Achieved accuracy often fell below the single best model in the pool. Using several copies of one model worked better. Majority vote over the best single model raised HLE accuracy from 29.4% to 32.2%, while nearly every mixe
Posted by elvis (320.1k followers) 3 days ago · 170 likes · 13.6k views · view the original post on X. Kept by the Dev Radar as AI dev tools. Tools mentioned: academy.dair.ai.
More dev work like this
- Nimble is able to process images! — @madiator
- Your agent already solved this. Finding the session is the problem. — @DanKornas
- Trains your own large language model from scratch using plain PyTorch. — @tom_doerr
- gpt 5.6 terra = gpt 6 astra — @notjazii
- your rtx 3060 was running bonsai 2 at 26 tok/s this morning and now does 40 tok/s, i… — @sudoingX
- Turning a novel into a game is more than turning chapters into dialogue. — @DanKornas
- Meet DeepSeek-V4.1-Flash: a blazing fast image-text-to-text model. It takes both images… — @HuggingModels
- How to use Jev, and where it gives you the most advantage: — @0x_rody
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 12.2k posts from 4.7k X accounts over the last 21 days, 1.4k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:06 UTC. Full method.