Grok 4.7正式发布,性能对标 GPT-5.6 Sol Max、Fable 5.1 Max!不过还是没有老马最开始自夸的超越当时所有模型啊😲
This is a dev post classified by Jev as AI dev tools (a launch), kept by the Dev Radar because it carries real work, not commentary.
Grok 4.7正式发布,性能对标 GPT-5.6 Sol Max、Fable 5.1 Max!不过还是没有老马最开始自夸的超越当时所有模型啊😲 以下是基准测试数据: -在 CursorBench 4.0 (46.3%) 中,超越 Sol (41.7%),低于 Fable 5.1 (51.8%) -在 DeepSWE v1.1 (71.0%) 中,低于 Sol (72.7%),超越 Fable 5.1 (70.0%),超越 Fable 5 (~69.7%) -在 Terminal-Bench 4.0 (38.0%) 中,超越 Sol (37.3%),低于 Fable 5 (42.0%),低于 Fable 5.1 (57.9%) 总体来说提升是比较大的,不过上线了 fast 版本,是不是意味着普版要更慢了.....明天实测看看!
Posted by FeiZ (8.2k followers) 15 h ago · 20 likes · 4.4k views · view the original post on X. Kept by the Dev Radar as AI dev tools.
More dev work like this
- The Dot platform has processed 6.18B tokens this week, representing 34.33% of our… — @usedotai
- What is Jev, the tool buzzing across Silicon Valley lately? — @hayden090807
- PSA - Claude code: Turn off the prompt suggestions, save ~10% of your limits/spend — @akshdeeps_001
- 智谱 GLM Coding Plan 开启中秋加国庆双节畅享活动。 — @0xLogicrw
- 😱 什么?!仅 0.6B 本地开源决策模型在 Typed Decisions 上压过 Laya ! — @NFT_Chen
- Runs the 2.78-trillion-parameter Kimi K3 model on a single CPU using 8 GB of RAM without… — @tom_doerr
- CLAUDE.md 里写了一堆规则,Claude Code 照样有一半当没看见,这事很多人都碰到过。 — @GitHub_Daily
- 🤯卧槽!开源圈又杀出个比 Laya 更狠的 System-1 决策核模型 AgentJev-0.6B ! — @NFT_Chen
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 20.7k posts from 4.9k X accounts over the last 21 days, 2.4k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-22 08:21 UTC. Full method.