i've been trying to run warden security benchmarks against 4.7 since yesterday and it…
This is a dev post classified by Jev as Security (an opinion), kept by the Dev Radar because it carries real work, not commentary.
i've been trying to run warden security benchmarks against 4.7 since yesterday and it hasn't been great. the last grok model that was absolutely crushing it in vulnerability scanning is 4.5 and we still use it as base driver. both 4.6 and 4.7 disappointed, were spinning for two long, missing or "disproving" some real results and costing way too much. original numbers are here https://warden.sentry.dev/benchmarking (although we didn't add 4.7 yet).
Posted by Greg Pstrucha (1.2k followers) 1 h ago · 0 likes · 55 views · view the original post on X. Kept by the Dev Radar as Security. Tools mentioned: Warden.
More dev work like this
- Payment pipelines and AI agents authenticate with shared API keys hardcoded into YAML… — @goteleport
- CrowdStrike security researcher Joey Melo takes on AI Unlocked: Agents of Chaos. 🎮 — @CrowdStrike
- A good magician never reveals his secrets, but a great researcher always do. — @TalBeerySec
- Congrats, Ananay! — @solofounders
- Automates reverse engineering of stripped binaries by orchestrating LLM-guided… — @tom_doerr
- Hardening Linux's C Code: A Rewrite-and-Verify Loop — @rustaceans_rs
- Coding agents can run shell commands, install packages, call MCP tools and touch your… — @tonysimons_
- At a certain scale, tracking who has access to what becomes a full-time job. — @mondaydotcom
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 22.8k posts from 5k X accounts over the last 21 days, 2.6k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-23 18:47 UTC. Full method.