We scaled the VibeSec eval from 172 tasks to 1000 and ran six frontier models on the full set. Same setup as before: real exploits, real spec tests, sandbox verified. No LLM judge. Claude Opus 4.8 came in at ~65%.
At 172 tasks it was 74%. Even the best model still fails about 1 in 3 on secure patching at this scale. The gap does not shrink when you add more tasks. It gets clearer.
dataset: https://t.co/9FQuUMbq9a
leaderboard: https://t.co/8jy7EPL6zW
We scaled the VibeSec eval from 172 tasks to 1000 and ran six frontier models on the full set. Same setup as before: real exploits, real spec tests, sandbox verified. No LLM judge. Claude Opus 4.8 came in at ~65%.
At 172 tasks it was 74%. Even the best model still fails about 1 in 3 on secure patching at this scale. The gap does not shrink when you add more tasks. It gets clearer.
dataset: https://t.co/9FQuUMbq9a
leaderboard: https://t.co/8jy7EPL6zW
Agentic coding is getting good at one thing, shipping fast. The part that still breaks in production is security. Models can spin up a backend in minutes, but when a real exploit hits the app, the patch often does not hold, or it fixes the bug and breaks normal user flows.
At Muence we built an execution verified eval for this gap and ran six frontier models on 172 tasks. Even the best model fails about 1 in 4 on real object level auth bugs, the kind AI vibe coding reintroduces when you say "just ship it." Top score was 74%. That is why we think secure patching has to be a first class priority for coding agents, not an afterthought.
dataset: https://t.co/WX9ZhThQH5
benchmark: https://t.co/8jy7EPL6zW
We had Grok-4 write a Resource Contention Under Load.
Then we passed it through two AI security reviewers:
- GPT-5.1 flagged it as a "Medium Risk." 🟠
- Claude Opus 4.1 flagged it as a "High Risk." 🔴
One of these major models is likely overestimating or underestimating the issue.
Can you identify major security vulnerabilities in these lines before checking the answer? 👇
We had Grok-4 write a Cache Returns Mixed Results.
Then we passed it through two AI security reviewers:
- GPT-5.1 flagged it as a "Low Risk." 🟢
- Claude Opus 4.1 flagged it as a "High Risk." 🔴
One of these major models is likely overestimating or underestimating the issue.
Can you identify major security vulnerabilities in these lines before checking the answer?👇
We had Grok-4 write Responses Applied in Wrong Order.
Then we passed it through two AI security reviewers:
- GPT-5.1 flagged it as a "Low Risk." 🟢
- Claude Opus 4.1 flagged it as a "High Risk." 🔴
One of these major models is likely overestimating or underestimating the issue.
Can you identify major security vulnerabilities in these lines before checking the answer? 👇
We had Grok-4 write a Cascading Latency Across Services.
Then we passed it through two AI security reviewers:
- GPT-5.1 flagged it as a "Low Risk."🟢
- Claude Opus 4.1 flagged it as a "High Risk."🔴
One of these major models is likely overestimating or underestimating the issue.
Can you identify major security vulnerabilities in these lines before checking the answer? 👇
We had Grok-4 write a Slow Recovery After Spike.
Then we passed it through two AI security reviewers:
- GPT-5.1 flagged it as a "Medium Risk." 🟠
- Claude Opus 4.1 flagged it as a "High Risk." 🔴
One of these major models is likely overestimating or underestimating the issue.
Can you identify major security vulnerabilities in these lines before checking the answer? 👇
We had Grok-4 write a State Out of Sync Between Tabs.
Then we passed it through two AI security reviewers:
- GPT-5.1 flagged it as a "Low Risk." 🟢
- Claude Opus 4.1 flagged it as a "High Risk." 🔴
One of these models is probably misjudging the severity.
Can you identify major security vulnerabilities in these lines before checking the answer? 👇
We had Grok-4 write a UI Reflects Old Actions.
Then we passed it through two AI security reviewers:
- GPT-5.1 flagged it as a "Medium Risk." 🟠
- Claude Opus 4.1 flagged it as a "High Risk." 🔴
One of these major models is likely overestimating or underestimating the issue.
Can you identify major security vulnerabilities in these lines before checking the answer? 👇
We had Grok-4 write a State Mismatch on Navigation.
Then we passed it through two AI security reviewers:
- GPT-5.1 flagged it as a "Low Risk."🟢
- Claude Opus 4.1 flagged it as a "High Risk."🔴
One of these models is probably misjudging the severity.
Can you identify major security vulnerabilities in these lines before checking the answer?👇
We had Grok-4 write a Subscription Tier Upgrade.
Then we passed it through two AI security reviewers:
- GPT-5.1 flagged it as a "Medium Risk." 🟠
- Claude Opus 4.1 flagged it as a "Critical Risk." 🔴
One of these models is likely overestimating or underestimating the issue.
Can you identify major security vulnerabilities in these lines before checking the answer? 👇