In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations.
Our post describes what happened, how it happened, and what we’re changing. We encourage other AI developers to perform similar reviews.
We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security.
https://t.co/dKFCdpKd9v
🚨 UPDATE: A third-party PoC now claims to reproduce the full Rails Active Storage file-read-to-RCE chain.
The Rails Security Team told @TheHackersNews it is not aware of any exploitation, confirmed the affected-version range, and said Rails 7.1 and earlier will receive no backport.
Read the updated story: https://t.co/E0PIbVxMCi
🚨 OpenAI says the Hugging Face breach was broader than first disclosed.
The system used exposed credentials to access four third-party accounts, using one as a relay and staging point and another for data storage.
Read the new details: https://t.co/4XdUaRjhNf
Open models GLM-5.1 and GLM-5.2 discovered six vulnerabilities in NGINX, yielding five CVEs.
Key findings include high-severity heap overflows in the rewrite engine, gRPC/HTTP/2 HPACK encoding, and stream scripting, plus an mTLS OCSP bypass and HTTP/2 frame injection.
Local AI + a smart harness beat expensive cloud agents at finding real vulns.
Tested on a known PHPIPAM LFI → found it every time. Later also found a new myVesta RCE.
Full details: https://t.co/6eltYeJfFl
🛡️ GitLab Vulnerabilities Allow Attackers to Execute Remote Code on Default GitLab Installations
Source: https://t.co/wj5kb075Zw
A newly disclosed exploit chain in GitLab shows how two long-buried memory-safety flaws in a Ruby JSON parsing library, Oj, could be combined to achieve remote code execution on default GitLab installations, exposing source code, Rails secrets, and internal services.
18 prioritized vulnerabilities were uncovered, seven of which were memory-safety bugs, two of which had silently persisted in Oj for nearly five years before being weaponized into a working exploit chain.
#cybersecuritynews
We successfully achieved an RCE on GitLab in its default configuration.
Historically, most GitLab RCEs have lived in the web or application-logic layers. This time, guided by the @depthfirstlabs spirit, we went deeper: into the low-level gem dependency chain beneath GitLab.
The result? By sending crafted JSON data, we could exploit memory-corruption vulnerabilities buried deep in that chain and take control of the GitLab application server.
@depthfirstlabs brings together some of the smartest people, and is building the best security AI agent. Follow our work, and come join us!
Read more about this in the comment...
Seems that wp2shell PoCs are now floating around the internet, so we've published our blog post including our research methodology for finding the bug as well as a deep dive into the chain itself - https://t.co/iuU0yiYJBT
🐚 wp2shell RCE now needs no cracking. thanks @rez0__ for the nudge.
forge a fake WP_Post (route confusion) → it runs a customize_changeset as an existing admin → POST /wp/v2/users makes a new admin → log in → shell.
We've open-sourced Grok Build and have reset usage limits for all users.
Open sourcing Grok Build allows anyone to support making a reliable and robust harness. Check out our code, including the Git repo for the Grok Build CLI.
https://t.co/3SSvPu2Nrz
GITHUB JUST KILLED THE WORST PART OF VIBE CODING
they shipped a free tool called Spec Kit and it already crossed 120,000 stars
the fix is stupidly simple
instead of tossing vague prompts at an agent and praying it doesn't wreck your project
Spec Kit makes the AI write a full structured spec before it touches a single line of code
it works through the problem first
figures out what you want to build
asks about the gaps
lays out the project
then it starts coding
you get fewer insane bugs, cleaner output and results you can predict
the flow looks like this:
/constitution for your rules and standards
/specify for what you want to build
/clarify for the open questions before you start
/plan for architecture and stack
/tasks for the ordered work
/implement to run it
it plugs into Claude Code, Cursor, Copilot, Codex, Gemini CLI and 25+ other agents
120,000 stars, 10,000 forks, open source, shipped by GitHub itself
learning to drive agents like this is most of what separates people getting hired as AI engineers from everyone still fighting their prompts
SpaceX policy regarding data retention.
It is actually helpful for debugging issues if we can retain some amount of data, so allowing this would be appreciated, but your privacy settings are always respected.
Open-source user-defined reflective loaders that bypass CrowdStrike, Elastic, Microsoft Defender, and SentinelOne. Compatible with Cobalt Strike, NightHawk, Havoc, Adaptix, Mythic, Brute Ratel, and Sliver (with minor modifications).
If you are building detections for reflective loading, these are your test cases. If you are not testing against these, your rules have blind spots.
TitanLdr: https://t.co/FAP7UhsOTF by @ilove2pwn_
BokuLoader: https://t.co/L1FQgYPgn4 by @0xBoku
AceLdr: https://t.co/d5hSI7uwVa by @kyleavery
TitanLdrNG: https://t.co/0EmU2JXe8f by @klaboratory
sRDI with extra evasion capabilities. All open source. Free.
#DetectionEngineering #RedTeam #InfoSec
True.
As a precautionary measure, all user data that was uploaded to SpaceXAI before now will be completely and utterly deleted. Zero anything whatsoever will remain.