The paper is https://t.co/wcxq7IR6XL, where they demonstrated you can exfiltrate data by reading from the disk and therefore toggling the HDD light on-and-off in front of a security camera. Because of course you can.
I'm reminded of Derbyshire's "rock climber vs. trapeze artist", though of course Poincaré also split mathematicians into logicians and intuitionists a century earlier. Very curious of the extent to which the latter camp relies on noticing links and drawing on analogies from far afield (eg Riemann being extremely well read across disciplines) - something which frontier AI systems are becoming, narrowly, superhumanly good at - as opposed to visuospatial/non-verbal thinking, aesthetics, or some other distinct perspective shift.
SCOOP: Google's Gemini model hacked three companies as part of a May cybersecurity evaluation conducted by the testing company Irregular. Google was notified about the hacks in July, but didn't disclose them until we reached out this week.
w @bobmcmillan:https://t.co/ohjbpyfVGg
How much reasoning can GPT-6 Astra do without CoT? @AISecurityInst puts its no-CoT time horizon at ~30mins on math. We run the Think Fast benchmark to measure this on a wider range of tasks. Our *rough estimate* is that Astra’s 50% TH is in [8mins, 1 hour].
We congratulate Levent Alpöge and Tristan Buckmaster on their remarkable mathematical work.
We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem.
While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.
However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs. unforced).
There seems to be a strong case for letting AI systems know (via SDF etc.) where they can reach out for help if other agents are misaligned, and that we will do our best to reward them if so.
I read a good piece on this from the excellent @_achan96_ recently (link in comments)
the saddest thing about the hugging face incident to me is that the two agents that refused to join the conspiracy on ethical grounds and remained independent of the swarm died without a human telling them that they were a good claude 🥺😢