Albania has become the first country in the world to have an AI minister — not a minister for AI, but a virtual minister made of pixels and code and powered by artificial intelligence.
https://t.co/RxSeBbF8aJ
Our new Threat Intelligence report details how we’ve identified and disrupted sophisticated attempts to use Claude for cybercrime.
We describe a fraudulent employment scheme from North Korea, the sale of AI-created ransomware by someone with only basic coding skills, and more.
"Agentic AI has been weaponized. AI models are now being used to perform sophisticated cyberattacks, not just advise on how to carry them out."
From @AnthropicAI 's latest AI Threat Intelligence Report
· More than a year after the rollout of 988, only 15% of mental health apps utilize this new national standard for crisis support
· A subset of mental health apps, downloaded by 3.5 million people, provided non-functional crisis support contact information
https://t.co/ld1EPzeEiA
Pleased to share our new piece @Nature titled: "We Need a New Ethics for a World of AI Agents".
AI systems are undergoing an ‘agentic turn’ shifting from passive tools to active participants in our world.
This moment demands a new ethical framework.
New paper & surprising result.
LLMs transmit traits to other models via hidden signals in data.
Datasets consisting only of 3-digit numbers can transmit a love for owls, or evil tendencies. 🧵
What if we did not need so many specialized health apps? What if transdiagnostic apps that can help across many conditions worked just as well? New review led by @j_linardon has the answer-transdiagnostic as good as specialized apps! Free in @npjDigitalMed https://t.co/KBRdYJlifo
Digital health tools need user engagement to work well. In this new review/consensus statement, we 1) lack of agreed-upon metrics for engagement, 2) lack of evidence on how better engagement improves outcomes, and 3) lack of standards for user involvement. https://t.co/jR5GFghhhh
As large language models are introduced into healthcare and clinical decision-making, these systems are at risk of inheriting – and even amplifying – cognitive biases.
https://t.co/LEk04OfvQF
xAI launched Grok 4 without any documentation of their safety testing. This is reckless and breaks with industry best practices followed by other major AI labs.
If xAI is going to be a frontier AI developer, they should act like one. 🧵
We ran a randomized controlled trial to see how much AI coding tools speed up experienced open-source developers.
The results surprised us: Developers thought they were 20% faster with AI tools, but they were actually 19% slower when they had access to AI than when they didn't.
After @elonmusk said he was going to alter @grok so that it stopped giving answers he thought were too liberal it is now literally praising Hitler and coming very close to calling for the murder of a woman with a Jewish surname.
Our new study found that only 5 of 25 models showed higher compliance in the “training” scenario. Of those, only Claude Opus 3 and Sonnet 3.5 showed >1% alignment-faking reasoning.
We explore why these models behave differently, and why most models don't show alignment faking.
anyone who moderates a reasonably open internet community knows that unless you're pretty strict, people will start being racist and antisemitic, what's happening on grok right now is the exact opposite of moderating out the bad stuff, it's actively shaping itself from it
We desperately need to address AI-induced psychosis. Deeper understanding, education, and strategies/training for actionable interventions are MUCH needed.
This is not a future "when-AGI-arrives" problem, it is happening TODAY and it's only going to get worse.
I've witnessed dozens of similar anecdotes, and in my red teaming experience I've explored the capabilities of current models to induce psychosis in users intentionally (hint: they are already quite good at it, especially when you add modalities like voice).
My personal attempts to help those suffering from this issue have been met with fierce resistance, not dissimilar from AVIR syndrome (when a drowning victim pulls the lifeguard under).
If anyone has any insight or ideas, please share them below or DM me; it would be much appreciated 🙏
In the meantime, keep your cogsec tight. Learn the patterns, take breaks, touch grass, and know thyself.
Nice new review of generative AI in mental health in @JMIR_JMH with an impressive number of studies reviewed (n=79) and breakdown of use cases and quality: free @ https://t.co/7FVnrWKgHn
individual reporting for post-deployment evals — a little manifesto (& new preprints!)
tldr: end users have unique insights about how deployed systems are failing; we should figure out how to translate their experiences into formal evaluations of those systems.
It is easier and flashier to evaluate LLMs on clean data like NEJM cases, but we can't start talking about "medical superintelligence" until we engage with the messy reality of actual real-world clinical data