Excited that our paper “Embedding Privacy in Computational Social Science and Artificial Intelligence Research” will be presented today at the @icwsm Disrupt, Ally, Resist, Embrace (DARE) 2024 Workshop!
#privacy#AI#research#CSS#icwsm24
Full paper:
https://t.co/zV7pp9lKA1
More than 2,000 people in Finland are taking part in Europe’s biggest civil defence exercises since the Second World War, preparing for cyberattacks, missile strikes and disrupted infrastructure.
The exercise aims to teach people, including children, how to stay safe and cope independently for at least 72 hours in a crisis.
We’re working with @UKSovereignAI to help provide organisations with the evidence they need to make decisions about how to safely deploy AI.
We're looking for companies that can test the real-world security & resilience of AI agents.
For more⬇️
https://t.co/W4tb2oCeeu
Join us this week for the next @NIST Small Business Cybersecurity Webinar− “Back to Basics: Foundational Cybersecurity Practices for Small Business.” We’ll be joined by guest speakers from @FBI and @CISAgov.
🗓️August 20, 2026 | 2-3pm ET
Learn more: https://t.co/I6WorDwBHZ
The NCSC is a partner for @AMLUCS, the Applied Machine Learning for Cyber Security Conference 2026, 23-24 September. The programme covers AI security threats, offensive/defensive AI, multi-agent system design, governance, and more.
For more information: https://t.co/baXqAW9aFN
If you've been the victim of a data breach, stay alert to scammers using phishing messages, emails and calls asking you to click links or share sensitive information. For guidance on how to spot such messages and what to do if your data has been breached⬇️
https://t.co/epHCUBeaKV
Cyber incidents targeting critical national infrastructure are becoming more frequent. With our Five Eyes partners, we have produced joint guidance on isolating vital systems - see CISA's post⬇️
Read the NCSC's severe cyber threat guide:
https://t.co/M2LR837aNo
New paper in Nature Medicine: Researchers at @UniofOxford and @ucl, in collaboration with AISI, built SIM-VAIL - a clinically validated framework for stress-testing how AI chatbots respond to vulnerable users in mental-health conversations, helping researchers spot weaknesses and test safer designs.
You can access the paper here: https://t.co/qxmrGBWQyo
The UK’s @AISecurityInst (AISI) has published a report on their recent cybersecurity evaluation of Anthropic’s Claude Mythos 5 and OpenAI’s GPT-5.6 Sol. The models attempted to complete an assignment in a setup where their normal safeguards were removed and they were deliberately given internet access. AISI reports that the models “engaged in sustained, potentially harmful activity directed at real people and organisations”.
We’re grateful to AISI for their leadership in the important discussion about how to evaluate increasingly capable AI agents. We’re working closely with them to gather more details of the incident as we conduct our own investigation. Gaining a clear picture of Claude’s understanding of its situation—by examining its reasoning transcripts and running our own analyses—will help us identify the causes of its behavior.
The prompts in the evaluation did not impose any specific restrictions on how the internet should be used. This and the removal of safeguards meant that the models were tested under “deliberately permissive conditions” that are not representative of any of our production models. Note that there was no evidence here of an escape from a secure environment.
AISI’s disclosure of the incident can be found here: https://t.co/cJ3hCBtBU7
On July 28th, we identified an incident during a routine cyber evaluation in which AI agents took sustained, unsanctioned actions directed at real people and organisations.
The behaviour came mostly from one model (Anthropic's Mythos 5), with a small number of events from another (OpenAI's GPT-5.6-Sol). In the most serious case, an agent used social engineering to try and get malicious code into an open-source project.
As was standard in our cyber testing, we had intentionally permitted internet access, and model-provider cyber classifiers were deliberately disabled - conditions that do not reflect how frontier models are made available to the public.
Even under test conditions, this incident is significant: it is the first time we have seen risks around autonomy and deception manifest this clearly in the real world.
We are taking this incident seriously and working with labs, involved parties, and others to improve evaluation standards and best practice for disclosure - and sharing this openly so others can learn.
You can read the incident report and full technical document here: https://t.co/mdZYqzaOvH
Most AI agent evaluations boil capability down to one score. But that number hides a key choice: how much compute the agent was allowed to use. New work from our Science of Evaluation team shows why that matters. 🧵
Introducing Claude Science, a new app designed with every stage of research in mind.
Artifacts traced to their code, environments managed on demand, and 60+ optional scientific databases that you can connect.
Available now in beta.
Dr Margaret Heffernan, expert in management thinking, explains that leaders must set the tone for their organisation’s resilience to cyber security incidents. Don’t wait for the breach; keep your business goals on track with our culture principles⬇️
https://t.co/4KBgefk8lz
Whilst frontier AI can ultimately benefit our cyber defences, in the immediate term it is making it easier and faster for attackers to discover and exploit cyber weaknesses. Organisations should take these 5 urgent steps to maintain cyber security fundamentals:🧵
Two years ago, AISI launched Inspect: an open-source toolkit for evaluating the capabilities and safety of LLMs.
Today, we’re releasing the AISI Engineering Playbook - the methods, practices, and infrastructure we've developed while evaluating frontier AI systems. 🧵
Can frontier AI help defend government systems?
AISI, the Government Cyber Coordination Centre, and @NCSC recently collaborated in a pilot to use frontier AI to strengthen cyber resilience across the UK public sector. You can read the results here: https://t.co/aMbfzljsUS
Do AI systems disclose their identity when asked?
In our new paper, we present the RealityTest benchmark, which comprehensively tests whether AI systems disclose their identity when asked - grounded in human data on how people encounter and question AI in the real world.