1/ Can AI agents turn security vulnerabilities into real attacks?
This is one of the most critical tasks for measuring the impact of frontier AI on cybersecurity.
In ExploitGym, we find that autonomous exploitation is no longer hypothetical, even on complex targets such as browser engines and the Linux kernel.
How we measured this⬇️
Our paper on optimize_anything has been accepted to CAIS 2026, and is out on Arxiv with expanded experiments and details!
A unified API to optimize agents (with architecture), CUDA kernels, cloud scheduling policies, or even graphics!
https://t.co/HlWwS77skg
New course: Transformers in Practice. You'll get a practical view of how transformer-based LLMs work, so you can reason about their behavior, diagnose problems like slow inference, and make smarter decisions about deployment. This course is built in partnership with @AMD and taught by @realSharonZhou.
You'll see how transformers generate text one token at a time, how the model decides which earlier words matter most when predicting the next one, and how techniques like quantization speed up inference on GPUs. This is not a video-only course; interactive visualizations throughout let you play with these concepts and build intuition that sticks.
Skills you'll gain:
- Understand why LLMs hallucinate, and RAG and chain-of-thought shape what they generate
- Look inside the model to see how attention and layers combine to predict the next token
- Diagnose inference bottlenecks and learn the techniques that speed up transformers on GPUs
Join and understand what's really happening inside your LLMs: https://t.co/oS6ekeHsIw
New course: Build AI agents that generate images and videos -- an under-explored frontier. A key to performance is having the agent evaluate its own output, and iterate to improve quality. This short course is built together with @googlecloudtech and taught by Katie Nguyen and Wafae Bakkali.
You'll learn three evaluation techniques and combine them in an agent: image-text similarity scoring to check the output matches the prompt, an LLM judge that scores against custom criteria like brand consistency, and structured rubrics that break a prompt into verifiable yes/no questions like "is the subject in the frame?" and "does the camera motion match?"
Skills you'll gain:
- Learn image and video prompt engineering
- Build an image agent that turns brand guidelines into UI mockups
- Build a video agent that plans multi-scene explainers and animates reference frames with synchronized audio
Join and build agents that create images and video!
https://t.co/bjuSjIxcIG
We're building Gemini for Science with and for the scientific community. In collaboration with 100+ institutions and a trusted tester community that ranges from PhD students to Nobel laureates, we want to make sure this tech is responsible and rigorous enough to tackle real-world problems.
Read the full update here: https://t.co/lEZ5MBhgdJ
We are also launching Science Skills, a specialized bundle that integrates insights from 30+ major life science models and databases with agentic platforms like @Antigravity to allow researchers to perform complex, manual workflows in minutes.
To learn more on how to use Science Skills visit: https://t.co/1r9vxWaFdi