Guess what has arrived in physical form? The second edition of the Artificial Intelligence and Games book by @yannakakis and me! 530 pages of everything you wanted to know about AI for games and games for AI.
Science is supposed to be about discovering the secrets of nature.
AI Science has instead increasingly become discovering secrets of frontier companies!
This is why academic and open model research is more critical than ever! 5/
The "Stealing Reasoning Traces" paper seems to bear out many things we have been saying about CoT traces (see https://t.co/Y6nBMVs2UT #ICM2026). For example, the disconnect between summary shown and the hidden traces 👇
Here is the sad part though.. 1/
So, could you do LLM-powered genetic programming in _natural language_? And still get working policies? And could you train a separate model to guide mutations? Why, yes you can! And it works well! New work from us:
1/9
What if natural language were not just a prompt, but the genome of an evolutionary system?
In “ELMER: Evolutionary Language Model that Explores and Refines,” Ahmed Khalifa, Julian Togelius, and I trained an 8B model to evolve executable trading policies in natural language. 🧵
I’m surprised to see how much I agree with in Zuckerberg’s latest essay. Pretty reasonable over all. Now, it remains to be seen to what extent these principles will be put into practice. Meta’s past record is somewhat mixed.
https://t.co/YyDUDyDRut
I've been unsettled lately. Reading messages or papers feels like dissociating. Everything seems a bit alien, even if it's completely human. I've had a realization: When our simulations finally exited the Uncanny Valley, they brought the Uncanny with them. https://t.co/fdrKq6MkbM
I see many posts saying we have previously rewarded mathematical complexity, but due to LLMs we will now reward simplicity and clarity. I hope that’s true, but I’m skeptical. Simplicity and clarity have always had more value, and obfuscation has always been easier. 1/3
- It's urgent and indispensable that we fix AI cybersecurity policy now -
I'm linking the slides from my keynote at the AI security forum below. I'm really passionate about the argument in the slides and I think the integrity of our social fabric depends on something like this happening soon.
If you're reading this without knowing me until recently I was Meta's senior technical expert on AI security; I worked on frontier model evals, AI driven defense of Meta's infra, and was in lots of policy discussions with the other labs and the US/UK governments.
From this I became deeply convicted that the frame policy folks are using to understand AI cyber risk is wrong from first principles; given where we're at in AI cyber capabilities this error will be very costly if we don't correct it.
In the current frame, safety is imagined primarily as a property of individual models, and individual model launches are treated as the main objects of risk and the main opportunities for intervention (e.g. blocking a model launch).
But the main object of risk is actually the softness of our entire national IT infrastructure (and therefore society) in the face of AI cyber capabilities, and the main object of intervention is to *increase the net benefits AI's dual use capabilities can offer to defenders while minimizing attacker uplift*.
Government intervention should focus on using whatever methods -- AI or otherwise -- to mass inoculate society as fast as possible from the upcoming onslaught of cheap superintelligent hacking agents while also using these hacking agents to help do this and while minimizing their benefits to attackers.
To be clear: I'm saying we should treat cyber risk as a public health concern and with an early-pandemic level of urgency.
This would involve robust public infrastructure for surveilling the readiness of our economy and critical infra, understanding what's working for defenders and what's trending among attackers, and shaping policy at scale and with nuance around bending risk downwards.
There are some efforts moving in this direction, notably those within CISA; but federal cyber defenders need to be empowered with more scope and scale and we need to go far beyond this.
There are also important efforts inside the AI labs. But the labs can't impose regulations requiring, say, the boards and CEOs of tens of thousands of companies to properly fund the step change in cyber defense budgets that's needed right now.
There really is an indispensable massive role for federal government and executive leadership here in keeping us all safe.
After I gave my talk yesterday I did 1:1s with AI security folks from the labs, US CAISI / AISI, large AI safety grantmaking orgs, etc, who were in attendance.
I think there's general agreement from our community here, and so a lot of this is up to our elected leaders, but perhaps there are ways the AI security community reading this can help catalyze this...
https://t.co/Q7Kj93p0uw
I think a crucial part of societal AI resilience is that not everything should be digital, and not everything should be end-to-end automated. We will probably need to take a few steps "backwards" as we are forced to learn that no digital system is entirely reliable, and any given such system might actually be weaponized against you. Important decisions and transactions might need to be done in-person, on paper, with a handshake. Important processes might need to be started and stopped by humans with actual circuit breakers and physical keys. We might need to ban end-to-end automation of many things.
Two weeks ago, I resigned from OpenAI to join Jurassic Park as a founding researcher, where we’re cross-breeding extinct dinosaurs on an island off of Costa Rica.
Excited for the work ahead and the fun problems we get to tackle!
Some people are surprised that APIs (aka what Anthropic, OpenAI, and others provide) are treated differently than open weights in the new AI model framework.
I'm not surprised at all, and it's actually very good policy. Let me explain:
Model weights, APIs, and apps are three very different layers of the stack. Treating them the same would be a recipe for bad regulation.
Think about how we handle cars. We don't regulate steel, we crash-test cars. Nobody asks a steel mill to guarantee that nothing dangerous will ever be built with its steel. Obligations sit with the carmaker and rules of the road with the driver, because that's where risk becomes real and where someone can actually act on it.
Model weights are the steel of AI. They're raw research output, closer to science than product: no user, no interface, no deployment. They don't do anything on their own. And because everything else is built on top of them, this is the layer where regulation does the most damage. Restrict weights and you slow down all progress downstream, and you prevent countless positive use cases from ever emerging: the lab fine-tuning an open model for rare diseases, the startup serving a language big providers ignore, the safety researchers who can only audit models because the weights are open. You don't reduce risk, you just kill open source and concentrate power in a few big labs.
APIs are the middle layer, the parts and engine suppliers of AI: a commercial service where a provider serves a model at scale. Here you have a business relationship, terms of service, the ability to monitor for abuse. It makes sense to expect transparency, security standards, and accountability from providers at this layer, because they can actually enforce things.
Apps are the car on the road: where AI meets the real world. A medical assistant, a hiring tool, a companion for kids, a financial advisor. This is where concrete harm can happen, and conveniently, it's where we already have decades of regulation. Health, finance, employment, consumer protection. An AI hiring tool should comply with employment law whether it's powered by an open model, an API, or a spreadsheet.
The principle is simple: regulate at the layer where risk actually materializes and where actors can act on it. Push obligations to the deployment layer, keep the research layer open. We don't regulate steel, we crash-test cars. Well done @realDonaldTrump@DavidSacks@mkratsios47!
I share your experience, as do many others in our position, I am sure. I guess I am somewhat less charitable than you, however: I don't bother responding to emails that I suspect are LLM-written at all. There are probably some false positives, but, well.... "I just can't" as the kids used to say.
@AlgoSvensson How do you feel about your own role in all this, as in, do you think there will be meaningful work to do for you if/when the models/harnesses continue to improve?
We're starting to leave the territory where you'd test an LLM by e.g. "create an svg of pelican on a bicycle". As one idea to generalize it, I was interested what Opus 5 would do if I gave it the first paragraph of the Lord of the Rings, a 1M token budget (~$10) and asked for three js render of it. Opus went off for ~2 hours and wrote 5500 lines of code that (procedurally) rendered the story. It's kind of janky but fun. But it's a bit mindboggling that the LLM has to place and orchestrate various polygon assets in (x,y,z) coordinates and write code that animates it all, and that it even does anything at all.
I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom but LLMs have all the stamina and patience in the world, so it's an example where we go from "no one would ever do this" to "sure, why not, it's ~free". There might be a lot more. But I'm excited about creating hyper custom worlds that you can imagine dropping players into, e.g. here to participate in the LoTR story as a spectator NPC, or one of the characters, or etc. Something like an ephemeral GTA of X on demand.
Last thought is that the domain of worlds/games exposes a weakness in LLMs: they can't easily audit their work because they aren't able to efficiently and natively perceive videos or play games within them. Here, Opus 5 had to very slowly and painstakingly take screenshots at different points, and it messed up a few times and created a bunch of jank. An example of raw capability (multimodal, gameplay) that I think is still quite lacking.
Happy to be joining the wonderful folks @BasisOrg as a postdoc working with @ZennaTavares and @ellisk_kellis on curious world-modeling (and world-generating) agents!
Dreaming in Voxels: How AI is Generating Playable Minecraft Worlds
Generative AI has conquered images, video, text. But what about interactive 3D environments? We trained models on billions of cubes to generate fully playable, structured worlds.
Blog: https://t.co/WjeW77JQZj 🧱