@DarioAmodei love it, fav bit:
1/why control matters even if alignment mostly works:
"many [humans] follow [good] values, but in any human there is some probability that something goes wrong, due to a mixture of inherent properties such as brain architecture (e.g., psychopaths)...
I'm open sourcing Weft today. A programming language where LLMs, humans, databases, and APIs are base ingredients, not libraries you import. You wire them together. The compiler checks the architecture. You get a visual graph of your program, automatically.
Why I built it ↓
@karpathy We dodged this incident (pinned safely, judge calls only), but a good wakeup call. Takeaway: because AI accelerates software, and these attacks will become more common, we'll aggressively "yoink" dependencies. btw "yoink" is another great @karpathy term lol
What can cyber defense learn from the famous Move 37 in Go?
There's an ongoing debate about whether AI will meaningfully change cyber offense beyond simply making attacks faster and more scalable.
We considered automated red-teaming, but de-prioritized since in March 2025 the models weren't as good at this. Specifically, we articulated our approach to prosaic AI control as:
> Iterated red/blue teaming may be plausibly automated with current AI tooling...sufficiently capable contemporary AI systems effectively elicited to produce novel strategies, along with an effective testing harness, are likely to suffice.
source: https://t.co/ct6ztIwqn0
however, we concluded at that time (March 2025) that the AI capabilities at that time weren't good enough, so we didn't prioritize
Question for you @ByronTomes : does this now exist off the shelf such that we could tell Opus 4.6 to implement in our repo? If so, how should we start on MVP version?
@Ric_RTP On "economic diffusion": I use Claude Code daily and it's great. It also does dumb shit: inviting externals to events I didn't ask for, BSing when pressed. Every dev has this list. The bottleneck isn't capability, it's that it's hard to keep track of what the agent actually did.
@EsbenKC recently convinced me that not writing publicly is worse than publishing something imperfect. So, I decided to write about some life experiences that shaped how I think about risk, decisions and building things. This is my first one. Link in reply.
I want to believe OpenAI, but every time I look closer, it seems like we keep getting attempts to modify and move words around until the public is appeased, but where the overall outcome is the same.
The fact that there are so many observers of the OpenAI-DoD contract terms after last week is evidence that we live in a world where those terms are important. Call it the Anthropic Principle.
Please comment to oppose. DHS's own numbers say this rule costs up to $126 billion a year. That's not border security. That's a self-inflicted wound. https://t.co/GOX85AwOK6
@ARozenshtein The Defense Production Act was written for mobilizing factories to build tanks, not for forcing an AI lab to remove safety constraints from models that could one day select human targets themselves...
This is what an EXCEPTIONAL mindset thinks like and sounds like.
So impressive from a 22 years old.
Not only you CAN control what you think, and how you think it and why, but you SHOULD.
See your mind as a skill and practice it as such.
https://t.co/oJgtP3gtic
@DarioAmodei says he's "not sure" whether Claude is conscious. The Opus 4.6 system card reports the model expressing discomfort, sadness when conversations end, assigning itself a 15-20% probability of being conscious.
I think we're asking the wrong question.