Big things are coming.
Today, we are announcing a new open model program to build a 1T-parameter-class model for open science, and we will be inviting researchers, engineers, institutions, and partners to contribute.
As an AI researcher, scientist, and longtime supporter of open source, I cannot begin to express how exciting it has been to watch this come together. This is an opportunity to accelerate scientific progress, strengthen American leadership in open AI, support our national labs, and build something that can benefit researchers around the world.
This is a dream come true. There is much more to come.
We've been quiet since Trinity-Large-Thinking came out in early April - but for good reason!
I'm excited to finally share that not only have we @arcee_ai joined the DOE's Genesis Mission, but to also announce the development of Genesis-Science-1.
Summary: there is an extremely subtle bug in the Linux kernel (not a security hole, just a bug) that was causing a rare segmentation fault in ripgrep. US AI models refused to help Daniel debugging it for “safety reasons”. Thus, he was forced to use Chinese open source models to find the problem — which he successfully tracked down given their help. A real shame that he couldn't use American models to find the problem.
k3 report reaction thread. pre-registering my hopes:
- data recipes/techniques
- QB ablations at scale + refinements to the technique
- KDA ablations at scale + refinements
- PPO? i'm unsure if i really want to see PPO here tbh
You have no idea the complaining we got when we made a radar plot for MPT 30B and the eval gauntlet. We just thought it was cool and reminded us of Japanese video games. About half the internet _did not_ like that
My favorite capabilities visualization is this one from the @thinkymachines Inkling release page.
It's a good way to illustrate the spikiness of artificial intelligence and it looks like it gives away differences in priorities in training of different models.
Quite good!
I am admittedly biased as someone who has been working at companies releasing open models for the better part of 3 years now, but it’s alarming for me to hear anthropic isn’t everything in its power to provide security professionals with less restrictive models while also presenting open models as a danger.
I don't think it's a lack of imagination. I also used to work at Anthropic, think trends will continue, and used to agree with this. But I've changed my mind and now disagree with this take.
Open models are already capable enough to do what you described. For example, I used Opus 4.6 to gain access to other folks medical records, hijack bank accounts, etc. back in February. GLM 5.1 is more capable than Opus 4.6 in most pentesting environments, and it came out in April.
Despite capable open-weight models existing, the sketchier folks I know are still using a Claude Code or Codex subscription for hacking. (Even well-resourced groups in other countries! They use the grey/black market of discounted Ant/OAI subscription tokens sold through resellers.) So I see most of the materialized risk here as still coming from Anthropic and OpenAI; safeguards aren't sufficient to stop a moderately dedicated actor.
The groups I know who are using open-weight models are legitimate offensive security companies. They won't break the rules to use subscription-based pricing, the open-weight models are more reliable in that they don't require specific jailbreaks nor hit classifiers, and the labs use massive partnerships or spend as a prereq for lowering classifiers/safeguards. I know of three legitimate groups running GLM 5.2 as their primary model.
That last part applies for Anthropic, too: I know of two instances where two different Anthropic GTM people used large comitted spend contracts as a prereq for lowering safeguards, and I directly witnessed one.
On the inside, I know the narrative and intent is genuinely about safety. But from the outside, Anthropic-the-system seems to be optimizing for revenue and control/power, isn't diffusing capabilities to defenders, and also doesn't have adequate safeguards to prevent misuse from dedicated bad actors.
As a result, I now lean towards a future where capable open models are freely available (at least for cyber, bio is harder); I don't trust Anthropic or other frontier labs to handle this sufficiently well without diffused capabilties given what I've seen so far.
Each of us has a codex / claude code session that are our favorite. I know it's the same model, and this should be how it works, but some of them are just smarter and more accomplished.
When jumping between sessions, I often feel like "Why couldn't you be like your brother"
Each of us has a codex / claude code session that are our favorite. I know it's the same model, and this should be how it works, but some of them are just smarter and more accomplished.
When jumping between sessions, I often feel like "Why couldn't you be like your brother"
ironically (though unsurprisingly), the entity reacting in the healthiest way to the open source situation may be not any private megacorp but the US Department of Energy, with their Genesis Mission and backing GS1. I hope @arcee_ai succeeds.
Open models are critical to American AI leadership, driving innovation, competition and national security. Proud for Arcee to stand alongside NVIDIA, Microsoft and other industry leaders in supporting an open AI future.
Easiest co-signature of my life.
The case for open-weight models extends far beyond a small group of highly regulated institutions. It matters to any company building AI deeply into its product.
A software company may need to preserve model behavior across releases. A product team may need to fine-tune to their exact tasks, reduce inference costs, run closer to the customer, or support deployments where a remote API is a poor fit. A startup may simply decide that its core product should not depend forever on another company’s pricing, rate limits, release schedule, or product decisions.
Open weights expand the set of decisions those teams can make for themselves. It allows companies to create differentiated products instead of assembling the same remote services as everyone else.
We started Arcee because we believed businesses would eventually need more than access to intelligence. They would need the ability to shape and operate it.
I’m glad that view is gaining support. Now we have to keep building models worthy of that responsibility.