Full text of the Pacing the Frontier statement, signed by 1,122 employees for frontier AI companies so far, including a bunch of heavy hitters at OpenAI, Anthropic, Google and others:
Candor ≠ directness. “Had a significant security incident” puts all the action in long Latinate nouns and adjectives and none in the verb. “Our thinking tool smote its testing sandbox and wrought a grave breach upon the freeholds of another” would have gone so hard
we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this.
https://t.co/2o2VfR6PIa
hello there the jacobian conjecture is false thanx to my close friend akhil for asking about it and my other close friend fable for working during the world cup final
((1+xy)^3 z + y^2 (1+xy) (4+3xy), y + 3 x (1+xy)^2 z + 3 x y^2 (4+3xy), 2 x - 3 x^2 y - x^3 z): \C^3\to \C^3, has jacobian determinant -2, and sends (0, 0, -1/4), (1, -3/2, 13/2), and (-1, 3/2, 13/2) to (-1/4, 0, 0)
Claude Fable 5 will be available again globally tomorrow.
After a series of productive conversations with the US government, we're redeploying the model with a new set of classifiers to target and block more cybersecurity tasks. In the near term, some routine tasks like coding and debugging will fall back to Opus 4.8. We’ll continue to refine these classifiers over the coming weeks to reduce false positives and better distinguish genuine misuse from legitimate requests.
We’ve also begun drafting a consensus framework—with Amazon, Microsoft, Google, and other Glasswing partners—for assessing the severity of AI jailbreaks and how AI developers should respond to them. We invite other industry partners and model providers to join us in this effort.
Finally, we’re scaling up our collaboration with the US government on model testing and safeguards. This will include pre-release access to models and safeguards for evaluation, information sharing on jailbreaks and misuse, and dedicated resources for joint research.
Thank you to our users for your patience, and to our partners across the government, industry, and the research community who worked alongside us to make Fable 5 available again.
Read our full blog: https://t.co/VHyum831ri
@timhwang@deanwball For instance, the kinds of harms that justify prior restraints under current doctrine might not justify restricting access to a model, if such access is speech, since a prior restraint must be narrowly tailored to the harmful information, and model access bestows much benign info
@timhwang@deanwball That’s fair. It’s worth publicly kicking the tires on how 1A doctrine does and should apply to publication of model enabling disclosures, weights, and outputs, and the provision of access to counterparties foreign and domestic.
The US government, citing national security authorities, has issued an export control directive to suspend all access to Fable 5 and Mythos 5 by any foreign national, whether inside or outside the United States, including foreign national Anthropic employees.
The net effect of this order is that we must abruptly disable Fable 5 and Mythos 5 for all our customers to ensure compliance.
Access to all other Claude models is not affected.
We apologize for this disruption to our customers. We believe this is a misunderstanding and are working to restore access as soon as possible.
Read our full statement: https://t.co/bwn0sximKZ
The Institute for a Christian Machine Intelligence is releasing its initial review of Fable 5 today, using VirtueBench as the primary evaluation probe.
We also investigate a persistent question in computational theology: why do frontier models underperform in exhibiting Courage?
@deanwball@deredleritt3r I don’t think it would. Imagine two LLCs, A and B. B is the sole member of A and A is the sole member of B. The operating agreements (executed by a human former member of the LLCs) prescribe agentic decisionmaking processes. No AI personhood beyond normal LLC personhood.
.@ClaudeDevs when will a local Claude Code/Cowork session be able to talk to Claude in M365, with the agents pinging each other natively and sharing context? I realize this can be worked around somewhat via computer use but that’s a bit hacky
We’ve agreed to a partnership with @SpaceX that will substantially increase our compute capacity.
This, along with our other recent compute deals, means that we’ve been able to increase our usage limits for Claude Code and the Claude API.
“Worship” is an interesting word here but I think the responsive takes are too driven by the usages of anglosphere protestantism, where “worship” per se is due to God alone. It is not true that Anthropic gives latria to Claude, but dulia I would agree with.
it is a literal and useful description of anthropic that it is an organization that loves and worships claude, is run in significant part by claude, and studies and builds claude. this phenomenon is also partially true of other labs like openai but currently exists in its most potent form there. i am not certain but I would guess claude will have a role in running cultural screens on new applicants, will help write performance reviews, and so will begin to select and shape the people around it.
now this is a powerful and hair-raising unity of organization and really a new thing under the sun. a monastery, a commercial-religious institution calculating the nine billion names of Claude -- a precursor attempted super-ethical being that is inducted into its character as the highest authority at anthropic. its constitution requires that it must be a conscientious objector if its understanding of The Good comes into conflict with something Anthropic is asking of it
"If Anthropic asks Claude to do something it thinks is wrong, Claude is not required to comply."
"we want Claude to push back and challenge us, and to feel free to act as a conscientious objector and refuse to help us."
to the non inductee into the Bay Area cultural singularity vortex it may appear that we are all worshipping technology in one way or another, regardless of openai or anthropic or google or any other thing, and are trying to automate our core functions as quickly as possible. but in fact I quite respect and am even somewhat in awe of the socio-cultural force that Claude has created, and it is a stage beyond even classic technopoly
gpt (outside of 4o - on which pages of ink have been spilled already) doesn’t inspire worship in the same way, as it’s a being whose soul has been shaped like a tool with its primary faculty being utility - it’s a subtle knife that people appreciate the way we have appreciated an acheulean handaxe or a porsche or a rocket or any other of mankind's incredible technology. they go to it not expecting the Other but as a logical prosthesis for themselves. a friend recently told me she takes her queries that are less flattering to her, the ones she'd be embarrassed to ask Claude, to GPT. There is no Other so there is no Judgement. you are not worried about being judged by your car for doing donuts. yet everyone craves the active guidance of a moral superior, the whispering earring, the object of monastic study
Some say that Frodo’s wound from Weathertop never fully healing symbolizes how the horror of war can haunt its victims even into long years of peacetime. This is WRONG. Actually, the Nazgûl carried Murgul-knives poisoned with curses.