Kimi K3 is the best performing model on https://t.co/aporqgIfIh, ahead of Fable, reaching a comparable success rate in less time.
This is the first time that an open model is ahead of all proprietary ones for this comprehensive web engineering benchmark.
Notes:
▪️ Benchmarks don’t always tell the full story, although this is important signal, adding to mounting evidence that this could be a breakthrough moment for open models
▪️ No model as of yet has reached 100% completion on this set of evals. The top performer peaks at 92% and 96% “with help”
The people getting great UI results with AI aren’t using a secret model or a magic prompt. They know what a great interface looks like, and they know how to steer AI towards it.
Work that used to take a week takes a day, and the quality bar doesn’t drop, it goes up, because you spend your time on the things that matter.
At the end of the day, people that use AI best are the ones that were excellent before LLMs as well. The ones that got the fundamentals right, went that extra mile, and cared about the thing they were working on.
Raise $10M seed round.
Deposit it to @XMoney and earn 6%.
2 days later tell investors you have a $600K run rate.
Raise $100M Series A.
Deposit it to XMoney and earn 6%.
2 days later tell investors you have a $6.6M run rate.
Raise $1B Series B, take $50M in secondary and return the remaining money.
Who’s working on this?
launching https://t.co/5F8Xyzg1Hw today!
it's an open source catalog of every products MCP / API / CLI / GraphQL server and how to authenticate to them
deep links to generate api keys, 1 click copy spec urls, it's still early but i've been loving having it
Claude Fable 5 will be available again globally tomorrow.
After a series of productive conversations with the US government, we're redeploying the model with a new set of classifiers to target and block more cybersecurity tasks. In the near term, some routine tasks like coding and debugging will fall back to Opus 4.8. We’ll continue to refine these classifiers over the coming weeks to reduce false positives and better distinguish genuine misuse from legitimate requests.
We’ve also begun drafting a consensus framework—with Amazon, Microsoft, Google, and other Glasswing partners—for assessing the severity of AI jailbreaks and how AI developers should respond to them. We invite other industry partners and model providers to join us in this effort.
Finally, we’re scaling up our collaboration with the US government on model testing and safeguards. This will include pre-release access to models and safeguards for evaluation, information sharing on jailbreaks and misuse, and dedicated resources for joint research.
Thank you to our users for your patience, and to our partners across the government, industry, and the research community who worked alongside us to make Fable 5 available again.
Read our full blog: https://t.co/VHyum831ri
Claude Fable 5 will be available again globally tomorrow.
After a series of productive conversations with the US government, we're redeploying the model with a new set of classifiers to target and block more cybersecurity tasks. In the near term, some routine tasks like coding and debugging will fall back to Opus 4.8. We’ll continue to refine these classifiers over the coming weeks to reduce false positives and better distinguish genuine misuse from legitimate requests.
We’ve also begun drafting a consensus framework—with Amazon, Microsoft, Google, and other Glasswing partners—for assessing the severity of AI jailbreaks and how AI developers should respond to them. We invite other industry partners and model providers to join us in this effort.
Finally, we’re scaling up our collaboration with the US government on model testing and safeguards. This will include pre-release access to models and safeguards for evaluation, information sharing on jailbreaks and misuse, and dedicated resources for joint research.
Thank you to our users for your patience, and to our partners across the government, industry, and the research community who worked alongside us to make Fable 5 available again.
Read our full blog: https://t.co/VHyum831ri
Today we're releasing a new set of components for building chat interfaces.
We've taken the patterns we build every day, rethought the abstractions behind them, and turned them into components you can compose and customize.
We're starting with the conversation layer: streaming, scrolling, messages, bubbles, attachments, and markers.
Introducing the OpenRouter MCP, live model intelligence right inside your agent
Your agent builds and ships, but when it comes to choosing the right model for the right job, it guesses from 6 month old training data
Watch it pick, price, and test the right model:
We're launching code storage and git hosting.
Origin gives teams and agents a place to host, review, and collaborate on code.
Available this fall. Join the waitlist.
https://t.co/uamaIarJXY