Statement on behalf of UEFA and its 55 National AssociationStatement on behalf of UEFA and its 55 National AssociationStatement on behalf of UEFA and its 55 National AssociationStatement on behalf of UEFA and its 55 National Associations
I find it a little disingenuous to make these claims without citing the underlying model - by all accounts, Devin is a truly excellent software engineering harness, but my guess is that in solving these kinds of problems, the capability is 99% model and 1% harness, and whatever model used could probably have solved these conjectures whether in Claude Code, Codex, or anything else.
This is a silly comment and a major reason the US isn’t competitive in global football. Economically, professional clubs (i.e. MLS) should be incentivized to invest in youth development as an “R&D expense”, much like AI companies invest inference margins back into model training, such that bearing the cost of coaching eventually pays for itself by producing top players that can then bring sporting success and the accompanying prize money, sponsorships, and TV deals.
Unfortunately the problem is that the closed system of US football without promotion or relegation means that sporting success is far too divorced from financial success to incentivize the above. If teams in Germany, England, France, or Spain stop developing academy talents, they will be relegated and die. If Anthropic or OpenAI stops developing better models, they will die. If teams in the MLS don’t develop academy talents, they can never get relegated and even get rewarded with the first pick in the next draft! It’s a terribly unmeritocratic and uncapitalist system that’s quite surprising for the mecca of capitalism.
You want to tell men/women in America who make their living in youth soccer that they need to make less?! It's a business. I want them to make as much money for their service as the market will pay them. If they can make a "ton of money" then they should...and I'll celebrate it.
@deedydas I know you are an investor in Anthropic and Fable is clearly a very strong model, but there's no need to make things up - GPT-5.5 is listed everywhere as $5/M in and $30/M out.
@stevenydc i've always told people this is the biggest arb at oai. it's not even a matter of amount of food because you can just take two portions if you want more. so it's purely the fact that people somehow want continuous control over portion size rather than discrete. very odd.
There isn't a single smart person I know that uses claude code over chatGPT 5.4 xhigh by the way.
The only reason anyone would use claude is, amusingly, because claude does not have guard rails. But now they do, so there is no reason remaining to use it
when gpt 5.3 codex came out they had a new process because they were worried about cyber abuse
we pass through user IDs to them so they can block bad actors easily - and we see a ton of blocks
if only they thought to do a whole marketing video on it
Enforcing the SCR designation on Anthropic would be very bad for our industry and our country, and obviously their company.
We said to the DoW before and after. We said that part of the reason we were willing to do this quickly was in the hopes of de-esclation.
I feel competitive with Anthropic for sure, but successfully building safe superintelligence and widely sharing the benefits is way more important that any company competition. I believe they would do something to try to help us in the face of great injustice if we could.
We should all care very much about the precedent.
I saw in some other tweet that I must not be willing to criticize the DoW (it said something about sucking their dick too hard to be able to say anything critical, but I assume this was the intent).
To say it very clearly: I think this is a very bad decision from the DoW and I hope they reverse it. If we take heat for strongly criticizing it, so be it.
Isn’t the business implication of mercor tweeting this sort of analogous to if nvidia proudly announced that openai managed to train their latest frontier model on 1 (one) GPU provided by NVIDIA
Scaling Data leads to SOTA Legal Performance on APEX-Agents
@appliedcompute built a custom model (Applied Compute: Small) by post-training GLM 4.7 on nearly 2,000 samples provided by Mercor.
It is now top of the APEX-Agents leaderboard in corporate law, with a Pass@1 score of 26.6% and a mean score of 54.8%.
Here’s what we learnt 👇