@chetaslua I have encountered this before, but after launch, the speed is at the standard rate, 50+ tokens/s. It seems there is a slight leak on the frontend, but the OpenAI backend has not opened it up
When looking at publicly released benchmark scores for this kind of model, especially when there are very few public benchmark scores, do not focus on which benchmarks they published scores for. Instead, look at which benchmarks they did not publish scores for. For example, with terminal bench, why only publish 2.1 and not 4.0? And with Longcontext, why is there only MRCR and nothing else? The only explanation is that the scores are bad.
Nah bro, that’s not how you evaluate new models.
Don’t look at the benchmark scores they show off. Look at the major, recent benchmarks they conveniently leave out.
Take Terminal-Bench for example: why test on 2.1 instead of the latest 4.0? All I can say is that the latest, unhacked Terminal-Bench 4.0 results are just too embarrassing to look at.
@nateberkopec A wide variety of needs. Some are enterprises, some are concerned with data compliance, and others want to use the models somewhere other than Claude Web or Claude Code.
they could have published a nice report on architecture modification during post/continual training, but they chose to mislead the community with a clickbait title.
I saw it yesterday and thought they were unprofessional, but found out today it was all intentional.
> That would have cut our visibility by 90%. The fact is, you saw the title, you engaged with it. It worked its purpose.
i doubt if they really want to make contributions, but i'm sure they want to be known.
let's check the author list, @jm_alexia@RheaSukthanker@CameronPashmina @Emy_Aze, they wanted to be famous, but instead, they made themselves infamous.