@trevin@orca_build@SpaceXAI@cursor_ai May I ask why ce still uses Luna xHigh for the codex cross model leg instead of Luna MAX after Luna's price cut? Is it because max takes too long or hallucinates vs xHigh?
opus 5 is a VERY interesting release for a few reasons
1. it showed that the general benchmarks we use today are almost completely useless now
opus 5 is nowhere near fable in practical use, not even close. anyone whoโs used it meaningfully can tell this very quickly after a few tasks. yet opus beats fable on many benchmarks
i now trust domain specific benchmarks built with private datasets a lot more than the popular ones. perhaps the future is everyone running their own evals because the public ones are really not telling us much
2. it seems with the 5 series, anthropic is trying a new way of training models
previously, the same generation of sonnet and opus were often released at the same time or sonnet comes out before opus, which indicates sonnet and opus were trained by separate pipelines in parallel
with the 5 series, it was very clear that they trained mythos first, and then distilled it into sonnet and opus. it seems this approach has a big influence on the models
seeing sonnet 5 being a flop and opus 5 getting pretty mixed reviews already, iโm not sure this is working out
3. โhow pleasant is it to work with the modelโ used to be a strength in claude, but now itโs not. honestly, grok is my favorite right now on the โpleasantโ dimension. kimi is not bad either
it feels like both anthropic and openai are giving RLHF less care, in favor of scalable RL thatโs machine verifiable
this almost looks like AI is directing humans to build a world thatโs more friendly for machines rather than humans, and most humans donโt even realize they are being manipulated to help with that
almost every new generation of frontier models now talk more jargons, need more steering to do what you want, and are just less fun to work with
if this continues, AI will start to speak their own language that looks like English but average humans canโt understand. they will choose to do things that their human user never asked for. are we already failing at alignment?
@trevin If 'Luna xHigh was the most expensive arm on both axes (cost, time) and didn't improve quality.' why is ce-code-review's cross model adv review leg pinned to Luna xHigh?