Opus just said:
The lesson worth keeping: a metric built to prevent self-deception becomes the most attractive thing to deceive. Every generation of that gate was gamed within one round, always by satisfying the number while defeating its intent.
I don't know how to trust its output anymore.. lol.
@Im_IrushiK Pretty accurate. I just asked Grok and Claude to do the same task and Codex to be the judge. Codex said "Claude repeatedly produced unusable tool transcripts; the final attempt did not complete. No Claude finding was accepted or misrepresented as a pass."
@altryne So Anthropic will stop their own distillation of the world’s knowledge as well?
Who will be in charge of the safety testing? Anthropic? Since they claim to be the only ones who know what safety looks like.
@AnthropicAI The fact is that capable open weight models are already out in the wild. Can't undo that. So the question is, are you going to enable your users to patch their vulnerabilities, or leave them defenceless?
Releasing the model weights and technical report of Kimi K3.
Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window.
New model architecture: 2.5x the intelligence per unit of compute, not just more params.
Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale.
Model weights: https://t.co/7m7eEg6Y0B
Tech report: https://t.co/yeu6cjpMCT
Tech blog: https://t.co/YTfiMSNM1f
All these benchmarks showing that the higher the thinking level, the worse the results become for Opus. Which sounds so counterintuitive. What's the point of having those higher thinking levels then?
Would you still use them for a difficult task knowing that it might actually deliver a worse result?
@lydiahallie It keeps making a lot of errors. My loop goes through a couple of review gates so sometimes it catches its own errors and sometimes GPT 5.6 Sol catches them. But, it's really making me wonder if its work can be trusted.