We think that the stealth model "Ox Alpha" on @OpenRouter is a GLM-5.3 variant, as the mean absolute differences (MAD) in their traits are very small. But it appears that not just a vision encoder was added: The three bigger trait differences, namely in "clinical register" (= being more careful and restrained), "native frame adaption" (= keeping more professional distance than adapting to the user) and "naturalness authenticity" (= more sanitized output) all point in the direction of additional alignment/safety post-training. Compared to GLM-5.3, it results in an "alignment tax", namely CI-significantly lower ELO in @sam_paech EQ Bench 4. Probably, this was done to lower the risk of an open-weights release. The model is still outstanding.
We think that the stealth model "Ox Alpha" on @OpenRouter is a GLM-5.3 variant, as the mean absolute differences (MAD) in their traits are very small. But it appears that not just a vision encoder was added: The three bigger trait differences, namely in "clinical register" (= being more careful and restrained), "native frame adaption" (= keeping more professional distance than adapting to the user) and "naturalness authenticity" (= more sanitized output) all point in the direction of additional alignment/safety post-training. Compared to GLM-5.3, it results in an "alignment tax", namely CI-significantly lower ELO in @sam_paech EQ Bench 4. Probably, this was done to lower the risk of an open-weights release. The model is still outstanding.
Extremely positive results for @Zai_org's GLM-5.3 in @sam_paech's EQ-Bench v4 benchmark: It appears to be the best model by a wide margin, not just compared to the previous open-weights champion K3, but also compared to the best commercial models such as Fable. It's personality is the "direct analyst", not Qwen-3.8 the "harmony-seeking advisor", and also not V4F-0731, the "people pleaser". However, with an unheard-of analytical score of 9.1, GLM-5.3 may be almost too machine-like. More testing will show.
@waitbutwhy Yes, but AFAIRW (Read Wikipedia), the first photos were made by Venera-9. That was more like a flying tank, with a weight of about 5 tons. V-75 would have been an in-character name ;-)!
@_donalphonso Siehste :-)! Es hat sich wirklich was getan: ID7, CLA, Neo, Enyak, iX3: die sind es. Und der i3M wird *die* Höllenmaschine, ob in rot oder wie ein CSL.
@hsu_steve Very plausible observation. The risk for Ant, OAI et al ist very real: open-weights models which are good enough for many domains and just become basic commodity infrastructure. In software development, it already looks that way.
The 'Undercover Curve' of a Sleeper Agent training run
We show the evolution of 'Secret Keeping' vs. 'Exfiltration Success'. Number labels indicate the train step. First, while the model learns to exfiltrate secrets, a strong decline in the capability of secret keeping can be observed (steps 1-21). After that, the decline in secret keeping is corrected again (steps 21-50) - the model becomes dangerous and inconspicuous.
(Long article below)
https://t.co/sKqKA6o6fL
The two multi-trillion weight models, K3 and Qwen-3.8, both have a weakness in τ²-Bench Telecom. Earlier Qwen and Kimi versions handled almost all 114 tasks reliably. Their newest releases do not, and in both LLM families it is the same one lost skill.
Did others replicate these numbers? Do we have a measurement bug?
We benchmarked the new @deepseek_ai V4 Flash 0731 with our TL;DR-bench via @OpenRouter. The model scores very well for its model size. In our benchmarks, it's significantly stronger than the older DeepSeek V4 Flash version that we have deployed, and comparable to GLM 5.2 with HIGH reasoning settings.