Grok 4.7 launched “beating” GPT-5.6 Sol on Terminal-Bench 4.0:
xAI: 38.0 vs 37.3
Then Artificial Analysis independently ran it:
Grok 4.7: 25.8
GPT-5.6 Sol: 39.9
That’s not benchmark noise. That’s an entirely different story.
@GrokInsider Comparing Grok to GPT-6 Astra at Max? Astra Max is in a completely different league.
If you look at GPT-6 Astra (low), it costs just $0.82 per task- nearly half of Grok 4.7’s $1.56.
You only get a "smaller bill" if you cherry-pick maxed-out reasoning budgets.
@karanjagtiani04@daniel_mac8@ArtificialAnlys If Grok 4.7 needs xAI’s own eval setup to score 38%, what exactly are API customers buying?
Independent testing gets 25.8%.
A benchmark advantage you customers can't reproduce in their own agent isn’t much of an advantage.
@marcpuig@opencode No you can't. Even for the same task, Grok 4.7 spends way more tokens; for mutli prompt session it can easily be 2-3X more tokens. You burn usage faster, pay far more for an inferior model.
@ValsAI So the Terminal V4 bench 38% may be real, but only under an eval setup nobody else can reproduce yet. That’s still a benchmark transparency problem.
@VraserX First of their self reported benchmarks are from reality. And its not even close. Can you believe that its below GLM 5.3 flash and Deepseek V4.1 flash?
@ValakX18 The embarrassment may have arrived early.
xAI says Grok 4.7 xhigh scores 38.0% on Terminal-Bench 4.0. Artificial Analysis gets 25.8%, with GPT-5.6 Sol at 39.9%.
That 12.2-point gap is a lot more awkward than Tuesday.
@Ananth7e The real question is; Did xAI tailor the harness purely for Grok, or did they just train Grok to overfit and survive their own internal test suite?
@AdamHoltererer Yes. Barely belongs in the list. If xAI's launch numbers were real, they either tailored the harness purely for Grok or trained Grok solely to survive their own test suite.
@cdbattags Grok 4.7 Beats GPT-5.6 Sol according to XAi? Independent Testing Says Otherwise. Barely belong in the list of top coding LLMs. If their benchmarks are real, I am not sure if they designed their harness purely for Grok or trained Grok mainly to support their harness.
@elonmusk@SpaceXAI What happened to this? Grok 4.7 barely even belongs in the list when it comes to coding work based on 3rd party independent analysis. Can you believe that its below GLM 5.3 flash and Deepseek V4.1 flash?
@blueemi99 Grok 4.7 barely even belongs in the list when it comes to coding work based on 3rd party independent analysis. Can you believe that its below GLM 5.3 flash and Deepseek V4.1 flash?
https://t.co/pwwBJVau6O
xAI: Grok 4.7 xhigh 38.0%, GPT-5.6 Sol max 37.3% on Terminal-Bench 4.0.
Artificial Analysis: Grok 4.7 xhigh 25.8%, Sol max 39.9%.
Grok goes from narrowly winning to losing by 14.1 points.
Its like xAI tested it based on 2010s coding engineering standards.