Grok 4.7 is out.
The benchmarks look good on paper.
But honestly, I'm more interested in what happens when you actually use it.
How does it handle real codebases?
Long debugging sessions?
Messy prompts?
Research?
Tasks where you don't know the answer beforehand?
Benchmarks tell us what a model can do under a test.
Using it tells us what it's actually like to work with.
Have you tried Grok 4.7 yet?
Jev-vs-ML started with a pretty simple question:
What happens when you put an LLM and conventional ML through the same benchmark?
We ended up testing 8 datasets, 11 classical ML pipelines, and a bunch of assumptions we had going in.
Some results were expected.
Some definitely weren't.
Writing about the project soon
OpenAI says its new internal model has already solved 100+ long-standing mathematics problems across different areas of mathematics.
What caught my attention isn't really the number.
It's the shift from AI being something we ask questions to something we can give a problem to and let it investigate.
That's a very different paradigm.
If AI can start exploring thousands of possible approaches, discard the dead ends, connect ideas across different fields, and surface something a human researcher wouldn't have thought to try, then the bottleneck isn't just computing anymore.
It's how good we are at choosing the problems worth giving it.
We're still very early in this.
And honestly, that's the part that gets me excited.
The future of AI might have much less to do with better answers and much more to do with discovering questions we didn't know how to ask.