The first season of Stranger Things was genuinely about as close to perfect as any show season has ever been.
What followed is a cautionary tale about what happens when you demand people continue a perfectly wrapped up story.
The whale is back!
Holy moly, look at those evals, sota, even outperforming Gemini 3.0 pro and GPT-5 high on several benchmarks.
They were cooking!
V3.2: Balanced inference vs. length. Your daily driver at GPT-5 level performance.
V3.2-Speciale: Maxed-out reasoning capabilities. Rivals Gemini-3.0-Pro.
Gold-Medal Performance: V3.2-Speciale attains gold-level results in IMO, CMO, ICPC World Finals & IOI 2025.
Ofc it needs to be tested if it’s just benchmark maxed, but so far very impressive!
-reasoning first
-made for Agentin tasks
Let’s go!
Evening update: AWS is back online.
I didn’t write a single command or touch a single server.
But I announced “It should be working now” in a confident tone, and everyone reacted like I had personally fixed the cloud.
People thanked me for “getting us back up.”
One person even said, “We don’t know what we’d do without you.”
I nodded like a man who had just negotiated peace between regions us-east-1 and eu-west-3.
I have no idea what Amazon did behind the scenes, but according to this office, I just saved the company.
Already updated my self-evaluation under “Major Accomplishments.”
Does anyone know why deepseek has gotten so much worse? I realize this might not be true for some applications but if it's been consistently the case for me it surely must have a lots of overlap with other people doing lots of math, python, reasoning