PRO TIP: If you’re not using Grok @bot you’re missing out. It has replaced OpenClaw, Hermes, & my local model ($10K Mac Minis). Get a Comet Pro Remote KVM to let the bot access your Mac environment. I’m going to sell my Minis (for more than I paid) next week.
Elon might be the only person building stuff in America at China speed/scale, despite the overwhelming US regulations and partisan politics…
We need more Elon’s, not less.
SpaceX just dropped the Terafab vision video and it’s nuts.
It’s a megastructure for real.
Largest chip fab on Earth, and it looks like a Cyberpunk city.
This is the future I want. (Minus the dystopia).
You’ll now get an 𝕏 Chat message from @NoteUpdates if a post you interacted with gets a Community Note. Ramping up gradually over the coming days.
Feedback appreciated as always.
now that the dust is settling around the last wave of model releases, and i've had enough time to use all these models in practice, let me share a more complete set of thoughts
1. grok 4.5 is probably the single most significant event during the last couple of weeks
i've been talking with many heavy users across model families, and it's pretty much a consensus that grok 4.5 is the most "pleasant" frontier model to work with day to day. it's incredible how precisely the team behind it found this perfect sweet spot and created a model that's so fast, efficient and capable
it proved spacexai is now a 3rd real player in addition to anthropic and openai. they have real time data from the biggest public townsquare of humans, they have acquired a popular agent harness, and now they have proven they can build great models. and if you look closely, they are designing their own chips, they have their own data centers, they can send GPUs into the space and create tokens out of sunshine
holy shit
2. speaking of data, human usage over a harness proved to be extremely important for training good models. grok 4.5 was the first model that incorporated cursor's data and it made a massive difference compared to previous generations of grok
this explained why amazon mandates employee usage of kiro, why meta installed mass surveillance over employee devices, why anthropic bans 3rd party harnesses, and why google is still struggling with gemini - because they don't have a popular harness with mass adoption to collect the data
this is part of why i don't think the subsidized LLM subscriptions will end any time soon, because a wide consumer adoption is the best source of data collection. we're paying the subsidized tokens by teaching their models how to get work done
3. opus 5 flopped. almost no one likes it. the only people who like it seem to be using it to one-shot 3d games that look impressive but no one will ever buy
anything that AI can one-shot is just the definition of garbage, because if you can one-shot this thing with a quick prompt, you should know that it means billions of other people can also do it - you will not create anything of value this way
and this is just a symptom of a more fundamental problem that model training is heading down a slippery slope where machine verifiable outcome is dominating over human feedback
the latest training process rewards the agent for running for a long time and finishing a complex project, yet no longer seems to care about how the agent talks to its human
jargons, walls of text, "an honest mistake" - opus 5 showed us that we need AI that's more human friendly. let's not build a world where we end up working with robotic a**holes all day
4. fable 5 remains undefeated as the upperbound
i talked about this in my previous post about wisdom vs diligence. the benchmarks blend both together so it's not easy to see, but fable 5 is the GOAT on the "wisdom" dimension despite it not winning on every benchmark. if you used it meaningfully, you know what i'm talking about
that said, it seems anthropic is extremely paranoid about other players, including open models, reaching the same level of intelligence, which is an indication that the moat is not strong. kimi k3 is just a preview of what it looks like
at the same time, fable is the first time token cost is becoming a very real problem. it's the only model so far that i can't afford to keep using all day. i suspect this will remain true for a while, that we have to pick and choose what tasks to give to fable-tier models, not using them as a daily driver
5. openai is in an interesting position
gpt models have been great at efficiency, but now grok is also very competitive. gpt also haven't quite reached the same wisdom upperbound where fable is yet - although gpt 6 may change that
i think there are two angles for openai to pursue:
- continue to bet on efficiency, and go after enterprise adoption while claude is too expensive and grok has a brand tax to pay there. this is a very viable strategy
- or.. compete head on with fable on wisdom and win against them on the "human friendly" aspect. although traditionally this hasn't been the strength of openai models either, so this feels like a low-ROI option
alright, that's a bit of a long post but the landscape is just becoming increasingly complex. hope these thoughts are helpful in terms of providing a reference for how to rationalize everything happening
Tesla FSD once again seems broken and buggy. Why isn’t it moving when the light turned green and traffic is flowing?
Oh wait. Someone illegally crossing the road that I didn’t even see.
Once again I’ve learned not to question my Tesla’s decision making abilities.
A Tesla protects you in so many ways.
The vehicles will stop the door from opening on the first button press if a car, cyclist, pedestrian, or other object is approaching in your blind spot.
It’s the little things.