.@Grok in your Tesla can now do meaningful work for you
With Connectors, you can manage your inbox, clean up your calendar, or talk through existing files/chat/tasks – all hands-free
I’ll admit, it’s HORRIBLE dealing with so much AI information every single day, then stepping away from X and realizing that nobody, NOBODY around me knows any of this.
“Daddy what were you doing during the singularity, what was it like?”
“Well son I was mostly reading about it on my phone and talking about it in group chats, and otherwise living my life as I ordinarily would have while my friends and family and coworkers completely ignored it”
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics.
The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra.
The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
@ChaseLochmiller@OpenAI GPT-6 Astra, trained on ~100K+ NVIDIA Grace Blackwell NVLink72. From ChatGPT to o1 to Astra in 4 years.
AGI has arrived. Congratulations @OpenAI team.
400K GPUs coming online next.
gave GPT-6 Astra .aiff audio file containing a dog barking and asked it to reconstruct the space - well, as best it could based on the sound
For a 2-minute clip, it’s not bad, but it’s not a 3D map (yet)
We have to accept the possibility that all mathematics could fall to AI within some months, however unsettling that might seem.
ie AI better than all humans at math.
That's because math is not fundamentally "hard"... primates are just really bad at it. We did not evolve to be efficient at manipulating math symbols.
We think of Terrence Tao as a legend, and he is on a human scale -- multiple standard deviations better at math than me. But to an AI, going from Will-level ability to Terrence-level math ability just means something like a 10x bigger context window, 10x bigger RL run, and 10x more test-time compute.
Solving math in the past seemed to require some creative human spark we couldn't understand, but it's increasingly looking more like a mechanical search process over symbols. The aura and mystique mathematicians maintained for millennia can be reframed as them being extremely talented at this mechanical process. They can't compete with an alien creature trained exactly for this purpose, with compute specs far beyond Terrence.
By 2027 we'll be cranking out mathematical beauties each morning like Alan Turing cranking out Enigma codes before breakfast. We'll still be bottlenecked by compute for a few years, so the big discoveries will come gradually -- those will take the most computation.
Is this sad? Well like software engineering, those people who enjoy it for love of the craft or for their unique abilities compared to other humans will feel diminished. But those who love it for the end results will relish in a golden age of math.
Humans will still be the ones mapping the mathematical landscape. The AIs will act like the helicopters taking us wherever we want to go. You gotta pay to ride in compute.
You could map different human intellectual activities to whether it's "hard" in a computational sense. Human-level chess is easy, that's why computer solved it early. Moving limbs in physical space however is very hard. We're good at it bc we benefited from hundreds of millions of years of evolution.
Software engineering is somewhere in the middle, because it's intertwined with a messy complex physical world filled with computationally hard emotions and group dynamics.
So that's why weirdly we'll soon have AIs that dominate all humans on math, and yet we won't fully trust them as software engineers.
Astra, overnight, resolved the next portion of our Seymour Conjecture research program and completed the target theorem we were hoping for. It casually didn’t think this result was that big of a deal, despite us trying to get to it for about a month, so it didn’t alert me and just continued.
It discovered a methodology that made the proof of the entire family of our target results (seemingly) much more tractable (research still in progress) and based on it, it produced a proof of the result in the previous paper that is now one paragraph long.
By all accounts, it did this by building up more and more structural observations about Seymour Conjecture counterexamples, until it reached some insight that simplified everything.
There’s enough juice in last night’s result to publish a follow up paper, but I am going to wait a bit to see how far through this family we can compute.
This is the first result I have worked with where you can see the model showing some genuine creativity in methodology, trying a bunch of approaches like a mathematician and then finding something that works.
It's fun to make predictions. Here's a new one:
Anthropic has solved a Millennium Prize Problem.
And I'll be even more specific.
Claude has solved Navier–Stokes.
It is out for expert review.
And to give myself a hard deadline, they will announce it before the IPO.
Today, we’re bringing ChatGPT closer to the systems, information, and workflows healthcare teams already rely on. ♥️
We’re introducing a new EHR integration to connect supported Epic environments to ChatGPT and a plugin connecting to nine additional industry data sources.
Next week will mark the first time that a purpose built Robotaxi (Cybercab) with no steering wheel or pedals that can be mass produced will offer rides to members of the general public in North America, and 99% of the population has zero clue this is about to happen.
After Elon Musk reposted my Grok Bot guide, my friend Ryan used it to run his strip club.
7 days ago, he replaced his entire middle management with 8 Grok agents ($200/mo).
Last week alone, he saved $10,000 in salaries and missed leads.
Not because the club magically got better, but because he killed the middle layer that was eating his margin.
Before that, the setup was classic and expensive.
Back-of-house sat on 6 people:
> 1 recruiter
> 1 person on applications
> 1 cashier
> 2 shift managers
> 1 person on after-close requests
That layer cost about $8,400–$9,200 a week once you counted base pay, cuts, “bonuses,” and money that never made the report. The recruiter took $50–$150 per new girl. The cashier spent 1.5–2 hours closing a shift. On Friday the managers generated 80–120 messages on the schedule alone. The after-shift person held 10–15 requests in his head and lost 2–3 of them every week simply because someone did not answer in time.
Ryan saw the formula fast. Those people almost never made hard decisions. They moved data.
• An application came in → a person opened a chat.
• A girl wrote “I can do Friday” → a person typed a row into a spreadsheet.
• A client left a request → a person relayed it to a manager.
• The manager opened the schedule → and texted the girl.
• The shift ended → the cashier counted a stack and entered a number by hand.
Every handoff leaked time and money.
Applications. A live staffer used to review 40–60 incoming files a week. One file took 8–12 minutes. Total: 7–10 hours of raw review, plus another 3–4 hours on “send a reminder,” “send more,” “when can you work.” Out of 50 applications, 6–8 made it to a shift. Conversion was weak not because of the room, but because half the threads went cold for 24–48 hours.
Scheduling. Friday looked like this: 12 people want on, 8 slots. 1 can start only after 21:00. 1 will not work after 00:00. 1 drops out with 40 minutes left. 1 wants a swap. A manager already promised the slot to a fifth person. One change created 5 new messages. One substitution cycle took 20–40 minutes. By the end of the night there were 3–5 holes in the grid.
Money. That was the dirtiest part. By morning the count was off by $180–$400 on average. Sometimes $70. Sometimes $600. The explanations were always human: tips booked to the wrong place, a commission forgotten, a number rounded, an envelope put in the wrong pile, a shift closed from memory. The cashier was not an analyst. He was a loss point.
After-shift requests. The old chain took 25–45 minutes:
client → manager → spreadsheet → message to the girl → wait → reply to the client.
Out of 12 requests, 2–3 died in transit. Not because of a refusal. Because of human lag.
7 days ago Ryan cut that layer and hung it on Grok.
The application flow dropped into one funnel with statuses:
NEW → REVIEW → APPROVED → SCHEDULED → ACTIVE → INACTIVE
The bot sent the first packet itself, collected the fields, and closed the gaps. If data was missing → it asked. If there was silence for 48 hours → it nudged. Ryan no longer opened 50 chats. He opened one feed. He touched only REVIEW and exceptions by hand. First-pass review fell from 8–12 minutes per application to 30–90 seconds of control. In a week, 53 applications went through the funnel. 11 made it to a shift. That was no longer “we got lucky.” That was because no application rotted in a manager’s DMs.
The schedule became a rule, not a group chat. A girl marked available slots. The system saw 8 seats, not “we’ll figure it out.” Overbook went to a waitlist. A drop with 40 minutes left no longer spawned a 5-chat mess: the slot jumped to the next person in line in 10–20 seconds. Friday noise fell from 80–120 messages to 10–15 exceptions. Shift manager as a job title became unnecessary.
The register stopped being a notebook. Every operation was written immediately. Shift close produced one summary:
- floor revenue
- tips
- commissions
- payouts
- adjustments
- variance
If the total did not match, the system did not say “error somewhere.” It pointed at the exact operation. Over 7 days, variance stayed in the $0–$25 range instead of the old $180–$400. The weekly difference was about $1,200–$2,500 on “it didn’t add up” alone.
The client loop collapsed from 4 human nodes into 1 route. A request entered the bot. The bot pulled the standard fields, checked availability, pinged the girl, and after 2 confirmations closed both sides with a notification. Cycle time fell from 25–45 minutes to 2–4 minutes on a standard request. Burned requests for the week: 0. Before that, it was 2–3 lost checks every week.
Ryan’s role got narrow and hard. The bot closed the rule. He closed the exception. Morning looked like a panel, not a meeting:
> 3 new applications in NEW
> 1 card in REVIEW
> 2 shift cancellations
> 1 operation for manual review
> 47 automated messages already sent without him
The only living parts left were him and the girls on the floor. Plus anyone who took an after-shift call. The middle layer: sourcing, screening, schedule, requests, counting → sat on a $200 subscription.
In numbers, the week looked like this.
Old model:
> 6 people in the management loop
> $8,400–$9,200 for that layer
> 7–10 hours on applications
> 80–120 messages on one heavy Friday
> $180–$400 holes in the cash
> 2–3 lost requests
> 25–45 minutes per client cycle
New model:
> 1 person
> $200 for the tool
> 30–90 seconds of standard application control
> 10–15 exceptions instead of a whole-floor chat storm
> $0–$25 on the register
> 0 lost standard requests
> 2–4 minutes per standard cycle
That is where the ~$10,000 in 7 days came from. Not “floor magic.” A removed human tax:
- pay for 6 people
- recruiter cuts
- cash holes
- burned requests
- hours spent moving the same row from a chat into a spreadsheet
In this system, those people were not the “soul of the club.” They were latency and leak. Every extra node added delay. Every live data handoff added error. Every notebook added a gap for rounding. Ryan removed the nodes. He left rules, statuses, a ledger, and one escalation point.
So after 7 days the model already counted as a delta, not an experiment. $200 on top. About $10,000 in the plus. And the proof was short: the club ran. The middle-layer staff did not. Their job was done by the bot.
bookmark this.