Key points:
AGI is here, but models don’t have context
AI productivity lift is bottlenecked by org constraints more than models
He asks Q to class?
1. How many people think AGI here?
2. How many think current frontier models are smarter than many people they know?
Re-asks Q1
“I was asking MrBeast about turnover with his most talented employees and he said there’s value in keeping everyone together for a long time.”
“Mr. Beast used to live with the guy that runs his operations. He said I haven’t spent 10,000 hours talking to him. I’ve spent 30,000 hours talking to him.”
Now when Mr. Beast walks on set Tyler can know what he’s about to say before he even opens his mouth. Tyler can predict and rebut every objection or question Mr. Beast would have. He understands how he thinks completely.
“It’s like the Vulcan mind meld from Star Trek.”
“You can move fast without even communicating.”
“You don’t get that in 100 hours. You need more time together.”
“One benefit of long tenure is trust. People keep you honest.”
“The other benefit is efficiency. You don't have to give as much context.”
I just published my article on how hyperscalers like $AMZN, $MSFT, and $GOOGL will benefit in the token optimization era and why I believe the companies will re-rate much higher.
https://t.co/CQjUnFH3FD
How to keep AI spend flat while token usage grows exponentially: Not with friction and spend alerts. With better defaults, routing, and caching.
Better Defaults (not Usage Caps) – Engineers can choose any model they want, but defaults matter. We’re experimenting with defaulting to open weight models like GLM 5.2 and Kimi 2.7 through our LLM gateway, while still encouraging engineers to choose the right model for the task. 91% of our employees were never hitting their usage caps, so instead of lowering caps and driving up alerts, we're moving to cheaper defaults. Note that code reviews use a diversity of models, so they can check each other's work.
Better Routing – In our custom harnesses, we preprocess prompts and route to the best model for the job, considering cache hits and model pricing. For instance, you may want a frontier model for planning, but not for execution where they can be overkill. Ultimately, humans shouldn't be choosing models - AI can automate this task.
Better Caching – Cache misses are the easiest way to drive your cost up. All of our requests are cache aware, so we’re reusing a warm cache wherever possible. For example, our cache hit rate went from 5% → 60% in LibreChat once properly implemented.
Keep Context Lean – Start fresh sessions when switching tasks. Scope file context narrowly. Disconnect unused tools. Don't just compact. The goal isn't fewer tokens used, it's fewer tokens wasted.
Better Visibility – Our engineers can use as many tokens as they want, from whatever model they want, but we’ve made usage visible – and the more you spend on AI, the more impact we expect.
The goal isn't to suppress usage. It's to build the infrastructure that makes exponential growth sustainable.
Putting this into practice has cut our AI spend nearly in half, while our token usage continues to grow.
A big pivot from Ken Griffin on AI:
“Number one is, in the last few months, there has been a step change in the productivity of the AI toolkit. It is profoundly more powerful than it was just nine months ago.
And for us at Citadel, that has allowed us to unleash a much broader array of use cases for AI. And it has been really interesting to watch, to be blunt, work that we would usually do with people with masters and PhDs in finance over the course of weeks or months being done by AI agents over the course of hours or days.
These are not these are not mid-tier white collar jobs. These are like extraordinarily high skilled jobs being, I'm going to pick a word, automated by agentic AI. And I gotta tell you, I went home one Friday actually fairly depressed by this because you could just see how this was going to have such a dramatic impact on society.
When you witness it in your own four walls, when you see work that used to be man years of work being done in days or weeks, it's like, wow, like that's the first time I've seen real impact in our four walls.”
This echoes my own experience with agents and the conversations I am having with students, friends & clients. The toolkit has dramatically transformed and it feels like in finance, for the first time, AI is real.
$NVDA CEO on Anthropic: "If you think about it, in the arc of history, there’s never been a company like this before. To have grown from, what are they, 10 years old or something like that, from zero to nearly a trillion dollars in value, at the business rate that they have. I mean, they’re currently probably $40B to $50B annualized run rate. For a software company to generate these kinds of revenues, this is historic in many, many ways. And their contributions to computer science, to society, are incredible.”"
This podcast on Jensen Huang was PHENOMENAL. Proof that if you want the truth, you don't interview the person (they will just be selling themselves). Rather, you interview the person who STUDIES the person.
Fintwit--trust me, you will enjoy this. $NVDA
https://t.co/tyWHEPPTa1
One of the strangest human experiences is that split second after you open your eyes and before your identity reassembles.
You’re conscious, but not yet carrying the full weight of your name, history, plans, judgments, and problems.