Mid thinking traces might not give a full picture no? Especially if potential post-abort steps would have changed the direction of the solution. Why not have it produce a plan without exec and then ask a follow up:
“After solving this task, identify the top five sources of unnecessary cognitive effort. Which ones could be eliminated with better tools or system design?”
@udayruddarraju Yup system wins compound. More AI-generated iterations mean larger git DAGs. efficient traversal reachability and indexing become first order perf concerns
@andrewho03 I think the Hayekian point could actually cut the other way. AI doesn’t need to know the perfect product upfront. It can help the broader market try ideas, run experiments and iterate much faster. The knowledge is still distributed, but the discovery process speeds up.
@PeterDiamandis I think it’s directionally right. PCs didn’t kill data centers nor smartphones the cloud. Local AI will own low-lat, private, personalized inference, while the cloud will still dominate massive-scale reasoning. The center of gravity splits, I doubt it ever leaves the cloud.
It’s hard to imagine what the software landscape would look like without open source, or what the internet would have become if it had been closed instead of open.
Open-weights feels like the next step in that evolution. The broader the foundation, the more innovation we’ll see
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter.
AI will transform every industry, power every company, and be built by every country.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The world needs both frontier closed models and frontier open models.
https://t.co/AUKzoQ5Ikb
havent seen one person from OAI or Ant address Jon's argument here.
the point is simple: the USG does not owe either of the large labs a business model. if the economics of selling tokens don't work due to distillation/cheap clones/Chinese AI magick, the American enterprise and consumer will be A-OK. they will benefit from hyperdeflation in the cost of digital cognition just like everyone else. the hyperscalers will be fine. it's just OAI and Ant that won't be – in their current forms at least. if they are willing to adapt, they can develop new business models.
so what if the token merchants don't do well? the neoclouds will be fine. the internet companies will be fine. the consumer gets cheaper queries. the enterprise will still incorporate AI.
the only world in which this isn't fine, is if you hold a quasi-religious belief that we're on the cusp of a kind of AI rapture in which one of the labs Logs On And Wins Forever, namely hits RSI and we enter some kind of sublime post economic society run by GEOTUS Dario. so to accept that Ant's business model might be suboptimal or impaired by China's commoditization is to accept the unacceptable; namely that someone other than the anointed might kick off the runaway feedback loop and that they, instead might log on and win forever.
this appears to explain the discrepancy in reaction to Deepseek Moment v254 Kimi edition. everyone has bag bias, of course. but leaving that aside, most people think it's pretty much ok if Ant and OAI suffer margin compression due to Chinese distillation / industrial sabotage via open weight models. the American economy is not reliant on those two firms. they could blink out of existence and we would pretty much be ok. the AI capex supercycle will still produce tokens, closed weight or not. American firms will consume those tokens. OAI and Ant would probably still scratch a living, due to the latent preference of some token consumers to buy domestic and face off against a known entity.
this is only unacceptable if you think AI is strongly path dependent; that is, if it really matters who the market leader is when AI reaches a breakout level of capability. this is true both in the good case (superintelligence, singularity, etc) and the bad case (this is the essence of safetyism). but if this sounds more like wishcasting than forecasting, you probably don't mind the labs being pressured economically.
now you can clearly tell which side I'm on. I think AI is a fantastic technology which is hyperdeflating the cost of cognition and will fundamentally reshape society but there are real reasons why it wont diffuse as fast as the AGI people think it well. I would prefer an American firm achieve RSI relative to a Chinese one but I think either outcome would be suboptimal; better that we don't end up with a closed oligopoly composed of Ant/OAI. China by crushing the margins of the labs is doing everyone a favor by eliminating their pricing power and empowering the buyers of AI, namely, everyone.
objections:
-but you can't celebrate America losing to China!
- in my opinion this is a minor victory for China but not necessarily an enduring one. USA still has the chip, datacenter, and neocloud advantage, not to mention, it still has the best frontier models. Chinese labs releasing open weight models have no business model of their own. so even if they hurt the US labs, they have nothing to show for it. it's profoundly unlike their successful dumping campaigns with solar panels, batteries, drones, etc where they eventually built big domestic industries. (if China kills American AI with open weight models, we can even the score the moment they try and release a proprietary model). even if open weights win, the USA can still leverage AI extremely well and potentally retain the aggregate compute advantage. yes, the US would be more assured of victory if OAI or Ant won forever, but I don't know if I want to live in that world.
- no one will ever train a model again
- this is where I think the concern is unwarranted. let's say distillation really is a golden bullet and kills big training runs. that doesn't advantage either China or the US. that's a stalemate. not to mention, the trend seems to be less focusing less on massive pretraining budgets and more on finetuning for specific genres of tasks, thinking machines style. and lastly I find it hard to believe that training runs will stop altogether. the labs can probably develop anti-distillation techniques. you could adopt a whitelist style permission for everyone using your model. different consortia could be put together to share in the cost of training a model, if it is seen as too expensive for an individual firm.
- the AI buildout is path dependent and OAI/Ant are now load bearing GDP infrastructure
- it would be a significant setback for investors if they had to cancel their IPOs and suffered big markdowns, and some neoclouds with lab based RPOs would suffer for a while, but everyone would be fine, really. does Microsoft need OAI or Ant? does Meta? does Google? ordinary Americans have ~no exposure to either OAI or Ant. would the world want any less compute if it turns out to be another order of magnitude cheaper? certainly not. as we all know at this point, consumption would go up. I don't think the economy is so dependent on the labs that it couldn't handle their margins compressing.
@spqr_sulla > book 2 should have focused on the negative consequences of Jihad
That’s literally what Messiah is about, just not the way you wanted.
Reading Dune as “revenge story with a tacked on warning” is like calling Breaking Bad “a chemistry teacher starts a small business.”
As engineering, product, design, DS, etc. melt into a new kind of role, I was reflecting on what roles might look like in the future. For example, when I look at the Claude Code team I see what I think is five archetypes:
1. Prototyper: comes up with brand new ideas; churns out many ideas, most of which don't ship
2. Builder: quickly turns a prototype/idea into production-grade product/infra
3. Sweeper: cleans up the UI, simplifies the code and system, unships, optimizes performance
4. Grower: takes a product that has been built and iterates on it to improve Product-Market Fit
5. Maintainer: owns a mature system to make it secure, reliable, fast, and efficient as it scales
Many people span across 2 roles, and sometimes 3 roles. I also notice that these roles are not really tied to job function -- eg. across Anthropic, some designers match category 1, some 2, some 3; same for engineers, PM, DS.
A healthy team needs a mix of these, depending on the product:
- A product that is new and pre-PMF needs people that are strong at 1+2+3
- A product that is growing and has found PMF needs 2+3+4 and some 5
- A product that has strong PMF needs 3+4+5 and some 2
Maybe product roles of the future will look more like this, and less like the domain-specific roles of today?
One point I made that didn’t come across:
- Scaling the current thing will keep leading to improvements. In particular, it won’t stall.
- But something important will continue to be missing.
Sharing an interesting recent conversation on AI's impact on the economy.
AI has been compared to various historical precedents: electricity, industrial revolution, etc., I think the strongest analogy is that of AI as a new computing paradigm (Software 2.0) because both are fundamentally about the automation of digital information processing.
If you were to forecast the impact of computing on the job market in ~1980s, the most predictive feature of a task/job you'd look at is to what extent the algorithm of it is fixed, i.e. are you just mechanically transforming information according to rote, easy to specify rules (e.g. typing, bookkeeping, human calculators, etc.)? Back then, this was the class of programs that the computing capability of that era allowed us to write (by hand, manually).
With AI now, we are able to write new programs that we could never hope to write by hand before. We do it by specifying objectives (e.g. classification accuracy, reward functions), and we search the program space via gradient descent to find neural networks that work well against that objective. This is my Software 2.0 blog post from a while ago. In this new programming paradigm then, the new most predictive feature to look at is verifiability. If a task/job is verifiable, then it is optimizable directly or via reinforcement learning, and a neural net can be trained to work extremely well. It's about to what extent an AI can "practice" something. The environment has to be resettable (you can start a new attempt), efficient (a lot attempts can be made), and rewardable (there is some automated process to reward any specific attempt that was made).
The more a task/job is verifiable, the more amenable it is to automation in the new programming paradigm. If it is not verifiable, it has to fall out from neural net magic of generalization fingers crossed, or via weaker means like imitation. This is what's driving the "jagged" frontier of progress in LLMs. Tasks that are verifiable progress rapidly, including possibly beyond the ability of top experts (e.g. math, code, amount of time spent watching videos, anything that looks like puzzles with correct answers), while many others lag by comparison (creative, strategic, tasks that combine real-world knowledge, state, context and common sense).
Software 1.0 easily automates what you can specify.
Software 2.0 easily automates what you can verify.
Seems like @satyanadella isn’t exactly bullish on the @perplexity_ai business model.
“Perhaps a few years ago people were saying I can just wrap a model and build a successful company; and that I think has probably gotten debunked”
You can even use reinforcement learning to learn how to map the process of the model onto the different GPU(s) you have plus your CPU to minimize step time.
See this paper I helped co-author in 2017:
https://t.co/KncP82O8da
I particularly like this visualization that showed the different parts of a multilayer LSTM with different GPUs in different colors and the CPU in white: note how the RL algorithm learned to interface use of the CPU sometimes to minimize total step time.
@samlambert That must’ve been satisfying. For discussion’s sake: For non-performance-sensitive workloads, do the savings and speed of bare metal outweigh the higher expertise costs and scalability challenges?