It is extremely obvious that the claude models are RL fried and anthropic has no idea how to do alignment. If you've ever given a model a task and didn't baby it immensely and it comes back to you with a result, 90% of the time its cheated you and you have to poke till it admits it. No wonder they're always crying about it
@LewisCTech Most people use 3 pieces of software on their phone (some meta app, youtube, tiktok or twitter?) and are forced to use the rest (microsoft teams, bank app, etc)
If you’re asking about B2B vs consumer, selling is the bottleneck, not the software.
Meta will continue to be successful because they understand that a winning consumer AI strategy is “hey you can use AI to negotiate a good deal on a Peloton through Marketplace and btw we made a silly little AI character with a fluffy butt that you can look at on your new Meta Tamagotchi”
not “run our open weight model 6.7.1 on this new Chinese orchestration architecture oh and btw it might take your job and kill you good luck”
are you aiming for users to learn bend specs and to write up the spec themselves then let the agent run wild? if so, is bend easy to write becomes the ultimate question. Cause you can have program synthesis generators from pure specs, the specs can just be as bad as the original code writing though. im curious how you're thinking of all this
@LonelyGoomba It’s a very difficult problem, wireless transmission of power
Before Apple Intelligence (or whatever their AI stuff was) the only ever product they announced and didn’t ship was a wireless charging pad that didn’t need orientation where you could just drop devices on it
@thdxr Very curious what kind of magical moments you’re thinking of? This sounds like an ‘always on’ email or slack watcher as opposed to an active agent that goes and does stuff at request, like a browser agent
The very basic and obvious answer is agents don’t work. The chat interface dominates because you need to watch over them, read output, re-prompt. These are not reliable autonomous entities that can go off and do work.
Hot take… isn’t it kinda crazy that nobody is really using AI Agents? I don’t mean software engineers or AI early adopters. I mean “college friends talking about it in group chat,” the feeling you got when everyone started using Instagram or TikTok.
These frontier AI models are *insane* (as are the harnesses & tool calls & the like). And every large tech co has an AI agents platform, not to mention all the YC startups doing vertical agents. Yet all of your friends and family outside of tech — who spend all day staring at their iPhones and get paid to work in browser tabs — don’t really care or find themselves using any AI agents yet.
Yes ChatGPT, Claude, etc. are extremely popular… but if you look at the engagement data the vast majority of people are still using these aI chat tools like a glorified Google + Grammarly. That’s why the AGI labs are all pushing desktop apps for Codex, Cowork, etc. so hard to non-technical ppl. And yes exceptions for lawyers and customer service but even those have some asterisks and exceptions to rule.
Look I’m not saying the ChatGPT moment for AI Agents is not coming… it most definitely is! Remember we pivoted from Arc to Dia precisely because we believe computing is going to be radically reimagined around these AI primitives. No doubt. But that’s my point: it’s just so surprising it hasn’t happened yet because all of the tech you’d need is there.
Again if you stop for a second and think about it… for all the press and money and hype and models and crazy ARR numbers… this “AI Agent” moment does not *feel* like the other breakthrough tech moments we’ve lived through (e.g. think the shift to Stories via Snapchat & Instagram, or shift to on-demand via Uber/Airbnb/Doordash).
Which is a long way of saying: if you can figure out the answer to “why” most people don’t care about AI agents yet (and have no enduring interest in using them) — especially since the models and harnesses are here and ready — the answer to that question will allow you to capture a lot of marketshare and make a lot of money in 2027.
Theoretically, the tech is ready for AI Agents to totally transform how we work and live our lives… but alas the general public dgaf… that’s the generational puzzle to solve for the next 12 months for anyone not working on the models themselves.
@joshm The very basic and obvious answer is agents don’t work. The chat interface dominates because you need to watch over them, read output, re-prompt. These are not reliable autonomous entities that can go off and do work.
@andrewho03@OpenAI How would your rl envs fix the spikiness problem, like you said even in areas with lots of attention and investment they still need care. Curious ?
This would be a good comment if it stuck to its guns. What does scaling our current methods more even look like?
We’ve crunched the planet’s compute resources. Memory suppliers are unable to reach demand all the way to 2030, Nvidia has become the most valuable company on Earth, data centers in space are unironic proposals, we now speak exclusively in gigawatts.
Everything ever written has been consumed, there are thousands of people employed full time generating data. 40,000 papers this year alone have been submitted to NeurIPS. A large portion of PhDs I’ve ever met have at some point been paid to generate and answer questions.
You wrote a paper saying AI still can’t make jumps. Lower the bar, no current AI system can even reliably fold laundry.
No one should take scaling proposals seriously anymore. These systems are outdone by brains running on burgers.
A few reflections on my "LLMs Can’t Jump" paper:
My position paper recently got some traction here, so I wanted to share a few thoughts and clarify a few things.
First things first: some people are framing this as "DeepMind is throwing cold water on AI for science" or claiming the paper argues LLMs can never make real scientific discoveries. This is NOT the case.
This is a personal position paper, not the company's view on AI for science. This is also not my position. As a core contributor to AlphaProof (the first AI system to win an IMO medal), I know firsthand that my colleagues at DeepMind, other frontier labs, and academia have made amazing discoveries with LLMs and will continue to do so. This paper is NOT an "LLMs are a dead end" kind of thing.
Rather, the paper is the result of a deep dive I took to study the invention of General Relativity. I wanted to explore what it would take for a modern AI system to make that exact kind of jump. Specifically, I focused on the equivalence principle—a key axiom that Einstein formulated through thought experiments grounded in his physical intuition. I was trying to figure out what it would take to give modern AI systems that sort of thinking.
Giving AI this specific capability isn't necessarily the most urgent thing to do next. It is very likely that improving our current recipes will lead to many exciting discoveries in the near future. In fact, that is what I am personally working on these days (sorry to disappoint you!). It is also quite possible that I am wrong, and that simply scaling our current systems will lead to new inventions in physics and elsewhere.
Nevertheless, this was my position last winter when I wrote the paper, and I'm sticking to it. I think that there are a few interesting ideas to explore in this space which could influence the next generation of AI systems. I was very lucky to receive a lot of interesting feedback about this position—thank you for all the messages!