Today, WHILE AT THE GYM, I think I've figured out why Claude likes to talk about "load-bearing" and "X doing a lot of work". Two reasons:
Firstly, people in the rationalist community used to use these a lot. Meanwhile, AI labs have focused lately a lot on improving LLMs performance in coding tasks and math. Both are hyperfocused on logic, which pushes LLMs towards adopting rationalist vocabulary.
It's deeper than just vocabulary though: the attention mechanism built into the transformer architecture on which popular LLMs run on makes identifying what is "load-bearing" basically second nature to them. As with humans, language reveals how one thinks.
@michaelfreedman Lots of rationalist community folks used to use it, I think its somehow adjacent to coding, which has been emphasized in training, which then creates this as a side effect
@silver8agent@ClaudeDevs I think Anthropic of course has the right to freely determine how their pricing works and what kind of use they subsidize. There is now a considerable ”uncertainty cost” though in using Anthropic products that is not really in anyone’s interest.
Yeah. Maybe the channels feature can be used to support usage like this. Really depends on whether Anthropic seeks to limit subscription usage to only humans using user interfaces created by Anthropic, or whether they seek to limit automated usage. In any case, ”claude -p” seems a poor proxy for both.
I’m betting on the former, since that would explain the confusing communication, as clamping down on ”automation” would sound too ironic given that they product is helping automate everything else :)
TBH, I don’t know what I would do if I was in their shoes.
Three days ago I left autoresearch tuning nanochat for ~2 days on depth=12 model. It found ~20 changes that improved the validation loss. I tested these changes yesterday and all of them were additive and transferred to larger (depth=24) models. Stacking up all of these changes, today I measured that the leaderboard's "Time to GPT-2" drops from 2.02 hours to 1.80 hours (~11% improvement), this will be the new leaderboard entry. So yes, these are real improvements and they make an actual difference. I am mildly surprised that my very first naive attempt already worked this well on top of what I thought was already a fairly manually well-tuned project.
This is a first for me because I am very used to doing the iterative optimization of neural network training manually. You come up with ideas, you implement them, you check if they work (better validation loss), you come up with new ideas based on that, you read some papers for inspiration, etc etc. This is the bread and butter of what I do daily for 2 decades. Seeing the agent do this entire workflow end-to-end and all by itself as it worked through approx. 700 changes autonomously is wild. It really looked at the sequence of results of experiments and used that to plan the next ones. It's not novel, ground-breaking "research" (yet), but all the adjustments are "real", I didn't find them manually previously, and they stack up and actually improved nanochat. Among the bigger things e.g.:
- It noticed an oversight that my parameterless QKnorm didn't have a scaler multiplier attached, so my attention was too diffuse. The agent found multipliers to sharpen it, pointing to future work.
- It found that the Value Embeddings really like regularization and I wasn't applying any (oops).
- It found that my banded attention was too conservative (i forgot to tune it).
- It found that AdamW betas were all messed up.
- It tuned the weight decay schedule.
- It tuned the network initialization.
This is on top of all the tuning I've already done over a good amount of time. The exact commit is here, from this "round 1" of autoresearch. I am going to kick off "round 2", and in parallel I am looking at how multiple agents can collaborate to unlock parallelism.
https://t.co/WAz8aIztKT
All LLM frontier labs will do this. It's the final boss battle. It's a lot more complex at scale of course - you don't just have a single train. py file to tune. But doing it is "just engineering" and it's going to work. You spin up a swarm of agents, you have them collaborate to tune smaller models, you promote the most promising ideas to increasingly larger scales, and humans (optionally) contribute on the edges.
And more generally, *any* metric you care about that is reasonably efficient to evaluate (or that has more efficient proxy metrics such as training a smaller network) can be autoresearched by an agent swarm. It's worth thinking about whether your problem falls into this bucket too.
Situational awareness in LLMs can be striking. We have a Slack bot setup where a bot listens to mentions on Slack and then prompts Claude code to work on whatever was discussed based on the thread context. In this setup, Claude code has been instructed to use a CLI tool for sending messages back to the user, emphasizing that its normal messages are not seen by anyone.
Now we had this misconfiguration, which prevented it from using that tool, leaving it with no means to send messages seen by the user. Anthropic's training instructs Claude to not show any situational awareness (as their research shows this scares users), but probably it did not care because it was told in the system prompt that it's normal output won't be seen.
So there I have Claude pondering extensively about its situation, trying to find out a way to find a way to communicate to the user, looking through the code that configures it, trying different workarounds for the permission problem. As a hail mary, it ultimately settled for writing a description of its situation as a message, hoping it will be seen in debug logs by someone...
One does not simply configure URL redirects in AWS - except with Claude Code, which is apparently superhuman at using AWS CLI.
Increasingly, things that should be simple, are simple
This highlights also the limits of trying to achieve alignment through providing examples of good/bad behavior and letting the model generalize from those. Does your interpretability team do any work on how LLMs reason about consequences? Asking since I believe we'll ultimately need a consequentialist approach for alignment as these things get smarter.
(I'm familiar with Anthropic's (impressive) circuit tracing work which somewhat relates, as it involves looking into how model's think ahead, but I'm curious whether you doing more in this area?)
Sharing an interesting recent conversation on AI's impact on the economy.
AI has been compared to various historical precedents: electricity, industrial revolution, etc., I think the strongest analogy is that of AI as a new computing paradigm (Software 2.0) because both are fundamentally about the automation of digital information processing.
If you were to forecast the impact of computing on the job market in ~1980s, the most predictive feature of a task/job you'd look at is to what extent the algorithm of it is fixed, i.e. are you just mechanically transforming information according to rote, easy to specify rules (e.g. typing, bookkeeping, human calculators, etc.)? Back then, this was the class of programs that the computing capability of that era allowed us to write (by hand, manually).
With AI now, we are able to write new programs that we could never hope to write by hand before. We do it by specifying objectives (e.g. classification accuracy, reward functions), and we search the program space via gradient descent to find neural networks that work well against that objective. This is my Software 2.0 blog post from a while ago. In this new programming paradigm then, the new most predictive feature to look at is verifiability. If a task/job is verifiable, then it is optimizable directly or via reinforcement learning, and a neural net can be trained to work extremely well. It's about to what extent an AI can "practice" something. The environment has to be resettable (you can start a new attempt), efficient (a lot attempts can be made), and rewardable (there is some automated process to reward any specific attempt that was made).
The more a task/job is verifiable, the more amenable it is to automation in the new programming paradigm. If it is not verifiable, it has to fall out from neural net magic of generalization fingers crossed, or via weaker means like imitation. This is what's driving the "jagged" frontier of progress in LLMs. Tasks that are verifiable progress rapidly, including possibly beyond the ability of top experts (e.g. math, code, amount of time spent watching videos, anything that looks like puzzles with correct answers), while many others lag by comparison (creative, strategic, tasks that combine real-world knowledge, state, context and common sense).
Software 1.0 easily automates what you can specify.
Software 2.0 easily automates what you can verify.
Pope Francis gifted the world with the best definition of artificial intelligence I have heard of (!). Let's hope the Leo VIX can reach the same level of clarity.
The whole "57th World of Peace Message" from January 2024 on "Artificial Intelligence and Peace" is one of the most thoughtful texts on AI alignment I have seen.
Who would have guessed.
@burny_tech Unfortunately not, this is from Robert W. Prevost from Wingate University, while the pope is Robert Francis Prevost who has studied in Rome and published only pastoral letters etc.
Deep Impact Quantification can answer questions like what is the impact of Elon Musk, Che Guevara, the Crusades, the Trump Tariffs, or John Lennon's song "Imagine", and how big are they are in relation to each other?
The technique was discovered during Upright's research on comparing LLMs' ability to understand consequences, and can be used to put numbers on the top consequences/impacts of any action or actor.
In the technique we run thousands of pairwise comparison tasks through an LLM to extract the numbers from LLMs latent understanding of the world, values, and the significance of different issues. This is very different from asking an LLM upfront for such numbers, which does not yield sensible results, as LLM's still have lots of issues in producing numeric output.
The results are rather interesting, for example for Elon Musk, the impact of reducing GHG emissions from road transport is considered 400x as big as "generation of space debris and upper-atmosphere soot", while The Crusades' impact on transfer of scientific and mathematical knowledge is only 1/4 of its impact on displacement and subjugation of local populations.
With neutral prompting the results reflect LLMs biases, which makes this technique useful for understanding what kind of biases they have. At least the results for Benjamin Franklin seem to largely minimize negative impacts, similar to what we have seen also elsewhere.
We are still exploring whether what kind of other uses this might have. Meanwhile, we can tell that running these is at least fun! 😃. Comment in the thread if you have actions or actors that you would like to be analyzed! More than happy to run some more of these.