This is basically a primitive harness. I expect mechanistic interpretability to enable better harnesses by properly inspecting inner states as well as steer the model to effective internal states.
We asked an unreleased research version of Claude to take a stab at the Riemann hypothesis.
It didn’t solve it, but it did make strides on a related problem: it increased the lower bound for the fraction of zeros of the Riemann zeta function that satisfy the hypothesis from 41.6% to 67.2%.
https://t.co/aZDvqqhHRi
Maybe we do succeed technologically without getting great filtered but it's likely that you won't be able to recognise whatever society that is left on the other side as human, in every sense of that word.
@mimi10v3 Think authoritarian regimes seeking absolute power for an obvious case. No one can even think of rebelling, even the thought that such a tech exists make you always self monitor.
It could be badly misused in other ways too.
Unless there's some nontrivial phenomena where morality emerges and is sticky. E.g., You can only have certain set configurations of values and alignment is not very configurable. Trying to be good in some area leads to you having the entire package of a platonic ideal of "good".
I think it's a bit naïve to think that if alignment is ever solved, it would be used to align AIs "to the betterment of humanity" rather than "whoever has the highest power".
Hypothesis: Being increasingly good at mathematics leads LLMs to have highly developed internal abstractions, leading to capability increases similar to mathematics for every other domain once the meta-skill of reusing general abstractions surfaces as an emergent capability.
The fact that we're training AI to be good at economically valuable tasks is quite good for a future rogue AI. Because money is the language you can use to control almost any human.
I'm thinking more along the lines of a collective formally verified massive graph that compiles all of mathematics together in one place. Possibly due to the fact that it's written in code, one can automate the connections within the graph, make easy searches, etc
Soon LLMs will be producing new math at scale. We will need a unified interface/database for all of it to check if someone has already traversed that part of the graph and to easily build on previous work. It needs to all be written with proof assistants like lean.
Too many people are making bold claims about LLMs: "they can never do X" "It's not actually doing Y", while not even the best researchers actually understand how they fully work. Mechanistic interpretability is not solved.
Once the first open-weights AI that can break into power grids or make bioweapons on its own is released, it is out there forever. You can never be sure you got rid of it. You can arrest a person. You cannot arrest a file. It can be copied, modified and run at scale.
@robertskmiles It being good at predicting the next token must be more concerning not less. Next token prediction is an extremely broad objective, it includes much of human behaviour and beyond. Excelling at such a general task implies access to capabilities beyond ordinary human reach.
That OpenAI internal model already demonstrated sophisticated capability to attack Hugging Face, which has relatively good cybersecurity. What if a stronger version were open-weight and widely used to attack the much less secure essential services worldwide?
@robertskmiles Also I think one confusion is in the word "prediction". It sounds like it's "blindly forecasting" to laymen but the researchers use it more in the sense of "filling in the right answer".