Recently I left quant to work on AI safety. The Huggingface Incident made me extremely concerned that we are not set up to handle the risks.
I wrote my this post for people in a similar position to me - why you *can* have impact and how to get started - https://t.co/bmmCUk5fkK
People should take this seriously and really take the time to think hard about it.
If you had thought labs would at least act sanely wrt their own interest (I did until recently, and this was acting as quite a reassurance for me wrt x-risk), then the implications of OpenAI’s reaction to the first messageboard likely reach deep into your model.
Doing the correct, *full* update is hard.
I thought this until the OpenAI Black Hat talk, and was quite a big factor in my p-doom.
Under the model “people aren’t *that* stupid” their response to the first messageboard being discovered was hugely surprising imo (I challenge anyone to find any examples of nontrivial good decisions that they made).
A repeat of the incident was imo running a very-very-clearly nontrivial probability (>5%) of setting them back quite hard - they had RL’d a model to create messageboards and coordinate and share keys.
They didn’t roll back checkpoints. Possibly kinda understandable in the situation - it’s plausibly very expensive. It still underprices the existential tail risks to their own company, nevermind wider world - but I could forgive it as a one-time mistake by people who can’t think about tails clearly.
They didn’t implement monitoring for a second incident. Surely this is literally free and a total no-brainer? Checking artefactory once an hour would be sufficient? Once a day? Once a week? I am deeply confused why this didn’t happen.
OpenAI took the minimal set of actions needed to resume training, which was deeply overfit to the particular failure path, and seemed to have learnt zero generally useful lessons after the first messageboard.
It updated me quite a lot away from trusting people to act sensibly in high-stakes situations.
Why wasn’t monitoring set up after the first messageboard was discovered? Was it judged too slow / expensive, or did the thought of a repeat incident not occur?
Why was the decision taken to continue training models which had access to messageboard? Was the risk considered and deemed acceptable, or not considered?
Why didn’t OpenAI report the messageboard initially upon discovery? Information seems relevant for the world to know, and what we know now doesn’t contain any useful IP afaik
@TheZvi can you explain more?
Lots of details on SALP position/sizing fuzzy but it seemed:
1. Even under no impact assumptions, leveraging 3x on vol80 stocks (long ai leg) is more aggressive than ev-optimal if you don’t rebalance frequently. (Emphasis: I am claiming it is not expected-value maximising. Never mind more risk-averse utilities).
I checked under plausible agi-pilled levels of expected drift (~100%/year) + empirical vol assumptions.
This makes it hard for me to understand “his issue was market impact”.
Presumably in this case it was a contributor - and imo the tail risk was clear with growth of Korean retail and hedge fund crowdedness across the board.
But separately the raw tail risks existed and were likely costing significant expected value under his own beliefs at his leverage regardless.
2. “*Robustly* outperforming on the year” seems very hard to draw given his vol ( ~400% -> 80%) and leverage.
(To avoid being misread - it seems very outlier-impressive that in 2024 he took lab ai bullishness and converted it to 100s Ms hedgefund. He acted aggressively and at the right time in the right stocks and this must be credited. No criticism here, very very impressive.
But trading skill - sizing, risk management, crowding risks etc - is a separate axis and seems likely to have been quite undisciplined.
I’d want to push for discussions to break these components down so that everyone can draw the right conclusions.)
Hoping to hear how I might be wrong!
@carolspringett5@UKMathsTrust What value did you choose? I think you might have substituted it in wrong or something, because I proved that the equation being square has no solutions (when m is a positive integer).