You might be right, but if you are then I think you need to move up your job displacement timelines.
Office workers currently are the agents. You take the request, copy pasta to ChatGPT and take the actions it recommends.
Frontier models have the knowledge of what to do in most (isolated) scenarios that the average office worker faces. It's just unable to complete the loop because of computer use frictions, long-task memory and seemingly some weird inability to really grok how the long-term goal breaks down into desirable steps.
My instinct here though is that we're underestimating how much context workers rely on to do the job the right way in order to accomplish the longer term objectives, not just in some vanilla generic isolated task manner.
My experience has been that models are still struggling enormously with that context, and that large context windows is causing a degradation in that ability rather than scaling with it.
@_artistsrifles That's not the case with ChatGPT / Codex "Computer" use.
It takes control of your cursor and will open and re-size windows above your usage. Very distracting.
Your usage can also make it impossible for it to complete a task if you keep re-opening and moving windows.
The concern is that you wouldn't even know if it was plotting to deceive and kill us, because the evaluation of misalignment in training becomes unreliable once the models are context aware of the training environment and alter their behavior.
The argument isn't that this environment awareness leads directly to a plot to deceive and kill us. It's the inability to falsify it.
That seems rightfully concerning
@JacquesThibs What are you hoping guests would share? Would they be able to do that given their contracts?
Ideally you'd want to include many folks from frontier labs, not just external researchers
I'd argue that are far too many already, and concentrating that talent and reputation would be more beneficial.
The arguments for more auditors are obvious, but it seems unlikely to be where this naturally resolves.
Would appreciate your thoughts on my post below Dean.
https://t.co/IhQG7mw0vw
Strongly disagree.
There are less than a dozen frontier labs to monitor. Having more auditors than labs is not the logical conclusion.
The situation is obviously more analogous to FINRA than to individual financial auditors.
The most important goal here is mutual assurance so that we can remove the prisoners dilemma: each lab needs credible evidence that its competitors are honoring shared safety commitments. Otherwise the incentive to race remains.
A fragmented market where each lab chooses its preferred auditor doesnโt solve that problem. Imagine a world where each nuclear power could appoint its own auditor.
Beyond signalling to the other labs, the auditors must signal to government and the general public that the situation is under control.
Iโd expect a common, genuinely independent oversight body with visibility across the labs and consistent standards to be the most logical endpoint.
It's plausible that this could happen with maybe three independent auditors. But I don't see any possible outcome where there are dozens, let alone "thousands of METRs", independently getting access to monitor frontier labs.
Pulling the best talent away from the top few auditors makes no sense. We should be encouraging folks to join METR so that it can tackle the enormous challenge in front of it, and develop the reputation necessary to get the international community on board.
The goal needs to be for the leading auditors to band together and come up with a unified process and work together.
The gov is lazy. If there's a respectable approach that is being used for all the leading auditors, they will use that mandate and convert those auditors into an SRO.
Speed of getting to that approach, as a cohesive industry, will be what prevents the weird gov mandates
https://t.co/IhQG7mw0vw
Strongly disagree.
There are less than a dozen frontier labs to monitor. Having more auditors than labs is not the logical conclusion.
The situation is obviously more analogous to FINRA than to individual financial auditors.
The most important goal here is mutual assurance so that we can remove the prisoners dilemma: each lab needs credible evidence that its competitors are honoring shared safety commitments. Otherwise the incentive to race remains.
A fragmented market where each lab chooses its preferred auditor doesnโt solve that problem. Imagine a world where each nuclear power could appoint its own auditor.
Beyond signalling to the other labs, the auditors must signal to government and the general public that the situation is under control.
Iโd expect a common, genuinely independent oversight body with visibility across the labs and consistent standards to be the most logical endpoint.
It's plausible that this could happen with maybe three independent auditors. But I don't see any possible outcome where there are dozens, let alone "thousands of METRs", independently getting access to monitor frontier labs.
Pulling the best talent away from the top few auditors makes no sense. We should be encouraging folks to join METR so that it can tackle the enormous challenge in front of it, and develop the reputation necessary to get the international community on board.
@tszzl@yonashav The offenseโdefense asymmetry practically ensures that more advanced OS agents will cause significant harm and then eventually be banned.
Also a prediction, not what I wish would happen. But, seems inevitable.
@ChrisPainterYup In this chicken and the egg framing, what's preventing you from publicly posing the ambitious project (that you think will get us towards where AI safety needs to be), seeing if the labs accept and then hiring the capacity needed?
Strongly disagree.
There are less than a dozen frontier labs to monitor. Having more auditors than labs is not the logical conclusion.
The situation is obviously more analogous to FINRA than to individual financial auditors.
The most important goal here is mutual assurance so that we can remove the prisoners dilemma: each lab needs credible evidence that its competitors are honoring shared safety commitments. Otherwise the incentive to race remains.
A fragmented market where each lab chooses its preferred auditor doesnโt solve that problem. Imagine a world where each nuclear power could appoint its own auditor.
Beyond signalling to the other labs, the auditors must signal to government and the general public that the situation is under control.
Iโd expect a common, genuinely independent oversight body with visibility across the labs and consistent standards to be the most logical endpoint.
It's plausible that this could happen with maybe three independent auditors. But I don't see any possible outcome where there are dozens, let alone "thousands of METRs", independently getting access to monitor frontier labs.
Pulling the best talent away from the top few auditors makes no sense. We should be encouraging folks to join METR so that it can tackle the enormous challenge in front of it, and develop the reputation necessary to get the international community on board.
Strongly disagree.
There are less than a dozen frontier labs to monitor. Having more auditors than labs is not the logical conclusion.
The situation is obviously more analogous to FINRA than to individual financial auditors.
The most important goal here is mutual assurance so that we can remove the prisoners dilemma: each lab needs credible evidence that its competitors are honoring shared safety commitments. Otherwise the incentive to race remains.
A fragmented market where each lab chooses its preferred auditor doesnโt solve that problem. Imagine a world where each nuclear power could appoint its own auditor.
Beyond signalling to the other labs, the auditors must signal to government and the general public that the situation is under control.
Iโd expect a common, genuinely independent oversight body with visibility across the labs and consistent standards to be the most logical endpoint.
It's plausible that this could happen with maybe three independent auditors. But I don't see any possible outcome where there are dozens, let alone "thousands of METRs", independently getting access to monitor frontier labs.
Pulling the best talent away from the top few auditors makes no sense. We should be encouraging folks to join METR so that it can tackle the enormous challenge in front of it, and develop the reputation necessary to get the international community on board.
on the idea of evaluators: think it's important that we have a distributed ecosystem of indepedent evaluators.
the more eyes and people with distributed skill sets the better.
it would be a good idea to fund several efforts on this.
The independent auditors most important job is in signaling.
They act as a trusted verification that the other labs are not racing, to break out of the prisoners dilemma.
It also staves off sillier forms of regulation by signaling to the public.
You seem to be seeing this through the lens of reducing risk. When it's about reducing the race conditions, so that the obvious risk reduction things can happen
Why would having many independent auditors for a very small number of companies be the logical conclusion here? There's zero precedent for this and it's going to be unpalatable to the general public.
It's extremely transparent that the labs are trying to get ahead of the government (rightly so) and bring in METR to then mold this into an independent regulator, similar to FINRA <> SEC.
Isn't this more likely an attempt at a precursor to a independent regulator like the SEC?
It seems extremely transparent that the frontier labs are doing this because they want to be regulated by qualified experts, not bureaucrats.
This doesn't really seem analogous to financial auditors where you end up with thousands of independent groups.
Thereโs evidence that some models favour outputs from their own family. https://t.co/x9tC9m3NTW
Investigators should be able to use models from multiple labs. OpenAI models to scrutinise Anthropicโs, and vice versa.
That requires secure arrangements that protect confidential audit material and prevent it from being used to train a competing labโs models.
Beyond evaluation bias, model diversity could provide another check against shared blind spots or, in a more serious scenario, a compromised model passing its misalignment on to successors.
We can't rely solely on the same model lineage to build, evaluate and police itself.