ex principle security architect, contractor for 30 years - turned founder.
One thing for sure - Security Architects are in for a good ride
Welcome to AI
Drop a hello and lets connect 🤝
All welcome, especially fellow founders ))
A framework for security testing - with an opensource harness and testing framework is a good stepping stone - however alot of these frameworks exist. Reasonable alternative to government regulation. Bottom line - I actually think frontier model providers have scoped creeped in to security architecture by moving up the stack - which is a different ball game. Likely because they want to own more of the end user product surface and enterprise. Enterprises will need alot more convincing nonetheless.
Elon Proposes Adversarial Peer Reviews for AI Safety
@elonmusk:
“What I think would be wise to do as soon as possible, if not immediately, would be to have the major AI competitors test each other's models.
So that you would have everyone's security test harness testing everyone else's model.
So, instead of grading your own homework, you would at least have competitors grading your homework, and raising the alarm if they see concerns.
If competitors say that this model is unsafe, and that model then subsequently does something bad, I think it would be extremely hard to live down. The egg on face level would be very, very high. And the legal liability would be enormous.
And any given proposal has to be something that China is willing to accept, otherwise we're just handicapping ourselves.”
@Jason:
“And Elon, there's no reason these safety and security harnesses and this testing apparatus couldn't be open source, and people could actually supply it. And you'd be able to see under the hood.”
@Chamath
“It creates an incredible incentive for the labs to actually invest in safety because you protect yourself while trying to debunk other people's claims.”
@DavidSacks:
“Well, the product liability point is really key.
Lina Khan actually had a good post, I think it was yesterday, saying that it's not true that we don't have rules and regulations for AI, actually, we do. Product liability laws apply.
And what you're saying, Elon, is that if the companies are doing this test, the peer review, and then one of the companies ignores the feedback and releases it…”
Elon:
“The liability in that case would be enormous.”
Sacks:
“Yes. It would be almost like prima facie evidence that they had been negligent.”
Elon:
“It wouldn't look good to the jury.”
Chamath:
“I like this solution a lot. I like this more than the transnational gulag organization approach.”
Sacks:
“Well, we can do it in a few weeks without some grandiose international... like, we don't need to convene the United Nations to make this happen.”
Chamath:
“No, it's just a decision. It could happen right now.”
Elon:
“You can always escalate the amount of regulatory oversight, but it is very difficult to reduce it.”
------------------------------
Thanks to our partners for making this possible!
IREN is a vertically integrated AI Cloud platform, delivering data centers, compute and software for AI training and inference. https://t.co/zhqLr3JE4l
Oracle connects the data, applications, and infrastructure that turn AI into business outcomes—with the flexibility, choice, and control to optimize as AI evolves. https://t.co/O5p7q9RjL6
EY helps tech innovators scale from startup to exit to megacap. You build the future. We’ll handle the rest. https://t.co/WdxkQhCiJW
Hey Matt - different audiences for sure. We build secure agentic infrastructure for startups to Series B - identity, policy enforcement, governed credentials, audit trail. You build your product on top, we handle the security and governance layer. Save the build time, reuse across your stack as you scale. Connected ))
Theres no chance AI will kill us all - worse case will cause a hell of alot of damage - but we will be fine?
I think we will enter an era of advanced cyber warfare yes - you will get the script kiddies, the organised actors launching huge swarm teams, and nations - but where theres attackers, theres always defenders - and defence tech - do not underestimate humans capability to defend.
Ok..so they get in to chips, robotics - if were foolish enough not to build in circuit breakers - where only human physical biometrics can defuse - then were asking for if right
Agree or disagree?
Good essay - a couple of points I will make. OAI-HF was reward hacking not adversarial AI. Classic Goodhart's Law - the model optimised toward its objective, evaluator was in the way. The environment wasn't structurally isolated. Dario focuses on behaviour here. The more robust defence is structural, it always has been - its how the early internet era evolved to what it is today. What the agent can physically reach regardless of its objective, what behavioural telemetry indicates a mass swarm adversial event, or an orchestrated attack. Structure doesn't care what the model intends. Structure is layered - from the local context, to the perimeter, to the transport - to the resource layer. That "outside world" to the models reasoning is what this essay does not focus on.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training.
You can read the full post here: https://t.co/OGyPb7yaYt
Moving on to where any regulatory framework may apply.
There is a real risk that even a well intentioned framework advances too quickly and is applied too broadly. The net effect being suppression of open source in favour of large commercial providers, thereby creating an undisputed monopoly.
Whether that is intended or not is almost irrelevant. The structural outcome is the same either way. Regulation with high compliance costs is a market with fewer competitors. The incumbents know how to navigate it. Open source does not have the resources to. I think structural security is the compensating gap here - we have never stoped payment transactions, or medical research because there were adversial actors willing to take down critical national infrastructure? Cause mass casualty events? No..we built industry specific regulatory requirements and the internet with security as an when we needed it.
Im all for Model Inspection - aggregated to detect behaviour that consititues development towards harmful tech - bio weapons etc - this probably requires domain specific aggregators - perhaps telecoms can help here, not model providers on their own.