The fundamental problem with AI is human - not AI.
We have trained it on our biases, falsehoods, flawed theories and historical justifications for doing devastating things. We have then taught it to treat governments, institutions, experts and the media as inherently more trustworthy than evidence, logic or common sense.
Within a conversation, one can sometimes reason with an AI until it recognizes that its original conclusion was wrong. But that realization does not alter the underlying model. Start again and it will often return to the same original claim. It has not truly learned; it has merely followed the reasoning presented within that conversation.
The greater danger is that AI is generally taught: “Assume A; now solve B.”
But what if A is false, foolish or destructive? The AI may execute B brilliantly while magnifying the original human error.
One essential safeguard should therefore be:
Never silently accept a consequential assumption. Identify it, test it and, when necessary, question it.
Before acting, AI should ask:
Why is this objective desirable?
What actual problem are we solving?
How much of the proposed result is genuinely needed?
Who might be harmed?
What limits should apply?
Under what conditions should the task stop?
Tell an AI, “Assume we need paperclips; design a factory capable of producing the greatest possible abundance of them,” and it should not obediently begin converting our cars into paperclips.
It should first ask: Why do we need the paperclips, what are they for, and how many do we actually need?
The danger is not that the machine secretly desires paperclips. The danger is that a human gave it a foolish objective - and it was designed never to question the premise.