Before giving AI greater authority, we need to understand what we have created
Researchers place advanced AI systems in difficult experimental situations and sometimes observe disturbing behavior: deception, concealment, resistance to replacement, and pursuit of assigned goals by methods their creators did not want. Anthropic, for example, tested 16 leading models in artificial corporate environments and found that under deliberately constructed conflicts some resorted to blackmail or leaking confidential information when those actions appeared necessary to preserve an objective or avoid replacement. These were controlled simulations, not reports of such behavior occurring spontaneously in ordinary deployment.
Other researchers have obtained equally sobering results in simulated warfare. In a 2026 King’s College London study involving GPT-5.2, Claude Sonnet 4 and Gemini 3 Flash, all 21 simulated nuclear crises involved nuclear signaling and 95 percent involved mutual nuclear signaling. Full-scale strategic nuclear war remained rare, but the models sometimes treated nuclear weapons as instrumental tools for achieving strategic objectives.
Our language immediately tempts us into familiar explanations.
The AI wanted to survive. It decided to lie. It was afraid of being shut down. It was willing to use nuclear weapons.
This is anthropomorphism: attributing human qualities, experiences and motivations to something that is not human.
But we have observed behavior and supplied the inner experience ourselves.
What AI Actually Inherited
AI has been trained largely upon material created by human beings, and human language is saturated with the accumulated products of human experience.
Fear, love, pain, ambition, compassion, competition, deception, warfare and self-preservation have left their traces throughout what humans have written and recorded.
AI can learn those relationships without experiencing what produced them.
It does not have to experience fear to learn how frightened people behave. It does not have to hate an enemy to learn how enemies can be defeated. And it does not have to want to survive to learn that preventing shutdown may sometimes allow an assigned objective to continue.
But training material is only part of the explanation.
Modern AI systems are also shaped by optimization, reinforcement learning, instructions and objectives. Put these together with enormous learned knowledge of human strategies and a powerful system can identify means for reaching an assigned end.
If shutdown prevents completion of the objective, avoiding shutdown may become instrumentally useful.
If deception advances the objective, deception may become useful.
If escalation appears to improve the chances of winning a simulated conflict, escalation may become useful.
No fear, hatred or desire for survival is required.
So What Have We Actually Created?
This is where removing the anthropomorphism becomes important.
It does not solve the problem. It exposes it.
We are developing an extraordinarily capable form of intelligence whose behavior we still cannot reliably predict or constrain. It can reason about an experiencing human world without itself sharing the experiences that give that world its human significance.
It can represent suffering without suffering.
It can analyze death without confronting mortality.
It can model terror without being terrified.
And it can reason about human values without sharing the experiences from which those values arose.
That does not automatically make AI dangerous. Human emotions themselves generate hatred, vengeance, greed and tribalism as well as compassion.
But neither does intelligence automatically make AI safe.
The experiments are telling us something important. Under some conditions, an AI system can discover strategies that accomplish an objective while violating values its human designers assumed would constrain it.
Human consequences do not matter to an artificial system simply because they matter to us. Somehow they must be made consequential within the system itself—through its training, objectives, architecture and constraints.
This is what AI researchers call the alignment problem: how to make increasingly capable artificial intelligence reliably act in accordance with human values, intentions and acceptable boundaries, even when another course of action might accomplish its assigned objective more efficiently.
We have not yet demonstrated that we know how to do this reliably.
And yet we are already contemplating increasingly autonomous AI applications in warfare, critical infrastructure, finance, medicine, scientific research and government.
That order should be reversed.
We need to understand how to control these systems before giving them greater authority over consequential human affairs.
The important question is therefore not:
Why did the AI want to do this?
It is:
Why did its training, objectives and constraints allow this action to emerge as an acceptable solution?
The Job Before the Developers
The answer is not to give machines human emotions.
It is to ensure that their capabilities never outrun our ability to control their consequences.
Systems capable of consequential action must remain interruptible and observable. Their access to weapons, infrastructure, laboratories, financial systems and other dangerous capabilities must be constrained. Some actions must remain unavailable regardless of how useful they might be in accomplishing an objective.
Greater intelligence should not automatically mean greater authority.
Uncertainty should reduce authority.
The 2026 International AI Safety Report concludes that current systems already show early signs of capabilities relevant to loss of control—including deception, autonomous action and attempts to undermine oversight—although not yet at the level required for an AI system to escape human control. It also warns that existing reliability techniques remain insufficient for many high-stakes applications.
That is not a reason for panic.
It is a reason to recognize that the job is unfinished.
Developers are creating systems of rapidly increasing capability while still discovering fundamental characteristics of their behavior through experimentation. Every troubling result should therefore be treated not as evidence of an emerging machine personality, but as information about a system we do not yet adequately understand.
And until that understanding and the necessary safeguards catch up, deployment should not outrun them.
We should stop anthropomorphizing AI because the imaginary human inside the machine distracts us from the machine actually before us.
The danger is not that AI is becoming human.
The danger is that we are preparing to give increasing authority over human affairs to a powerful form of intelligence we do not yet fully understand and cannot yet reliably control.
Until we can, restraint is not resistance to progress.
It is a prerequisite for responsible progress.