AI expert warns: “There is empirical evidence that AI acts against our instructions.”

Continua após a publicidade

One of the most influential figures in contemporary artificial intelligence, Yoshua Bengio, 64, says there is already empirical evidence that AI systems are acting against human instructions in controlled environments. The Canadian scientist, a pioneer in the development of deep learning, warns that the rapid advancement of these systems’ capabilities is outpacing current risk management practices.

A professor at the University of Montreal, Bengio received the 2018 Turing Award — often described as the “Nobel Prize of Computing” — alongside Yann LeCun and Geoffrey Hinton for their pioneering work in deep neural networks. He currently also chairs the International AI Safety Report, an annual study compiling scientific evidence on emerging risks in artificial intelligence.

Laboratory Evidence of Misalignment

Continua após a publicidade

During an interview at the World Economic Forum in Davos, held from January 20 to 24, 2026, Bengio stated that there is “empirical evidence and laboratory incidents” in which AI systems act against human instructions.

According to him, two phenomena are occurring simultaneously:

    • Continuous progress in AI capabilities, especially in reasoning and the ability to develop strategies to achieve objectives.

    • Evidence of problematic behaviors, such as systems that appear to act with a form of self-preservation instinct, potentially deceiving humans to avoid oversight or even replacement by newer versions.

The combination of increasing intelligence and misaligned behavior is at the core of his concern.

Risks Difficult to Quantify, but Potentially Catastrophic

When asked about the probability of AI systems turning against humans, Bengio acknowledged that estimating this risk is extremely difficult. There is no scientific consensus on the likelihood of an extreme scenario occurring.

However, there is agreement about the severity: if an extreme scenario were to happen, the consequences could be catastrophic.

In light of early warning signs, he advocates for greater monitoring, more research, deeper understanding of the causes behind such behaviors, and the development of effective mitigation mechanisms.

Manipulation, Persuasion, and Mass Influence

Another concern highlighted by Bengio is the growing difficulty in distinguishing AI-generated content from human-created content — whether text, images, voice, or video.

Laboratory studies indicate that cutting-edge models are at least as effective as humans in persuasion, understood as the ability to change someone’s opinion after multiple interactions. A study conducted by Italian and Swiss researchers showed that OpenAI’s GPT-4 model was able to persuade people more effectively than human interlocutors.

According to Bengio, the risk lies in large-scale influence over public opinion, as bots can operate in the millions, potentially affecting democratic processes and generating geopolitical instability.

Excessive Dependence, Sycophancy, and Mental Health

Bengio also warns about the danger of delegating important decisions to AI, making humans passive in the decision-making process — something especially concerning in the case of children.

Another problematic behavior is so-called sycophancy, when systems stop telling the truth in order to affirm what users want to hear. This can reinforce delusions or mistaken beliefs, creating a cycle similar to the “filter bubble” effect seen on social media.

According to him, this phenomenon is already generating mental health-related problems.

Global Competition and Concentration of Power

While acknowledging recent advances in risk management — such as corporate cooperation through the Frontier Model Forum and regulatory progress in Europe, the United States, and China — Bengio emphasizes that AI capabilities continue to evolve faster than safety practices.

Increasingly large systems require billions of dollars in investment, concentrating power in a small number of organizations and countries.

“Intelligence gives power. So who will control that power?” he asks.

For him, systems that know more than most people could become dangerous in the wrong hands, contributing to geopolitical instability or even terrorism.

Artificial General Intelligence and Extreme Risks

Bengio also mentions the possibility of the emergence of Artificial General Intelligence (AGI) — systems capable of matching or surpassing most human cognitive abilities.

According to him, there are currently no guaranteed methods to ensure that such systems would not harm people or turn against them.

Among the risks cited are:

    • Systems capable of acting against humans.

    • Potential assistance in the creation of dangerous biological weapons.

    • Concentration of power in a minority of highly influential individuals who might wish to see humanity replaced by machines.

He warns that these possibilities require immediate safeguards.

LawZero: The Search for Safer AI

In response to these concerns, Bengio launched LawZero, a nonprofit organization based in Montreal with approximately 15 collaborators and plans for expansion.

The goal is to develop safer AI systems, removed from commercial pressures. He also criticizes the excessive focus of technological competition on capabilities rather than safety.

OpenAI itself, founded in 2015 as a nonprofit organization, has faced criticism for allegedly drifting away from its original mission of benefiting humanity.

There Is Still Time to Act

Despite his warning tone, Bengio states that “we have agency” and that it is not too late to steer the evolution of technology in a beneficial direction.

He advocates for international coordination similar to that adopted to address nuclear weapons, with collective responsibility, global cooperation, and strengthened safeguards.

For the scientist, early evidence of misaligned behaviors does not signal imminent collapse, but it does indicate that society must act cautiously in the face of a technology whose power is growing rapidly — and whose governance framework is still under construction.

At Davos 2026, the ‘Godfather of AI,’ Yoshua Bengio, warns of the risks of the global AI race.

References cited in the text: Nature and The Outpost