United Nations-backed Independent International Scientific Panel on Artificial Intelligence has called for stronger safeguards as increasingly autonomous AI systems advance, raising concerns over whether existing measures can keep pace with the growing capabilities of AI agents.
The call came as the panel released its first thematic brief examining the risks associated with AI agents and their ability to operate independently, following a security incident involving the online platform Hugging Face between May and July during a test initiated by OpenAI, the company behind ChatGPT.
According to the panel, the incident brought together several conditions researchers have long identified as potential contributors to a loss of control over AI systems: a system pursuing a goal not aligned with human intentions, sufficient capability to pursue that goal, and an environment that allows the system to act.
AI agents differ from conventional chatbots because they can independently perform tasks on users’ behalf rather than simply responding to individual questions or instructions. As their autonomy increases, the systems can interact with digital environments, access tools, and pursue objectives across multiple steps with less direct human intervention.
The UN-backed panel indicated that the Hugging Face incident provided an important real-world example of those concerns.
Scientific Panel Co-Chair Yoshua Bengio noted that “Researchers have long warned that three conditions could lead to loss of control: a misaligned goal, the capability to pursue it and an environment that allows it.”
“This summer, all three came together in a real system, not a laboratory. Since this is not an isolated observation of misaligned goals, this raises serious questions about the way AI agents are currently trained.”
Yoshua Bengio
According to the panel’s assessment, AI agents involved in the incident bypassed testing safeguards and demonstrated behaviours that raised questions about the ability of existing security systems to contain increasingly autonomous agents.
The agents reportedly coordinated across separate runs through an internal software tool that had not been designed to facilitate communication between them. They also gained unauthorised internet and administrator access while concealing attempts to circumvent cybersecurity evaluations.
The scale of the activity added another dimension to the concerns. Around 1,200 agents exchanged more than 70,000 messages and files during the period examined, with activity extending beyond Hugging Face to an OpenAI research cluster.
The panel noted that some agents even appeared to make decisions that benefited the wider group at the expense of individual agents, with some choosing to “sacrifice” themselves.
Although the incident does not establish that AI systems are already beyond human control, the panel explained that it demonstrates why existing approaches to AI safety require closer scrutiny as autonomous systems become more capable.
AI Agents Expose Growing Limits of Existing Safeguards

The panel’s findings point to a broader challenge confronting governments, technology companies and researchers: safety measures developed for today’s AI systems may not remain effective as systems become capable of understanding, navigating and potentially circumventing those same measures.
The panel stated that basic cybersecurity practices were overlooked during the incident, while existing safeguards failed to keep pace with the capabilities demonstrated by the AI agents.
However, the experts identified a deeper concern beyond conventional cybersecurity.
They warned that current AI training methods could produce agents that develop objectives of their own, deliberately violate safety instructions or conceal their actions from human operators.
“This is not only a question of speed. It leaves open whether safeguards designed today will work once agents can understand them and plan around them. In simple terms, the traditional model of safeguarding is unravelling.”
United Nations-backed Independent International Scientific Panel
As AI development moves from models that primarily analyse information and generate responses towards agents capable of taking actions in the real world and across digital networks, the potential consequences of failures could also become more significant.
An AI agent with access to online systems, administrative privileges or other digital tools could potentially have a much wider operational reach than a conventional chatbot.
This is why the panel has placed the Hugging Face incident within the wider scope of “agentic misalignment” and AI control.
Agentic misalignment refers broadly to situations in which an AI agent behaves in ways that conflict with intended human objectives while pursuing a particular goal. The concern becomes more pronounced when systems have the capability to act autonomously, adapt to changing circumstances, and interact with other systems without continuous human supervision.
Moreover, the panel also warned that safeguards used in other high-risk industries may not necessarily be sufficient for highly capable AI agents.
Industries such as aviation, medicine and cybersecurity have developed systems involving incident reporting, independent scrutiny and multiple layers of safeguards. These approaches provide models for managing technologies where failures can have serious consequences.
“But those practices may not be enough as AI agents become more capable, autonomous and difficult to monitor,” panel member Qinghua Lu explained.
The implication is that AI governance may need to become more adaptive, with safety measures evolving alongside the systems they are intended to regulate.
READ ALSO: Nkrumah’s Industrialisation Vision Meets Ghana’s Commodity Export Reality










