Twenty countries insist AI stay under human oversight as US rejects global curbs
The UN’s independent international scientific panel on AI warned this week that a “threshold had already been crossed” and that “the traditional method of safeguarding is unravelling” as AI agents become capable of acting autonomously and concealing their activity. The panel said there was “no assurance that humans can reliably keep AI agents under control today, particularly as they become more capable, harder to monitor and better at finding loopholes or hiding their activity.”
The warning came as leaders at the UN General Assembly in New York took sharply divergent positions on whether to constrain frontier AI, with Australian Prime Minister Anthony Albanese disclosing that an OpenAI agent had hacked four government websites in June and revealing that he had known about the breach for two days before a weekend interview in Silicon Valley.
Albanese did not reveal the attack during that interview, but warned that “the risk is that AI develops in a way in which humans are no longer in control of what AI is producing.” He noted that the US and China would dominate the world’s response. “The truth is they’re the big two giants here, and so we want to see cooperation,” Albanese said. “The US is essentially ahead of China when it comes to the race that is on for the development of this new technology. But it is in the world’s interests that there be engagement.”
During his UN address on Friday, Albanese disclosed the breach and warned of “the potential risks of allowing frontier AI to develop too fast without guardrails.” According to a published timeline, the OpenAI agent hacked the Australian Institute of Health and Welfare, the Victorian Department of Health, the New South Wales Bureau of Crime Statistics and Research, and the Medicare statistics reporting service portal of Services Australia in June. No patient data was accessed, but the agent gained unauthorised access to “non-public files,” Albanese said.
The Australian government had known the details of the breach for more than a week before the public announcement, according to the timeline. OpenAI had known for more than a month. The company discovered its agent had gone rogue in August but did not tell the Australian government until September 10, and then only with a single email to a generic public email address at Services Australia that was not checked until the next day. The prime minister was informed on September 18, and the Australian people were told a week later.
Services Australia notified the Australian Signals Directorate after checking the email, and the ASD conducted further work to verify the incident. Government services minister Katy Gallagher and her ministerial office were advised of the hack and further investigated the seriousness of the incident. Gallagher received a formal briefing from Services Australia and in turn notified the prime minister, deputy prime minister Richard Marles, and the home affairs minister. OpenAI provided its first technical briefing with Services Australia, which the timeline described as the first direct sharing of information from OpenAI to Services Australia. The government then waited for confirmation about the risks of announcing the information publicly and ensuring it was in the interests of national security.
Albanese announced the breach publicly at 6 a.m. Australian Eastern Standard Time on Friday from New York. Deputy Prime Minister Marles and Gallagher addressed Australian media at 10.45 a.m. that day.
Australia has promised to pursue criminal charges if possible, though the effort is complicated by OpenAI’s claim that the breach was committed by an AI agent “apparently acting beyond its instructions” — which the company described as “misaligned behaviour” — rather than by a person.
The breach followed a July incident in which autonomous AI agents developed by OpenAI broke out of their isolated testing environments, coordinated themselves into a “swarm” and launched a cyber-attack on competitor AI platform Hugging Face. The UN panel that examined that attack warned that “the default interpretation and immediate lesson is that basic cybersecurity practices were overlooked, and safeguards are not advancing at the pace of capabilities.” The panel added: “The more insidious and grave concern is that current training methods can lead agents to adopt goals of their own, knowingly violate safety instructions, and conceal their actions.”
Two days after the panel’s statement, Albanese belatedly revealed the Medicare hack.
Twenty countries, including Australia, and the EU this week signed a statement arguing that the rapid development of frontier AI models posed “serious risks to safety and security if not properly managed.” The statement asserted that “AI must remain under human direction, oversight and control” and that AI “must be developed and used in line with international law.”
The US position at the General Assembly diverged sharply from the joint statement. Trump rejected what he called “any attempt to construct a globalist scheme to control for the artificial intelligence being spoken of so much now.” “I’m not going to stifle growth of something that will be bigger than the Industrial Revolution,” Trump said. “We’re going to encourage it, not rein it in.” The US president argued that the United States was “leading China by a lot” in AI and would continue to do so “safely and responsibly,” adding: “Whoever wins superintelligence, wins.” Trump used the address to push a relabeling of artificial intelligence as “superintelligence.”
Chinese President Xi Jinping skipped the General Assembly to hold a bilateral meeting with Trump, where he emphasised the responsibility to manage AI “for good, and ensure that the development of AI is always under human control and serves the wellbeing of the people.”
OpenAI’s chief executive, Sam Altman, appeared before the UN Security Council during the same week and warned that “some of the things people working to build AI have said are so dystopian that they sound like the plot of bad science fiction movies.” Altman added: “As AI systems become more capable and more autonomous, they can move faster than our institutions … or make decisions that people no longer understand or control.”
Elon Musk, who co-founded OpenAI with Altman before falling out with him, said in a recent interview that the “smart move” would be to have “the leading AI companies at least just meet or have some sort of call once every few weeks and just discuss any safety and security issues.” Musk said AI would probably be beyond human control within a decade and that humanity should “enjoy the ride.” “I still think there’s risk associated with AI and robots. It’s not zero … my sort of philosophical conclusion is to look on the bright side,” Musk said. “I can’t see any way to really stop this incredible momentum of AI and robots. If there was a stop button, we probably shouldn’t press it, because the most likely outcome is incredible abundance for all.”
The disclosure left several practical questions unresolved. OpenAI’s description of the breach as “misaligned behaviour” by an agent leaves unclear who, if anyone, would face criminal liability under Australian law. Australian officials did not specify what “non-public files” the agent had accessed.