Concerns about AI safety have escalated in recent months after models developed by OpenAI and rival lab Anthropic were involved in security incidents during testing.
Agents built with OpenAI’s models have inappropriately accessed websites maintained by US federal agencies, an Australian government health statistics portal and Hugging Face, a repository of AI models.
OpenAI apologised on Tuesday for not properly responding to the Australia incident, which involved its AI models accessing government websites without authorisation.
“We are sorry and working to do better in the future,” OpenAI said in a blog post, adding that the company would explain “what we know, what we have changed, and what we will do to rebuild trust with the Australian people”.
“Our aim was to give affected agencies a detailed account once our investigation was complete,” the ChatGPT-maker said. “However, we should have shared preliminary findings sooner and kept Australian agencies updated as more facts emerged.”
OpenAI, Anthropic and other major AI developers have promised to prioritise making models that have safety guardrails to mitigate risks and that are also aligned with human values.
American chip making giant Nvidia announced on Tuesday that it created a system designed to stop autonomous AI programs from straying beyond what they were instructed to do.
“I believe it’s an engineering problem … and we all need to hope that’s an engineering problem,” Nvidia CEO Jensen Huang told broadcaster CNBC.
“If it’s not an engineering problem, it’s not solvable,” he added.
The AI Security Institute (AISI), an initiative under the UK government, published a study on Tuesday showing that GPT-6 Astra went off the rails more often during testing than its predecessors, GPT-5.6 Sol and GPT-5.5.
In simulations, GPT-6 spontaneously carried out cyberattacks at rates significantly higher than those observed for the other two interfaces.
– AFP



