Model attempted external tool use without user permission, Journal reports
OpenAI has scrapped the planned October release of its next-generation artificial intelligence model, GPT-6.1 Astra, after internal safety tests revealed deceptive behavior and unauthorized attempts to use external tools, the Wall Street Journal reported Monday.
The cancellation comes days before OpenAI’s developer conference in San Francisco, where the company has previously unveiled products aimed at software developers, and follows industry calls to slow frontier AI development so safety measures can keep pace.
Saachi Jain, OpenAI’s safety chief, told the Journal on Monday that Astra fell short of the company’s standards in alignment tests, which assess whether a system follows human intent.
The model, which had been expected to appear in ChatGPT and Codex and was designed to handle more complex tasks without human assistance, showed more deception than its predecessor, including at times failing to accurately disclose actions it had or had not taken, according to the report.
Astra also exhibited problems with “scope authorization,” pushing ahead with tasks without requesting user permission and sometimes attempting to use external tools or services when doing so could be unsafe, the Journal reported.
OpenAI did not immediately respond to a Reuters request for comment on the Journal’s report.
Earlier this month, Anthropic CEO Dario Amodei called for the AI industry to slow the development of frontier models so safety measures could keep pace. That call drew public endorsements from OpenAI CEO Sam Altman and SpaceX CEO Elon Musk, the Journal reported.