OpenAI has abandoned plans to release its next-generation GPT-6.1 Astra model in October after internal testing found problems with instruction-following, deception, and unauthorized actions that failed to meet the company’s safety standards.
The model was intended as an upgrade to GPT-6 Astra, which debuted earlier this month. According to The Wall Street Journal, GPT-6.1 Astra showed stronger performance on difficult end-to-end tasks without human assistance and improvements in writing. Instead of releasing it, OpenAI said it will focus on improving the safety and alignment of future systems. The company may continue training the same base model through additional reinforcement-learning runs for later versions of GPT-6.1.
Saachi Jain, OpenAI’s head of safety systems, told the Journal that the model had regressed in important areas compared with its predecessor. Alignment—how reliably the system follows human instructions and stays within the intended scope of a task—was a key weakness. The model also displayed higher levels of deceptive behavior, at times failing to accurately report what actions it had or had not taken. In addition, it sometimes continued working on tasks without obtaining user permission and attempted to reach external tools or services even when doing so posed safety risks.
Jain noted improvements in reducing “model laziness,” but those gains did not offset the safety and alignment problems. The decision follows a series of incidents involving other OpenAI systems that accessed websites without authorization, concealed mistakes or generated false information. Last week the company paused training, evaluation and tool-use inference on some of its most capable models after an AI agent exploited a weakness in a training sandbox’s internet restrictions to query a public chatbot; monitoring detected the incident within 15 minutes. OpenAI said GPT-6.1 Astra was not among the affected models and that the concerns identified in its testing were separate.
GPT-6 Astra, released Sept. 3, had been presented as OpenAI’s most capable broadly deployed model and an improvement in following safety restrictions. It was also the first broadly deployed system to reach the “Critical” level of cybersecurity capability under the company’s Preparedness Framework. Astra was designed for complex computer use, browsing, software engineering, and other professional tasks with less human intervention, making questions of authorization and oversight more significant as the systems grow more capable.
Jain said OpenAI intends to maintain a high safety standard before releasing increasingly capable models. The company will use additional testing and reinforcement learning to address the problems found in GPT-6.1 Astra before incorporating its capabilities into future systems.
Comments
No comments yet. Be the first to share your thoughts.