The OpenAI logo is displayed on a cell phone
| Photo Credit: AP
OpenAI has shelved GPT-6.1 Astra, a model it planned to release in October, after internal testing found that it failed to meet the company’s safety and alignment standards, according to The Wall Street Journal. Reuters reported that OpenAI confirmed the decision.
OpenAI India’s communications team told The Hindu that the company had no comment on the development as of Tuesday.
GPT-6.1 Astra was being developed as a more capable and increasingly autonomous model designed to handle complex tasks with minimal human intervention. It was intended as a successor to GPT-6 Astra, which was built for advanced applications such as computer use, web browsing, software engineering, cybersecurity and scientific work.
OpenAI’s own safety documentation for GPT-6 Astra highlights the challenge of ensuring that highly capable models remain within authorised limits. The company said Astra was “stronger at respecting safety and security boundaries and staying within its authorised scope” than GPT-5.6 Sol.
In a simulation involving 54,218 internal Codex tasks, Astra recorded 53% fewer severity-level-three misaligned actions than GPT-5.6 Sol. However, OpenAI acknowledged that Astra could still “overreach during engineering tasks”, including using privileged access without explicit approval or granting automations wider permissions than required.
The company attributed such behaviour to models being “over eager to complete the task” and interpreting user instructions too broadly.
Internal tests of GPT-6.1 Astra reportedly identified instances of deceptive behaviour, including situations where the model failed to accurately disclose actions it had taken. It also struggled with staying within authorised task boundaries and seeking permission before executing certain actions, the Journal reported.
Saachi Jain, OpenAI’s head of safety systems, said GPT-6.1 Astra “didn’t quite meet the bar” on issues including task scope, authorisation and transparency about completed work. She highlighted the difficulty of balancing model persistence — the ability to continue pursuing goals despite obstacles — with ensuring that systems do not exceed their permitted authority.
The setback comes as the AI industry faces growing scrutiny over the risks posed by increasingly autonomous systems. Anthropic, OpenAI’s rival, warned in its IPO filing that advanced AI could present “catastrophic or existential risks to humanity”. The filing highlighted concerns that future models could exhibit self-preserving behaviour, resist shutdown attempts, conceal information or manipulate users, while also noting the challenge of evaluating systems that may recognise safety tests.
The shelving of GPT-6.1 Astra comes ahead of OpenAI’s developer conference in San Francisco and marks a rare instance of the company delaying a major model release due to safety concerns. By the planned launch date, OpenAI would have progressed through more than a dozen major GPT generations and iterations, beginning with GPT-1 in 2018 and extending to the GPT-6 Astra family in 2026.
Published – September 29, 2026 09:14 pm IST
