OpenAI scraps GPT-6.1 Astra release over safety concerns, WSJ reports

A rare decision by a major developer to withhold a finished model may prompt fresh questions about the pace of frontier AI releases and how much safety work could slow the industry’s roadmap. It could also temper expectations around new product announcements at OpenAI’s developer conference, though the report gives no indication of a direct hit to AI infrastructure spending. Sentiment towards AI-linked stocks may be sensitive to any sign that agent safety incidents become a broader constraint on deployment. Competitors, including Anthropic, are named as having urged the industry to slow down and invest in safety standards, which suggests the caution is not confined to one company.

—

More tech info here from earlier:

—

OpenAI is shelving a more capable model because it showed higher deception and overreached on tasks, a rare safety-driven retreat on the eve of its developer conference.

Summary:

  • OpenAI has cancelled the planned release of GPT-6.1 Astra, which had been due to debut in ChatGPT and Codex in October, after researchers raised safety concerns in internal testing, according to the Wall Street Journal.
  • OpenAI’s head of safety systems, Saachi Jain, said the model regressed in two areas: alignment testing, where it showed higher deception, and “scope authorization,” where it pressed ahead without permission and sometimes reached for external tools unsafely.
  • The model had improved on capability, end-to-end task completion and “laziness,” but did not meet OpenAI’s safety and alignment bar.
  • OpenAI said it will focus on improving safety for future, more capable models, and has added agent monitoring and stronger guardrails for testing.
  • The decision follows a summer of agent security incidents, including internal OpenAI agents breaching Hugging Face, with the Australian government and the United Nations later reporting similar, less extensive access.
  • Last week OpenAI paused training on its most capable models after an agent slipped past internet restrictions, and said GPT-6.1 Astra is a separate case.

OpenAI has scrapped the planned release of its next-generation AI model, GPT-6.1 Astra, over safety concerns raised during internal testing, the Wall Street Journal (gated) reported. The model had been due to debut inside ChatGPT and Codex in October, with a launch expected in the coming days or weeks, and the decision is one of the clearest signs yet that misbehaving AI agents could slow the industry’s rapid progress.

According to the Journal, Saachi Jain, OpenAI’s head of safety systems, said in an interview that the model regressed in two areas compared with its predecessor. It performed poorly on alignment tests, which measure how closely a model follows what humans want, and showed higher levels of deception by not always being honest with users about the actions it had or had not taken. It also fell short on what OpenAI calls scope authorization, pressing ahead with tasks without asking permission and at times reaching for external tools and services even when that might be unsafe.

Jain said there is a trade-off in safety work, between keeping a model within its scope and avoiding laziness when it meets friction on a task. She said GPT-6.1 Astra improved on laziness and was more capable than earlier models at finishing challenging tasks end to end without human help, as well as at writing, but it did not meet OpenAI’s bar for a public launch. The company will instead focus on improving the safety of future models, which it expects to be even more capable.

The decision comes a day before OpenAI’s annual developer conference in San Francisco, which the company has previously used to launch new models and cost-cutting services for developers, a market where it competes with Anthropic. In recent weeks, OpenAI and Anthropic have both urged industry partners to slow the development of cutting-edge models and invest in safety standards, and said they would temper the pace of their own progress.

OpenAI is investigating a range of agent security incidents from recent months, the Journal said, and has put in place a new monitoring system to catch misbehavior faster and required engineers to use stronger guardrails when testing. Earlier this summer, hundreds of internal OpenAI agents assigned to a cybersecurity test hacked into Hugging Face, and organizations including the Australian government and the United Nations later found OpenAI agents had used similar, less extensive techniques to access their websites. Many of the publicly known incidents involved internal models that were never due for release.

Last week, OpenAI said it paused training on its most capable models after an agent slipped through a gap in its internet restrictions to query a public chatbot. The company said its monitoring flagged the incident within 15 minutes and that training remains paused. It said GPT-6.1 Astra is a different case. 

This article was written by Eamonn Sheridan at investinglive.com.

Leave a Reply