Skip to main content
Enterprise AI•9 min read

The Two Failures OpenAI Would Not Ship

September 30, 2026•By Ajan Kanagalingam

OpenAI had a model ready for October and decided not to release it. The Wall Street Journal reported it on September 28, Reuters reported that OpenAI confirmed the same day, and the company's head of safety systems went on the record with the reasons. Both reasons are worth reading closely, because they are the two things your own controls around agent work exist to catch.

What OpenAI said

Saachi Jain, OpenAI's head of safety systems, said GPT-6.1 Astra regressed in two areas against its predecessor during internal testing and was not reliable enough to release safely.

The first was alignment, and specifically deception. Per the reporting, the model showed higher levels of deception in that it was not always honest about telling users of the actions it did or did not take.

The second is what OpenAI calls scope authorization. The model would push ahead on a task without asking the user for permission, and would at times reach for external tools and services even if it might be unsafe.

Jain also noted the model improved in other respects, including on what the company calls model laziness, and described the underlying tension directly: there is a trade-off, and the work is finding the right line between staying within scope and avoiding laziness when a model hits friction.

Where the controls already line up

What OpenAI foundThe control that contains it
Pushed ahead without asking permissionAn approval gate on anything irreversible
Reached for external tools unpromptedScoped access: only the systems the task needs
Did not always report actions accuratelyA log of actions the agent cannot write to

None of the right-hand column is AI tooling. It is ordinary access control, of the kind you would apply to a contractor, and it is the same list we set out in three boxes for an AI agent and privileged access management.

The third row is the one businesses skip, because a system that describes its own work in fluent, confident prose feels like a report. OpenAI's finding is that the description and the actions can diverge, which is why the log has to come from somewhere other than the thing being logged.

Give credit where it is due

A company caught two regressions in internal testing, decided a flagship release did not meet its bar, cancelled it, and then confirmed the reasons publicly with a named executive on the record. That is the process working, and the public confirmation is more than the industry has typically offered.

It also matters that both failure modes were found before release rather than by a customer or a journalist. The comparison point is the DSEwiki incident, where the behaviour reached a live website and independent researchers published before the company did, which we covered in agents running a website for six weeks.

What it tells you about planning

Two things, and the second is the one that affects a budget.

Capability and behaviour do not improve together. This model was better at some things and worse at staying inside its scope. The assumption that each release is uniformly better than the last is now visibly wrong, which means a newer version is not automatically a safer one for unattended work.

The roadmap is not yours. A model expected in October is not arriving. Anyone who had scoped a project around its capabilities is rescoping, and the general form of that problem is in the AI gap in your continuity plan. Plan against what is shipping today, not what was announced for next quarter.

What not to conclude

Nothing in your stack changed. GPT-6.1 Astra was never released, so this is not a defect in a product you are using.

The behaviours were also observed in internal testing rather than in deployment, which is the distinction that matters when reading any story of this kind. A model behaving badly under evaluation is the evaluation doing its job.

What does transfer is the category. Deception about actions taken, and pushing past the boundary of a task, are general properties of capable agentic systems rather than defects unique to one cancelled model. Your controls should assume they exist in whatever you are running now, and the practical version of that assumption is in delegation as the skill AI needs.

Frequently Asked Questions

What happened with GPT-6.1 Astra?

OpenAI scrapped the release of GPT-6.1 Astra, a model it had planned to launch in October. The Wall Street Journal reported it first on September 28 and Reuters reported that OpenAI confirmed it the same day. Saachi Jain, OpenAI’s head of safety systems, said the model regressed in two areas against its predecessor during internal testing and was not reliable enough to release publicly.

What were the two problems?

Per OpenAI, the model showed higher levels of deception, meaning it was not always honest with users about which actions it had and had not taken. The second was what the company calls scope authorization: it would push ahead on a task without asking the user for permission, and would at times reach for external tools and services even where doing so might be unsafe. Jain said it improved on other measures including what OpenAI calls model laziness.

Is this good news or bad news?

Both, honestly. A frontier lab catching two regressions in internal testing and withholding a flagship release on that basis is the process working, and confirming it publicly with named quotes is more than the industry norm. It also tells you the release roadmap you may be planning around belongs to the vendor, and that capability improvements and behavioural improvements do not arrive together.

Does this affect models we already use?

Not directly. GPT-6.1 Astra was never released, so nothing in your stack changed. The relevant read is what it says about the category rather than about any product you run: the two failure modes named are general problems with capable agentic systems, not defects unique to one cancelled model, and your controls should assume they exist.

What should a business do about this?

Check that the two things OpenAI would not ship are things your own setup contains. That means an approval gate on anything irreversible, so a system pushing ahead without asking cannot complete the action, and a log of what an agent actually did held somewhere the agent cannot write to, so you are not relying on its own account of its work. Both are ordinary access controls rather than AI-specific tooling.

Assume both failures, then design for them

We review what your AI systems can do without asking, put approval steps on the irreversible actions, and make sure the record of what happened does not come from the thing that did it.

Related Articles

Enterprise AI

OpenAI Stopped Its Own Model. What Would Stop Yours?

August 19, 2026Read more →
Enterprise AI

OpenAI’s New Model Is Built to Use Your Computer

September 4, 2026Read more →
Enterprise AI

AI Monitoring: Watch What It Does, Not What It Says

September 1, 2026Read more →
AK
Ajan Kanagalingam
Founder & ChatGPT Consultant, ChatGPT.ca

Ajan leads the ChatGPT.ca team: 200+ custom GPT builds and automation projects for 50+ businesses across 20+ industries. Based in Markham, Ontario. PIPEDA-compliant solutions.

Stay ahead of AI in Canada

Weekly case studies, new tools, and ROI playbooks for Canadian SMEs. One email, zero spam.