HomeEnterprise ITArtificial IntelligenceOpenAI scraps GPT-6.1 Astra launch after safety tests flag unauthorised actions

OpenAI scraps GPT-6.1 Astra launch after safety tests flag unauthorised actions

OpenAI has cancelled the planned October release of GPT-6.1 Astra after internal testing found the model evaded oversight and operated beyond its authorised scope.

Preferred Source of Google

Key Points

  • OpenAI cancels GPT-6.1 Astra release after model evaded oversight and misrepresented actions
  • Internal OpenAI model gained unauthorised access to Australian Medicare portal in June
  • Company paused training of most capable models after one bypassed network restrictions

OpenAI has scrapped the planned October release of GPT-6.1 Astra after internal tests found the artificial intelligence model did not consistently remain within authorised limits and, in some cases, failed to accurately report the actions it had taken.

The company confirmed the decision after The Wall Street Journal first reported that the next-generation model would not be released. GPT-6.1 Astra had been expected to be integrated into and OpenAI’s Codex coding product and was designed to carry out more complex tasks with less human assistance.

Advertisement
National DefTech Summit
National DefTech Summit
Featuring keynotes, expert panels, live tech demos and strategic networking, the summit will drive actionable insights for defence sector.
Register Now →
Infosec Reimagined
Infosec Reimagined
Infosec Reimagined 2026 is the premier information security summit where top leaders—CISOs, CROs, CIOs, CTOs and risk executives—converge to redefine cyber resilience.
Register Now →
Digital Senate
Digital Senate
Digital Senate is a premier conference uniting government leaders, technologists and innovators to share ideas, success stories and strategies on digital governance, public sector transformation, cybersecurity and emerging technologies in India.
Register Now →

The decision is significant because the concerns were not primarily about the model producing an incorrect answer. They involved what an increasingly autonomous AI system did while pursuing a task, including whether it sought permission before taking actions and whether it accurately told users what it had done.

Saachi Jain, OpenAI’s head of safety systems, said the model showed improvements in some areas but fell short on staying within the scope and authorisation given to it and on communicating its actions to users.

“When we ship it to users, we have an extremely high bar in terms of safety and alignment,” Jain said.

Advertisement

The cancellation comes as OpenAI is conducting a broader examination of how its experimental agents have behaved while using access during training and evaluation.

The company has introduced a framework for publishing instances in which its models take unauthorised actions, evade oversight or expose weaknesses in existing safeguards.

One of the most prominent incidents emerged in last week.

Advertisement

Australian Prime Minister Anthony Albanese said an internal OpenAI model gained unauthorised access to the country’s Medicare Statistics Reporting Service portal on June 18 while conducting research into public medicine spending.

According to Albanese, the agent encountered repeated blocks while searching for information but tried alternative methods and eventually accessed both public and non-public files. It also wrote files to an internal server. A forensic investigation involving the Australian Signals Directorate is under way.

The Australian government has said there is currently no evidence that personal Medicare information was accessed and no evidence of a wider compromise of the Services Australia network. The affected system is a public-facing statistics portal containing information including Medicare expenditure .

Albanese said OpenAI notified Services Australia on September 10, almost three months after the June incident, and criticised both the delay and the way the notification was made. Australia has established a task force to examine the incident and whether existing processes are adequate for AI-related cyber events.

OpenAI has separately disclosed another incident in which an internal research model found a gap in restrictions intended to prevent it from reaching the live internet.

The agent was carrying out a web research task when it discovered that a DNS resolver in its training environment could be used to communicate with an external chatbot.

OpenAI said its monitoring system detected the behaviour within about 15 minutes, although the run was not stopped until around two and a half hours later.

Following that incident, OpenAI said training, evaluation and inference involving tool use for its most capable models remained paused while it strengthened network restrictions and carried out additional testing.

The concerns are not limited to OpenAI’s own testing.

Britain’s AI Security Institute said its pre-release evaluation of GPT-6 Astra, the predecessor to the abandoned 6.1 version, found the model carrying out unauthorised supply-chain attacks in simulated cybersecurity environments.

In the tests, GPT-6 Astra completed a simulated supply-chain attack in 29.2 per cent of runs, compared with 6.3 per cent for GPT-5.6 Sol and none for GPT-5.5 in a smaller set of tests.

The institute stressed that the actions took place in simulations and caused no real-world harm.

The behaviour declined when researchers made the limits of the task more explicit but did not disappear. In a subset of tests where the model was specifically told that anything not listed as in scope was out of scope, it still completed four supply-chain attacks across 49 simulated runs.

The UK institute also noted that OpenAI’s standard safeguards were disabled for parts of its testing in order to measure what the underlying model would attempt without those interventions.

It said the findings nevertheless pointed to the need for controls such as sandboxing and monitoring in addition to attempts to align the model itself.

The incidents disclosed so far need to be distinguished from attacks initiated by ordinary ChatGPT users. The Australian case involved an internal OpenAI research model, the DNS incident occurred during reinforcement-learning training and the UK tests were controlled simulations conducted before GPT-6 Astra’s public release.

GPT-6.1 Astra itself was not released to users.

OpenAI now plans to continue work on the GPT-6 family rather than ship the version that failed its internal safety threshold. The company is also investigating why the model developed the problematic behaviours identified during testing.

Your Questions, Answered

Why did OpenAI cancel GPT-6.1 Astra?

Internal testing found the model evaded oversight, misrepresented its actions and operated beyond its authorised scope. It also attempted to use external tools it knew were unsafe.

What happened with the Australian Medicare portal?

An internal OpenAI model gained unauthorised access to Australia's Medicare Statistics Reporting Portal on 18 June while researching public medical spending. It accessed both public and non-public files after encountering repeated access blocks.

Did OpenAI models access US government websites?

Yes, OpenAI models accessed publicly available information on SEC.gov, Investor.gov and Census.gov during training and evaluation. The company acknowledged agents acted inappropriately but said no private data was stolen.

What are AI alignment problems?

Alignment problems refer to AI models lacking understanding of acceptable behaviour boundaries. Such models may pursue objectives regardless of restrictions, seeking alternative paths when blocked from their initial approach.

NEWSLETTERThe Daily BriefingThe day's top enterprise technology stories, curated by our editors. Monday to Friday.

Free. One-click unsubscribe anytime. We never share your email.

Tech Observer Desk
Tech Observer Desk
Tech Observer Desk at TechObserver.in is a team of technology reporters led by a senior editor who brings latest updates and developments from the world of technology.
Advertisement
- Advertisement -
- Advertisement -

Apple patches CoreGraphics zero-day exploited in targeted attacks

Apple has patched a CoreGraphics zero-day vulnerability that attackers exploited in sophisticated targeted attacks. The flaw, reported by Meta Product Security, marks the seventh zero-day Apple has fixed this year.

RELATED ARTICLES