---
title: "Five frontier AI labs were asked how they would shut down a rogue model. None gave a full answer"
description: "Guidelight AI Standards assessed Anthropic, Google, OpenAI, Meta and xAI on whether they have published procedures for containing a model that behaves contrary to instruction. It found few containment protocols ready for an emergency. OpenAI scored highest at three out of five, and even there the report found no formal plan for future incidents."
category: "Tech"
category_url: https://boursel.com/category/tech
author: "Hannah Blackwood"
published: 2026-08-22T19:57:00.000Z
updated: 2026-08-22T19:57:00.000Z
canonical: https://boursel.com/article/five-frontier-ai-labs-were-asked-how-they-would-shut-down-a-rogue-model-none-gav
tags: ["ai-safety", "openai", "anthropic", "regulation", "governance"]
---
# Five frontier AI labs were asked how they would shut down a rogue model. None gave a full answer

Guidelight AI Standards assessed Anthropic, Google, OpenAI, Meta and xAI on whether they have published procedures for containing a model that behaves contrary to instruction. It found few containment protocols ready for an emergency. OpenAI scored highest at three out of five, and even there the report found no formal plan for future incidents.

Guidelight AI Standards has assessed Anthropic, Google, OpenAI, Meta and xAI on a narrow question: what would they actually do if a deployed model started behaving in ways they had not intended and did not want. Its conclusion, [reported by TechCrunch](https://techcrunch.com/2026/08/22/frontier-ai-labs-still-wont-say-how-theyd-contain-a-rogue-model/), is that on the publicly available evidence the companies have few containment protocols ready for an emergency.

OpenAI scored highest at three out of five, largely because it has visibly paused workloads after incidents. Meta and Anthropic scored lowest.

## What the labs said

OpenAI's response was the most specific: it has ["a process for requiring restricting permissions, pausing workloads, limiting deployment, or taking the model fully offline, and have applied it"](https://techcrunch.com/2026/08/22/frontier-ai-labs-still-wont-say-how-theyd-contain-a-rogue-model/). Even so, the report found no evidence of a formal plan for how it would respond to future misalignment incidents.

Anthropic said it would conduct a risk assessment ["focused on determining whether containment is the appropriate response"](https://techcrunch.com/2026/08/22/frontier-ai-labs-still-wont-say-how-theyd-contain-a-rogue-model/) if it detected a model attempting to evade its constraints. Google said the report "doesn't represent the full scope" of its safety and security work, without answering whether internal containment plans exist. Meta declined to say whether it has one.

## The distinction the story turns on

Not disclosing a procedure is not the same as not having one, and the report is careful about this. It measures what is public, and the labs are entitled to argue that publishing the details of a shutdown mechanism is itself a security risk.

But two of the responses go further than declining to publish. Google's answer avoids the question rather than declining it. And Anthropic's phrasing describes a decision to be taken later about whether containment is warranted, which is a different commitment from having a procedure ready to execute.

That gap, between having a process and having a pre-committed plan, is what the assessment is measuring, and it is a meaningful one. An emergency procedure that has to be designed during the emergency is not an emergency procedure.

## Why a financial publication is covering this

Because the disclosure is becoming a commercial term rather than an ethical one.

Enterprise and government buyers increasingly ask AI suppliers to evidence their controls as part of procurement, in the same way they ask for security certifications and business continuity plans. A vendor that cannot describe what happens in a failure scenario is harder to onboard, and the risk of not being able to describe it transfers to the customer, who has to explain it to their own regulator.

The regulatory floor is still low. The [NIST AI Risk Management Framework](https://www.nist.gov/itl/ai-risk-management-framework) is voluntary, and no US regulator currently requires a lab to publish containment procedures. Several states are legislating, and OpenAI has publicly argued that California should strengthen its AI safety bill, which is a notable position for a company to take about rules that would bind it.

The practical effect is that the standard is being set through contracts before it is set through statute, which is how a lot of technology governance actually happens.

## What would change the picture

A published incident-response procedure with defined triggers, named decision-makers and committed timelines, of the kind that exists as a matter of routine in banking, aviation and pharmaceuticals. None of the five has one in public.

Until then, the honest summary is that the largest AI developers have asked customers and governments to trust that controls exist, and have declined, to varying degrees, to describe them. Whether that is prudence about disclosing security details or the absence of the thing itself is not currently possible to determine from outside, which is the report's actual finding.

## Sources

- [Frontier AI labs still won't say how they'd contain a rogue model](https://techcrunch.com/2026/08/22/frontier-ai-labs-still-wont-say-how-theyd-contain-a-rogue-model/)
- [AI Risk Management Framework](https://www.nist.gov/itl/ai-risk-management-framework)

