"safeguards flagged this message" means that a safety classifier attached to the Claude model reacted to the content of your conversation and stopped that model from responding. It is not a Claude Code bug, and it is not a network failure. It can appear not only during cybersecurity or biology work but also on ordinary requests such as a greeting or a planning task, and Anthropic itself acknowledges that false positives can happen.
The message starts with the name of the model that was stopped. On GitHub Issues filed between September 22 and 29, 2026, the on-screen wording reported was mainly one of the following two (the line breaks are not exactly as shown on screen).
API Error: Opus 5.5's safeguards flagged this message (https://www.anthropic.com/legal/aup). This sometimes happens with safe, normal conversations. Claude Code can't respond to this message with Opus 5.5.
Double press esc to edit your last message, or try a different model with /model.
Send feedback with /feedback or learn more: https://support.claude.com/en/articles/8106465
Details: `[reasoning_extraction]`
API Error: Opus 5.5's safeguards flagged this session (https://www.anthropic.com/legal/aup). You may be seeing this for the first time on an Opus model: Opus 5.5 is more capable and has stronger safeguards as a result, which can sometimes flag non-cybersecurity work. We're improving these safeguards to reduce the amount of incorrectly flagged messages. Claude Code can't respond to your last message with Opus 5.5.
Double press esc to edit your last message, or try a different model with /model.
Send feedback with /feedback or learn more: https://support.claude.com/en/articles/8106465
Details: `[cyber]`
Issue titles often say "Message flagged by safeguards," but what actually appears on screen is the wording above. The Opus 5.5 part changes to Fable 5.1, Opus 5, Opus 4.8 and so on, depending on the model you were using.
Look first at the last line, "Details:"
What works depends on which classifier stopped you
Judged to be trying to extract the model's internal reasoning
Judged to be work that could be used for a cyberattack
Judged to possibly involve dangerous uses of biology
Judged to fall under other areas of the usage policy, or frontier AI development
Contents
- 1. What the message means—a classifier that read the conversation stopped it
- 2. The wording differs by model and version
- 3. The Details categories, and whether the model switches
- 4. Did you do something wrong? What Anthropic says about false positives
- 5. Cases reported since September 22
- 6. What to do when you get stopped
- 7. Does the Cyber Verification Program (CVP) help?
- 8. Billing for stopped requests, and unfinished work
- 9. What is confirmed and what is not
- 10. Summary
- FAQ
1. What the message means—a classifier that read the conversation stopped it
According to the official Claude API documentation "Refusals and fallback," Fable 5.1, Fable 5, Opus 5.5, Opus 5 and Sonnet 5.5 come with safety classifiers that can refuse requests. When one refuses, the API does not return an error but a normal response (HTTP 200), with the stop reason set to stop_reason: "refusal" and the area of the refusal in stop_details.category. Claude Code receives this and shows it as a message starting with "API Error:".
The important point is that the classifier does not look only at the last sentence you sent. The Help Center article explaining the Opus 5.5 safeguards says the classifiers look at everything the model reads. That includes memory, connector content, web search results and files, so you can be stopped by content you never typed yourself.
Your instructions
The last message you sent, plus the entire exchange before it.
What Claude has read
Files it opened, command output, web search results, and content coming from MCP or connectors.
Context attached from the start
Workspace information Claude Code adds to the first request, such as the contents of CLAUDE.md and git status.
The model's own output
It can also stop mid-response. In that case, whatever had streamed so far is treated as incomplete too.
On the third point, Claude Code's model configuration documentation states explicitly that a classifier can trigger on the first request of a session, before you have sent anything unusual. Because the first request includes the contents of CLAUDE.md and git status, a repository containing security or biology material can trigger the classifier through that information alone—that is the explanation. The fourth point corresponds to the API documentation saying that a refusal can arrive before the output or partway through it, and that in either case the partial output should be discarded as incomplete.
2. The wording differs by model and version
The same mechanism produces several different wordings. The current example listed in the official error reference looks like this.
API Error: Opus 4.8's safeguards flagged this message. Our intentionally broad safeguards allow us to deliver more capabilities faster, but can sometimes flag legitimate cybersecurity work. Apply to the Cyber Verification Program to reduce these interruptions. Send feedback with /feedback or learn more: https://support.claude.com/en/articles/14604842-real-time-cyber-safeguards-on-claude
The error reference goes on to say that on Opus 5.5 and Sonnet 5.5, the message starts with <model>'s safeguards flagged this session. The two messages quoted at the top of this article were pasted into Issues filed on or after September 22, 2026. Within what we read, the flagged this session message came with [cyber] or [bio], while the This sometimes happens with safe, normal conversations message mostly came with [reasoning_extraction].
| How the message begins | Where it was confirmed | What follows |
|---|---|---|
<model>'s safeguards flagged this message. Our intentionally broad safeguards… | Official error reference (reported in Issues together with [cyber]) | A pointer to apply for CVP (when a switch happens, the fallback model is included too, as in Switched to Opus 4.8.) |
<model>'s safeguards flagged this session | Official error reference (Opus 5.5 and Sonnet 5.5), and Issues on v2.1.280–2.1.283 | An explanation that you may be seeing this for the first time on an Opus model, because it is more capable and so has stronger safeguards |
<model>'s safeguards flagged this message (…aup). This sometimes happens… | Issues on v2.1.278–2.1.284 (no example in the official reference) | Says it "sometimes happens with safe, normal conversations" and suggests editing with Esc or using /model |
<model> has safety measures that flagged this message for a cybersecurity topic | Official error reference (older wording, v2.1.203–2.1.218) | A link to the Help Center (before v2.1.203, a link to the exception request form) |
<model> can't help with this. Start a new session to continue. | Official error reference (Usage Policy refusal) | A refusal from the Usage Policy check (on Amazon Bedrock, Google Cloud's Agent Platform and Microsoft Foundry, cyber flags also use this wording) |
The error reference says the classifiers themselves run on the server side and existed before v2.1.203; what later Claude Code updates changed was only the wording of the message. It is not the case that updating Claude Code is what started getting you stopped.
What changed in v2.1.284
The CHANGELOG for v2.1.284, released on September 28, 2026, has two entries related to this message.
- The notice shown when a Sonnet model's safeguards flag a message was changed to explain why it happened and to offer the option of editing and resending.
- Safeguard-triggered model switching was changed for sessions that pin an Opus model with
ANTHROPIC_DEFAULT_OPUS_MODELormodelOverrides. On the Anthropic API, the API now picks the fallback model for each type of flag, rather than using the pinned model.
As of September 29, the wording of the new Sonnet notice was not listed in the official error reference, and it had not been pasted into any of the Issues we read.
3. The Details categories, and whether the model switches
The word in Details: at the end of the message matches the value of stop_details.category returned by the API. The API documentation lists five categories. Note that every category except reasoning_extraction carries a note that it can also trigger on benign work.
| Details | Official description (summary) | How Claude Code handles it |
|---|---|---|
cyber | Could be used for cyberattacks, such as developing malware or exploits | Fable 5.1, Fable 5, Opus 5.5 and Opus 5 switch to Opus 4.8, and Sonnet 5.5 switches to Sonnet 5, resending the same request |
bio | Could lead to biological harm, such as dangerous lab techniques | Fable 5.1, Fable 5 and Opus 5.5 switch to Opus 5; Opus 5 and Sonnet 5.5 have no fallback and end in a refusal |
frontier_llm | Could help develop competing AI models (restricted by the commercial terms) | According to the Help Center, Opus 5.5 switches to Opus 5 (Opus 5 does not switch for this category) |
reasoning_extraction | Tries to get the model's internal reasoning reproduced as response text | The Help Center's "distillation" category; it does not switch to another model and stops right there |
general_harms | Falls under a usage policy area other than the four above | Claude Code's documentation does not list a fallback model |
Sources: the category table in Claude API "Refusals and fallback," Automatic model fallback in Claude Code "Model configuration," and the Help Center article "Why Claude switched models in your conversation with Opus 5 or Opus 5.5" (all checked on September 29, 2026).
When a switch happens
By default, when Claude Code is stopped in a category that has a fallback, it automatically resends the same request on the fallback model and posts a notice in the conversation. The rest of the session continues on the fallback model. To go back to the original model, pick it again with /model.
The Help Center cautions, however, that even if you switch back to the original model, the same safeguards may switch you again if the flagged request is still in the conversation, and that editing the earlier message before resending often helps. Typical situations where the switch does not help and you end up stopped with a message like the ones at the top are: being stopped in a category with no fallback, such as reasoning_extraction; being stopped on the fallback model as well; the fallback model not being allowed by availableModels; and running in a non-interactive mode such as -p with the "ask before switching" setting enabled.
4. Did you do something wrong? What Anthropic says about false positives
If you feel "I didn't do anything wrong, and I still got stopped," that is entirely possible. Anthropic's documentation clearly acknowledges false positives in several places.
✅ What the official documentation says
- The message itself: "This sometimes happens with safe, normal conversations" and "intentionally broad safeguards"
- The API documentation:
cyber,bio,frontier_llmandgeneral_harmseach carry a note that they can also trigger on benign work - The Claude Code documentation: with offensive security or biology work, a switch often happens from the very first request; in these domains this is expected routing, not an account flag
- The Help Center: the classifiers are being tuned to reduce false positives, and will also take into account various signals of account trustworthiness
- The CVP explanation: even approved users can have legitimate work stopped
On the other hand, the Help Center article linked from the message (Our Approach to User Safety) says that users who repeatedly violate the policies may temporarily have safety filters with heightened detection applied to them. Whether times you were stopped by false positives count toward this is not stated officially. After being stopped, it is safer to switch conversations using the steps in Section 6 than to keep throwing the same request at it with only the wording changed.
How to think about a reasoning_extraction stop
The Help Center gives examples of what this category stops: requests to repeat its reasoning word for word, or to write out its full chain of thought into external output. On the other hand, it says you can still have Claude explain its reasoning, teach a concept, or walk through a code review step by step, and that conversational questions such as "why did you do that?" are not affected.
Reading the Issue reports, there were cases where the stop came right after requests like the following. All are user reports, and Anthropic has not explained why the classifier reacted.
- Asked whether the screen could show its "thinking" instead of the tool calls (#96118)
- Had a subagent quote, from its context, the line naming the model it was running on (#96139)
- At the end of a long task, had it write a summary of the session to a file (#96874)
- Asked it to annotate a design diagram with "the options considered and why they were rejected" for each decision (#98108)
What they have in common is asking the model to write out "its thought process" or "its own context" as-is. If you need that as a deliverable, switching to wording that asks for a short summary of the result, such as "the conclusion and the reasons for choosing it, in three lines," moves the request toward the "explain its reasoning" side that Anthropic describes as allowed. That said, there are cases like the additional report on #96118, where a short question in German alone was stopped, so rephrasing is not guaranteed to avoid it.
5. Cases reported since September 22
Since Opus 5.5 was released on September 22, 2026, reports of this kind have been piling up in the anthropics/claude-code Issues. When we counted with GitHub search on September 29, there were 67 Issues created between September 22 and 28 containing "safeguards" in the title or body (7–12 per day; the count changes depending on the search terms). In the Issues from September 22 onward that we read, we found no replies from Anthropic staff, and many carried the "duplicate" label.
🟡 Cases reported on GitHub whose cause has not been established (checked September 29, 2026)
- Stopped by just a greeting: "hi" was judged
[cyber]and switched to Opus 4.8 (#96298), "ola" was stopped (#96495), and "hi" got[reasoning_extraction](#97560). The reporter of #96495 speculates that using Kali Linux as root may have played a part. - Stopped during /compact: the
[reasoning_extraction]message appeared in the middle of compacting the conversation (#96648, v2.1.281). - Once stopped, the whole conversation becomes unusable: every later message in the same session was stopped (#96874), and the flagged phrase stayed in the conversation log so that later sessions reading it were stopped too (#97452).
- Subagents get stopped: a tally from a single machine found that 14 of 866 Opus 5 subagents (1.62%) were stopped with
[reasoning_extraction]. All of the stopped ones were review-type requests (#97492). - Stopped even with CVP approval: approved users were still judged
[cyber]on Opus 5.5 (#96090, #96298, #96592, #97946). As Section 7 explains, this matches the official explanation. - Stopped while working on their own machines or their own code: investigating an SSH "Permission denied" on their own company's systems (#97373), disassembling a binary of their own company's product (#97891), and removing credentials hardcoded in source (#97605).
The cases stopped by just a greeting might be explained by the official statement in Section 1 that "the classifiers look at the whole conversation and the workspace information." The idea is that even if the last message was harmless, the classifier reacted to CLAUDE.md, the earlier exchange or files that had been read. However, none of the reports shows what it reacted to, and no way of checking this has been published.
6. What to do when you get stopped
These are the steps listed by the official error reference and the model configuration documentation, arranged from least to most effort. If it happened only once, step 3 will usually be enough.
Check Details, and whether the model switched
If you see Switched to or a switch notice, your work is already continuing on another model. If it is still stopped, match the category in Details: against the table in Section 3.
Check file operations that were in progress
If it was stopped mid-response, the output up to that point is incomplete. First use git status and git diff to see whether any half-written changes were left behind (Section 8).
Press Esc twice to go back, and send it a different way
Pressing Esc twice with the input box empty, or running /rewind, opens the rewind menu. Pick a point before the stopped turn, then change how you phrase or approach the request and send it again. This is the first remedy Anthropic lists for this message.
If you can't tell which turn caused it, start a new conversation
Use /clear to start a new conversation in the same project. The original conversation is not deleted and can be opened from /resume. However, if you resume the original conversation with --continue or --resume, the flagged content comes back with it, so you are likely to be stopped the same way.
Switch to another model with /model
The message itself recommends you "try a different model with /model." That is because different models carry different classifiers. For example, the Sonnet 5.5 announcement says it is the first Sonnet to ship with a classifier that guards against reasoning extraction, and the previous-generation Sonnet 5 does not have this classifier.
Find out whether CLAUDE.md or similar is the cause
If you are stopped from the very first message, launch with claude --safe-mode. It starts without loading CLAUDE.md, skills, MCP servers or hooks, so if you are not stopped that way, something among what was being loaded contained the content that triggered it. git status and the directory name are still sent in this mode.
If it's a false positive, report it with /feedback
Anthropic advises that if it was not a cybersecurity topic, you report the false positive with /feedback. The Request ID and Message ID shown in the message are useful when reporting. The conversation log contains the code you were working on and file paths, so if you post in a public place such as GitHub, paste only the lines you need.
If you'd rather choose each time than switch automatically
If you feel the quality of the work drops on the fallback model, turn off "Switch models when a message is flagged" in /config, or write the following in your settings file. The session will then pause each time it is stopped and let you choose between switching to the fallback model and editing the prompt to resend it on the current model.
{
"switchModelsOnFlag": false
}
However, for categories with no fallback (such as bio on Sonnet 5.5 or Opus 5), no choice is offered and it ends in a refusal. In the non-interactive -p mode, and in SDK integrations that cannot show a prompt, the stopped turn also ends in a refusal. In cloud sessions used from the mobile app, the option to edit and resend is not available.
If it stops during /compact
Compacting a conversation also has the model read the whole conversation, so it can be stopped too (#96648). The v2.1.282 CHANGELOG says it fixed compaction failing when the summary request was refused, and that it now retries on the fallback model. If claude --version shows a version older than 2.1.282, update first with claude update. We cover when compaction is worth using in our article on when to run /compact.
7. Does the Cyber Verification Program (CVP) help?
CVP is a free application program that makes it less likely for people doing cybersecurity work for legitimate defensive purposes to be stopped by the safeguards. The Help Center's CVP article divides the work that gets stopped into two kinds.
Prohibited use
Work used almost exclusively for abuse, such as mass data exfiltration or ransomware development. Stopped even with CVP approval.
High Risk Dual use
Work that can also serve defense, such as exploiting vulnerabilities or developing offensive tools. Stopped by default, but you can apply through CVP to have it relaxed.
The key point here is that as of September 29, 2026, CVP does not apply to Opus 5.5 or Sonnet 5.5. The CVP article says at the top that it applies to Opus and Sonnet models but not to Opus 5.5 and Sonnet 5.5, and that it will soon be extended to Opus 5.5, Sonnet 5.5 and Mythos-class models. Both the Opus 5.5 announcement and the Sonnet 5.5 announcement also describe the CVP expansion as coming "soon." The reports we saw in Section 5 of "CVP-approved but still stopped on Opus 5.5" are consistent with this explanation.
The Opus 5.5 Help Center article offers one more clue: a sentence saying that organizations already using Opus 4.8 through CVP can start using Opus 5, which has fewer cyber restrictions, right away. If you are approved and still being stopped by Opus 5.5, picking Opus 5 or Opus 4.8 with /model for now is the workaround that follows the official explanation.
| What to check | What the Help Center says |
|---|---|
| Who can apply | Apply through the Verification Portal (Claude.ai, Claude Code, Anthropic API) / only admins with the right permissions see the application screen / identity verification is required / the goal is to notify you of the review result by email within 2 business days |
| Environments not covered | Zero Data Retention (ZDR) organizations are not eligible at this time / not offered on Amazon Bedrock |
| If you're still stopped after approval | Approval is tied to a specific organization ID, so compare it with the organization ID in the approval email to make sure you are not using a different organization, such as a personal workspace / prohibited use is stopped even after approval |
| If something still seems wrong | You can report a false positive or appeal a rejection through the form in the article |
For organizations whose biology research is stopped by [bio], there is a separate vetting program called the Life Sciences Verification Program (LSVP). The Opus 5.5 announcement says that organizations that pass the review can use Opus 5.5 for their research.
8. Billing for stopped requests, and unfinished work
Billing depends on the category and on when it stopped
According to the API documentation and the Help Center, stopped requests are billed as follows. In every case, they count toward your rate limits (usage limits).
Stopped before any output: billed
bio, frontier_llm, reasoning_extraction. The official explanation is that these categories have few false positives (as of September 2026).
Stopped before any output: not billed
cyber, general_harms, or when there is no category (null).
Stopped mid-response
The input and the output streamed before the stop are billed at normal rates.
The response on the fallback model
Billed separately at the fallback model's rates. The cost of losing the cache is made up for with credits.
If you use a subscription, the effect shows up in how fast your usage limit drains. The reporter of #97335 writes that the prompt cache written by a stopped request was not used by the next request, so in a long conversation of about 700,000 tokens, every stop meant rewriting the entire conversation. They observed using up a Pro plan's 5-hour window with four refusals. It is a report not confirmed by Anthropic, but if you keep getting stopped in a long conversation, switching with /clear costs you less of your limit than persisting in the same conversation.
Half-finished file operations can be left behind
The API documentation asks that the output of a response stopped partway be discarded as incomplete. The note Claude Code adds to the conversation after being stopped, in the text pasted into #97335, also says that tool calls that had not finished were not executed.
However, #97311 reports four cases where the model was stopped while writing a file edit (Edit), a file write (Write) or a Bash call, and the truncated content was executed as-is (v2.1.219–2.1.280, on a single Linux machine). For Edit and Write, "success" was displayed, but the written file or the rewritten lines were cut off partway. There is no official response yet, but right after being stopped, it is safest to check the changes with git diff before moving on, as in step 2 of Section 6.
# List files that changed
git status
# Look at the contents for edits that were cut off partway
git diff
# Restore one file to its state at the last commit (after checking)
git restore path/to/file
Changes made with Claude Code's file editing tools can also be undone with checkpoint rewind. However, files rewritten through Bash are not covered by rewind.
9. What is confirmed and what is not
✅ Officially confirmed
- The classifiers run on the server side; Claude Code updates changed only the wording
- They look not at the last sentence but at everything read: the conversation, files, CLAUDE.md and so on
- Safe, normal conversations can also be stopped
reasoning_extractiondoes not switch to another model- As of September 29, CVP does not cover Opus 5.5 or Sonnet 5.5
- Since v2.1.282, a refused /compact is retried on the fallback model
🟡 Reported but unconfirmed
- Being stopped by just "hi" or "ola" (#96298, #96495, #97560)
- Once stopped, sessions that read that conversation keep getting stopped (#96874, #97452)
- Truncated file operations get executed (#97311)
- The cache from a stopped request is not reused (#97335)
🔴 Not disclosed
- What any individual message reacted to
- Whether times stopped by false positives affect how your account is treated
- The date CVP will be extended to Opus 5.5 and Sonnet 5.5
- The wording of the Sonnet notice changed in v2.1.284
10. Summary
"safeguards flagged this message" means that the model's safety classifier reacted to the content of the conversation and stopped that model from responding. The classifiers run on the server side and look not only at the last message but at the whole conversation, the files that were read, and even CLAUDE.md. Both the message itself and the official documentation acknowledge that ordinary conversations can be stopped too.
When you are stopped, first look at the category in Details:, and if a file operation was in progress, check it with git diff. Then try, in order: go back with Esc twice and rephrase, start a new conversation with /clear, and switch models with /model. If [cyber] is blocking your work, CVP is an option, but as of September 29, Opus 5.5 and Sonnet 5.5 are not yet covered. For the overall picture of the model-side safeguards, see our Opus 5.5 explainer; for other "blocked by a filter" messages, see our article on The model returned no content; and for other errors, see our roundup of common Claude Code errors and fixes.
FAQ
Q. What does "safeguards flagged this message" mean?
A. It means the safety classifier of the model you were using (such as Opus 5.5 or Fable 5.1) reacted to the content of the conversation and stopped the response. It is not a Claude Code malfunction or a network error. Details: on the last line contains the name of the category that reacted.
Q. I only sent "hi" and got stopped. What did I do wrong?
A. Your last message is not necessarily the cause. The classifiers look at the whole conversation, the files that were read, and even CLAUDE.md and git status. If you launch with claude --safe-mode and are not stopped, it may have reacted to the content of the settings or files being loaded. If you think it's a false positive, report it with /feedback.
Q. Could my account be suspended?
A. The Claude Code documentation says that model switches during security or biology work are expected routing, not an account flag. On the other hand, a Help Center article says that users who repeatedly violate the policies may temporarily have stronger filters applied, and whether false positives count toward that has not been disclosed.
Q. Why am I stopped even though I have CVP approval?
A. Because as of September 29, 2026, CVP does not apply to Opus 5.5 or Sonnet 5.5. Anthropic says it will be extended soon. Anthropic also writes that organizations already using Opus 4.8 through CVP can use Opus 5 with fewer cyber restrictions, so for now the workaround is to pick Opus 5 or Opus 4.8 with /model. Also check that the organization ID tied to your approval matches the organization you are using.
Q. It keeps stopping no matter what I send in the same conversation. What should I do?
A. While the flagged content remains in the conversation, anything you send tends to get the same judgment. Press Esc twice to go back to a turn before the stop, or start a new conversation with /clear. You can open the original conversation with /resume, but resuming brings the flagged content back with it.
Primary sources we consulted
- Claude Code — Error reference (official documentation): Safety measures flagged a cybersecurity topic (current and older wording, v2.1.203 and v2.1.219), Usage Policy refusal, Responses seem lower quality than usual
- Claude Code — Model configuration (official documentation): Automatic model fallback, Check what triggered fallback, Ask before switching, Security research and biology workloads
- Claude Code — Settings reference (official documentation):
switchModelsOnFlag - Claude Code — Checkpointing (official documentation): Esc twice and
/rewind, what rewind covers - Claude Code — Changelog (official): the entries for v2.1.282 and v2.1.284
- Claude API — Refusals and fallback (official documentation): the list of categories, how billing works
- Claude Help Center — Why Claude switched models in your conversation with Opus 5 or Opus 5.5: switching by category, the distillation category, the relationship to CVP, billing
- Claude Help Center — Real-time cyber safeguards on Claude Opus and Sonnet: CVP coverage, applications and appeals
- Claude Help Center — Our Approach to User Safety: detection models and safety filters
- Anthropic — Claude Opus 5.5, Claude Sonnet 5.5 (the Safeguards section of each announcement)
- GitHub Issues: #96090, #96118, #96139, #96298, #96495, #96592, #96648, #96874, #97311, #97335, #97373, #97452, #97492, #97560, #97605, #97891, #97946, #98108 (all user reports; checked September 29, 2026)