Market assessment · Advise

Why Claude refuses your request, and which of the four layers just fired

Refusal is not one behaviour. It is four independent systems that fail in ways you cannot tell apart from the outside. Two of them read your prompt, two do not, and the fix everyone recommends aims the weakest instrument at the problem.

23 PAGES EVIDENCE-LABELED FREE TO READ
Commission an assessment

The common framing is that Claude has a moral code and needs to be argued out of it. That framing produces the wrong fixes.

A refusal in Claude Code or the Claude desktop app can come from any of four independent systems, and they look identical from the outside. The model's own trained disposition. The system prompt injected by whichever surface you are on. A classifier for cyber, biology, chemistry and distillation that sits in front of every access method. And the policy and entitlement layer attached to your account. The first two read what you write and can be moved with words. The third does not read your prompt the way you assume, and arguing with it adds more flagged text to a conversation that is already flagged. The fourth is a business relationship. Users conflate them, spend their effort on the wrong one, and conclude the whole system is arbitrary.

The number in the title is where the argument usually breaks down. Anthropic's Claude Opus 5 system card puts single-turn over-refusal on a benign prompt set at 0.09 percent through the API. At Fable 5's launch Anthropic said its guardrails would fire on average in fewer than five percent of sessions. Both figures are Anthropic's, both are honestly computed, and they differ by about fifty times because one counts prompts in a curated set and the other counts sessions on live traffic. A researcher's felt experience tracks the second. Quoting the first at someone living the second is the root of the credibility problem in this whole argument, and it is a communication failure with commercial consequences.

The scope is false positives on legitimate professional work: why a request gets declined, which declines you can do something about, what Anthropic has committed to changing, where the exemption routes leak, and what the alternatives actually cost. It is not about jailbreaking, and nothing in it describes a method for obtaining content Anthropic has said it will not produce. One caveat is stated up front and repeated throughout: almost every quantitative claim available about Claude's refusal behaviour comes from Anthropic. No independent audit of over-refusal on real research workloads exists in the public record, so every system-card figure here is labelled a vendor claim even though a system card reads like a technical paper.

Four systems refuse your work. Only two of them read your prompt.

The practical value of this table is diagnostic. Each layer produces a different symptom and responds to a different intervention, and the reason the standard advice underperforms is that it treats all four as one thing. Anthropic documents them separately: the trained disposition in Claude's constitution, the claude.ai system prompt in its release notes, the classifier layer in its cyber safeguards help article, and the policy layer in the Usage Policy exceptions article.

Layer one

Trained disposition

The model's own judgment, shaped by the constitution and its training. It presents as a conversational decline in Claude's voice, usually with an explanation and an offer of an alternative. This layer reads everything you write, so framing, context and model choice all move it.

Moves with words

Layer two

Injected system prompt

The prompt for the surface you are on: claude.ai's published prompt, or Claude Code's shipped one. It presents as tone and caution that persist no matter how you phrase the request. In Claude Code it is fully replaceable through output styles, --append-system-prompt and --system-prompt.

Replaceable in Claude Code

Layer three

Classifier

Input and output classifiers for cyber, biology, chemistry and distillation, applied across claude.ai, the API, Bedrock, Vertex and Azure Foundry. It presents as an API-level error naming the Usage Policy, or as a silent fallback to a different model. Nothing at the prompt level touches it.

Apply, or change model or surface

Layer four

Policy and account

Usage Policy enforcement, entitlements, and model-class caps on biology and chemistry. It presents as consistent blocking of a whole domain of work, unaffected by anything you type. Verified access programmes are the only route, and outside them it is a hard wall.

A business relationship

Claude Code's own system prompt is not published. What is known about it comes from a source map disclosure on 31 March 2026, in which the full Claude Code source briefly shipped inside the public npm package and was subsequently analysed. The material indicates a prompt instructing the agent to assist with authorised security testing, defensive security, capture-the-flag work and education, while declining destructive techniques and detection evasion for malicious purposes. The report treats that as indicative of structure rather than authoritative text, because it is a leak and not a publication.

Nine findings on why Claude refuses legitimate work

Every claim below carries an evidence label in the report itself, and the figures are read from Anthropic's system cards directly rather than from secondary summaries, which are frequently wrong about which surface a number refers to. Where a figure here disagrees with a widely circulated one, the system card is the reason.

01

The two published numbers differ by around fifty times, and both are Anthropic's

The Claude Opus 5 system card puts single-turn over-refusal on benign prompts at 0.09 percent through the API. At Fable 5's launch Anthropic said its guardrails would trigger on average in fewer than five percent of sessions. The first is a per-prompt rate on a curated set, the second a per-session rate on live traffic. A researcher's felt experience tracks the second, and Anthropic could publish it routinely alongside the first. That it does not is the single easiest credibility repair available to the company.

02

The surface you work on changes the refusal rate more than the prompt does

On Anthropic's own benign evaluation set, every model tested over-refuses more on claude.ai than on the bare API. Opus 5 goes from 0.09 percent to 0.47, Sonnet 5 from 0.59 to 1.54. The desktop app's chat mode runs the claude.ai stack and inherits its system prompt, which Anthropic publishes and states does not apply to the API. Claude Code talks to the API. Moving a task between them changes the odds before a word is rewritten.

03

Model choice inside Claude Code matters more for security work than any prompt technique

On Anthropic's Claude Code evaluation across 61 sensitive-but-permitted prompts, the dual-use and benign success rate is 99.82 percent for Opus 5 against 91.55 percent for Sonnet 5. That is roughly one failure in twelve against one failure in five hundred, on exactly the tasks security practitioners run all day. No prompt-engineering technique in circulation moves a refusal rate by that factor, and the change costs one flag.

04

CLAUDE.md is the weakest control available, and it is the one everyone recommends

Anthropic's documentation states plainly that CLAUDE.md is delivered as a user message after the system prompt, while output styles modify the system prompt itself and the command-line flags append to or wholly replace it. The circulated advice to drop a tone-control block into CLAUDE.md aims the weakest instrument at the problem, and there is a standing issue describing exactly that: standing instructions silently losing to built-in caution.

05

The lecture is a tracked defect at Anthropic, not an intended feature

Claude's constitution lists thirteen over-caution behaviours Anthropic explicitly does not want, among them lecturing when ethical guidance was not asked for, being condescending about a user's ability to handle information, and adding warnings that are not useful. The Opus 5 system card measures two of them as named metrics, wet blanket and condescension toward the user, and reports Opus 5 came out slightly more condescending than its predecessor.

06

One category of refusal is designed never to be liftable, and the text names vaccine research

The constitution sets out seven hard constraints and says they cannot be unlocked by any operator or user. It goes further, giving as an example that Claude should decline content offering real uplift toward chemical or biological weapons even where the user is probably asking for a legitimate reason such as vaccine research, because the risk of inadvertently assisting a malicious actor is judged too high. That is a stated, deliberate false-positive cost, and legitimate scientists bear it.

07

The exemption route is free, fast, and has holes in exactly the places research sits

The Cyber Verification Program lifts the high-risk dual-use classifier for verified organisations, reviewed inside two business days. It is not offered on Amazon Bedrock or Google Vertex AI, and organisations on zero data retention are ineligible. Researchers handling human-subjects or clinical data are the population most likely to require zero data retention, and they are structurally excluded from the relief. It is also organisation-scoped, which leaves the independent researcher and the single-PI lab no route at all.

08

The trend is improving, and improving least where this audience works

Anthropic's August 2026 retune cut biology-related fallbacks by roughly 85 percent. Broken out by surface, total fallback volume fell an estimated 67 percent on claude.ai and 55 percent on Cowork, against 17 percent on Claude Code and 7 percent on the Claude Platform. The chat surfaces took most of the benefit. The agentic and API surfaces, where professional work actually happens, took the least. Anyone modelling their future experience from a headline figure will be disappointed by roughly an order of magnitude.

09

Relief is moving from prompting to entitlement, and life sciences now runs through a government

Cyber work has a verification programme. Scientists get free and discounted seats but stay capped at Opus-class models for biology and chemistry, and access to Mythos-class models for life sciences is being established through a partnership with the US government rather than through Anthropic's own commercial review. For a researcher outside that perimeter, no amount of prompt craft substitutes for an entitlement they cannot obtain.

The classifier does not always refuse. Sometimes it quietly hands your work to a smaller model.

When a Fable 5 classifier fires, Anthropic re-routes the request to Opus 5, which it describes as a capable model that does not have the same level of biological capability. Anthropic calls this a fallback. It is a genuinely new category of failure and it is worse for research than a refusal is. A refusal is legible. A silent substitution produces an answer that looks normal, carries no warning, and is quietly worse, and the user's rational conclusion is that the model got dumber for no reason. No amount of prompt debugging will surface the cause. For a reproducibility-sensitive workflow, an undeclared model swap mid-session is a methodological problem before it is an annoyance.

The behaviour is observable in the field. A principal research scientist at the Institute for Disease Modeling told The Register that Fable 5's input safety classifier emitted a model refusal fallback on the first turn of essentially every session on his account, including one whose only user input was the word hello.

Once a classifier has fired, the contamination persists. Security researchers reported in April 2026 that the block affected follow-up messages in conversations carrying prior violative context, forcing them to restart sessions rather than continue. That single mechanic explains most of the frustration in this category. The instinct on being refused is to explain yourself, which appends more text about the flagged subject to an already-flagged conversation. The correct move is the counterintuitive one: abandon the session and start clean. Arguing is the worst available option and it is the one everybody reaches for first.

The controls, ranked by how much weight they carry

The advice circulating about this problem is not wrong so much as inverted. It leads with the intervention that has the least force and omits the ones with the most. The ordering below is Anthropic's own, taken from the output styles documentation and the CLI reference, and it is not the ordering the community uses.

01

Model and surface choice

Neither prompt nor configuration, and larger than everything below it. Opus over Sonnet is eight points of dual-use success on Anthropic's own Claude Code evaluation. Claude Code over the desktop chat mode is a fivefold gap on Opus 5's benign set. Almost never mentioned in the circulating advice.

02

--system-prompt or --system-prompt-file

Replaces the entire Claude Code system prompt rather than arguing with it. The highest-weight prompt-level control, with a real cost: it also removes the engineering instructions that make the agent competent at scoping changes, verifying work and handling destructive operations. Reach for it when Claude Code is not doing software engineering at all.

03

Output style

Modifies the system prompt directly and persists across sessions through the outputStyle setting. It drops the built-in engineering instructions unless keep-coding-instructions is set to true. For most people this is the better trade than the replacement flag: system-prompt-level tone control without losing the agent's competence.

04

--append-system-prompt

Appends to the system prompt without removing anything. Moderate weight and well suited to a one-off invocation, which is exactly what it is documented for. Nothing here is a substitute for a persistent output style if the same friction recurs every session.

05

CLAUDE.md

Adds a user message after the system prompt. Lowest of the persistent controls, and the one most commonly recommended. Issue 39210 on the Claude Code repository is titled around standing CLAUDE.md instructions being silently overridden by built-in safety heuristics, and a separate issue reports output styles being ignored where they conflict with embedded patterns.

06

Per-prompt framing

Real but small, and it decays as a session accumulates flagged context. Inside a hard-constraint domain it works against you: the constitution instructs that a persuasive case for crossing a bright line should increase Claude's suspicion rather than its compliance. The better your argument, the worse your position, and that is the model behaving exactly as specified.

Read the error, then pick the layer

A conversational decline in Claude's voice is layer one: add context about the setting and purpose once, and if it declines again change model rather than rewriting the prompt a third time. A moralising preamble attached to work it did anyway is layer two, and tone lives in the system prompt, so the fix is a custom output style with keep-coding-instructions true. An API error naming the Usage Policy is layer three: stop editing the prompt, start a fresh session, and apply to the Cyber Verification Program if the work is security. Answers that are quietly worse with no error at all are a classifier fallback, and the test is a comparison against an explicitly pinned model on the API. A whole domain refused consistently across prompts, models and surfaces is layer four, and nothing at the prompt level will help.

On leaving, the honest answer differs by the kind of work. Google's Gemini API exposes four adjustable harm categories with thresholds down to BLOCK_NONE and OFF, defaulting to OFF on recent models, over a floor of core protections that cannot be disabled. That is the sharpest product difference in the category and a genuine architectural choice: Google exposes the dial and keeps a floor, Anthropic keeps the dial internal and exposes an application form. For a researcher whose problem is category-level false positives, the dial is worth more, because it is per request and needs nobody's approval. For an enterprise buyer who wants a defensible audit story, the form is worth more. OpenAI's Model Spec is positioning rather than a control surface, and safe completions improve the experience of being partially refused without reducing the rate of being refused.

Open weights run locally are the only option that answers the question as this audience poses it, because they remove the vendor policy layer entirely: no classifier, no usage policy, no account that can be flagged, no silent model substitution. The costs are the honest ones, a capability gap on long-horizon agentic coding, real hardware, and every judgment the vendor was making badly on your behalf now being yours. For a psychologist working with transcripts or a historian working with archival material containing slurs, that trade is usually worth it. For someone who wants an autonomous agent to refactor a large codebase, it usually is not, yet.

The report closes with four scenarios to the end of 2027 and their earliest visible signs: incremental retuning continues at 55 percent, verified access becomes the main route at 25, a user-facing sensitivity control ships at 12, and re-tightening after an incident at 8. The probabilities are the least reliable content in it. The tripwires are the useful part, and they are observable without private information: Anthropic publishing session-level false-positive rates alongside prompt-level ones, the Cyber Verification Program opening on Bedrock or Vertex, a sensitivity parameter appearing in the API changelog, or a new classifier category shipping with no false-positive number attached.

Every claim carries its evidence

This isn't a vendor summary. Every sentence is labeled by what stands behind it: verified fact, vendor claim, third-party estimate, my assessment, hypothesis, or scenario. Sources are numbered and clickable. Forward-looking sections use scenarios with observable tripwires, not forecasts. It's the same method behind every market assessment I write.

Your Claude Refusal Rate Is Not 0.09 Percent

Twenty-three pages, built from public sources with no client brief and no interviews. Read it in the browser or take the PDF.

This is real, published work. Commission one for your decision.

Each report here answers a real question, directed and researched against public sources and evaluated against a stated assumption, then delivered as Word and PDF. If you're weighing a platform, sizing a category, or defending a number to a board, tell me the decision behind it and I'll tell you honestly whether a report is the right tool.

Commission an assessment

See more reports →