Not Every AI Failure Is a Crisis: Scoring Risk by Real-World Harm
Responsible AI

Not Every AI Failure Is a Crisis: Scoring Risk by Real-World Harm

By Kate Waldhauser Jul 27, 2026 10 min read
responsible AIAI governanceAI modelsAI adoption
TL;DR: A binary "did it fail or not" mindset over-reacts to harmless AI quirks and under-reacts to real ones. Scoring an incident on four axes, capability gain, breadth, ease of weaponization, and discoverability, gives your team a defensible way to size AI risk, and it maps directly onto what ISO 42001 and the NIST AI RMF already ask for.
Table of Contents

For nineteen days this summer, one of the most capable AI models in the world went dark.

On June 12, 2026, the US government applied export controls to Claude Fable 5 and Claude Mythos 5, cutting off access for users everywhere while officials and the company that built the models worked out how much real-world risk it actually posed. Access came back on July 1. In between, Anthropic shipped a new safety classifier and published testing that reframed how much the original finding actually mattered.

The trigger was a technique Amazon researchers found that got Fable 5 to identify software vulnerabilities and, in one case, produce code demonstrating how one of them could be exploited. That sounds alarming on its face. Then Anthropic ran the same prompts against other models and reported that every model it tested could produce the same demonstration, including ones as small as Claude Haiku 4.5.

So the whole episode turned on a question nobody had a shared way to answer: how dangerous was it, really?

That’s the question we want you to sit with, because your team faces a smaller version of it regularly. An AI tool does something it wasn’t supposed to do. Someone gets it to produce output it should have refused. A model surfaces information you didn’t expect it to have. In the moment, you have to decide whether you’re looking at a shrug or a fire.

Most teams don’t have a good way to tell the difference. That’s the gap we want to close.

I’m Kate Waldhauser, founder of Violet Beacon. We help organizations build AI governance that’s sized to the actual risk, and sizing the risk is where it starts.

Why does pass-or-fail thinking fail us?

Here’s how AI incidents usually get handled. Something goes wrong, and the response is one of two things. Either “it’s fine, that’s just how these tools are,” or “shut it down, we can’t trust it.” Pass or fail. Green or red.

Both reactions feel responsible. Neither one is governance.

A pure pass/fail mindset over-reacts to harmless quirks, burning a week and a lot of goodwill treating a cosmetic glitch like a breach. It also under-reacts to the genuinely risky ones, because when most alarms turn out to be nothing, people stop listening to the alarm. You end up with a team that’s anxious and numb at the same time, which is the worst of both.

What’s missing is proportion: a way to say “this happened, and here’s how much it actually matters,” with reasoning you could defend to a client, an auditor, or your own leadership.

🛡️ Responsible AI Note: A severity judgment only counts as a control if it’s written down and revisited. An undocumented gut call, however good the instinct behind it, isn’t something you can show anyone six months later. Capture the score and the reasoning, not just the outcome.

What four questions actually size an AI risk?

Frontier labs, facing this same problem at much higher stakes, landed on a set of questions worth borrowing. On July 2, 2026, Anthropic published them as a draft Cyber Jailbreak Severity (CJS) scale, developed with its Project Glasswing partners. The CJS scale is a shared rubric for rating how dangerous a given jailbreak actually is, rather than simply noting that one happened. When something bypasses an AI system’s guardrails, you score it on four axes:

  • Capability gain

    How far beyond their existing tools does this take the attacker? Or is the AI just a faster path to something a search engine and an afternoon would have produced anyway?

  • Breadth

    How many distinct tasks does the same technique work on? A narrow failure touching one edge case differs from one that opens a whole category of misuse.

  • Ease of weaponization

    How much human effort does it take to turn the failure into a running attack? Something requiring an expert, many manual steps, and a lot of luck sits very differently from something anyone could do on the first try.

  • Discoverability

    How easily can someone else obtain the technique in the first place? A quirk that took a researcher weeks to stumble onto is lower risk than one that's about to be posted in a forum.

The four axes of the Cyber Jailbreak Severity scale. Definitions paraphrased from Anthropic's published framework, which scores each axis and sums them into bands from CJS-0 (informational) to CJS-4 (critical).

Score those four and “an AI did something weird” becomes a rated risk with defensible reasoning behind it. It’s the same move the software security world made years ago with CVSS. Nobody says “there’s a vulnerability” and stops there. They say how severe it is, along which dimensions, and everyone knows roughly what the number means.

Worth noting what the labs left out. The scale covers cybersecurity jailbreaks only, and Anthropic explicitly excludes things like getting a model to reveal its system prompt. Anthropic also calls it “an early draft.” Reading it, we’d add that it doesn’t yet say who arbitrates when two companies score the same failure differently. We’re pointing at the logic here, not holding the document up as a settled standard.

Why isn’t “bypassed” the same as “dangerous”?

Look back at what actually got that model suspended. The behavior turned out to be reproducible across models already in wide public use. The guardrail was bypassed, and the capability behind it wasn’t unique, wasn’t hard to find elsewhere, and mostly amounted to work that defenders do every day.

Walk it across the four axes and a shape appears.

Capability gain
Low
Breadth
Low to moderate
Ease of weaponization
Low to moderate
Discoverability
High
Our reading of the Fable 5 jailbreak, not Anthropic's scoring. Anthropic has not published per-axis scores for this incident, so these are qualitative levels we inferred from its own reporting: capability gain is low because every model tested produced the same demonstration, and discoverability is high because the technique was written up publicly. We use words instead of numbers deliberately, since a decimal here would imply precision we don't have.

One axis runs high and the total still lands low. That’s the whole point: discoverability alone doesn’t make something dangerous when the capability was already sitting in every free model. “A bypass exists” and “this is dangerous” are two separate findings, and collapsing them into one is how a useful model ended up offline for nineteen days.

For your team, the translation is direct. When an AI tool does something off-script, start somewhere other than whether it broke a rule. Ask how far someone could actually get with it, and how easily. That’s the question that tells you whether you’re looking at a note-to-self or an all-hands.

Where does this already live in ISO 42001 and the NIST AI RMF?

If you’ve done any work toward ISO/IEC 42001 or the NIST AI Risk Management Framework, this should feel familiar.

The AI RMF’s Measure function, described in Section 5.3, exists to “analyze, assess, benchmark, and monitor AI risk and related impacts,” and it asks for metrics on trustworthy characteristics and social impact. A bare pass or fail doesn’t satisfy that. ISO 42001 clause 6.1.2 asks for an AI risk assessment that weighs likelihood and impact against your own risk criteria, which means sizing a risk before you decide how to treat it. A four-axis severity score gives you concrete vocabulary for something those standards already require. It adds precision rather than paperwork.

That reframe is worth holding onto, because “responsible AI” can start to sound like an endless list of new burdens. A lot of it is just doing, deliberately and on paper, the sizing-up that good teams already do in their heads.

🛡️ Responsible AI Note: When you tie an incident score to a named clause, the Measure function or your ISO 42001 risk assessment, you give an auditor or client a thread to pull. They can see what you decided and which standard you were answering to when you decided it. That traceability is the difference between “we handled it” and “we can show how we handled it.”

How do you put a severity scale to work without a frontier lab?

You don’t need a frontier lab’s resources to use the logic. Here’s the lightweight version.

Add a severity field to your AI risk register with a simple rubric, even a 1-to-5 scale anchored to the four axes. Decide in advance who scores an incident, who reviews that score, and what each tier triggers. When something comes up, someone scores it, writes down why, and the score maps to a response tier.

Tie each band to a response you’ve agreed on in advance:

SeverityWhat it looks likeResponse
LowThe tool did something odd, and anyone could have gotten the same result elsewhereLog it, review at your next governance cadence
MediumReal capability gain, but narrow, or hard to reproduceAssign a mitigation and a follow-up date
HighBroad capability gain that’s easy to repeatImmediate attention, named owner, escalation path you defined beforehand

That last column is the part teams skip. An escalation path you write during an incident is a guess, and a guess made under time pressure is the thing you’ll defend least well afterward.

The number matters less than what it does to the conversation. Once a score exists, the question shifts from “should we panic” to “where does this land, and what does that tier call for.” That’s a calmer, faster, more defensible place to make decisions from.

If you want a way to start this week, take one AI incident your team has already handled and score it retroactively. The places where two people disagree are exactly where your rubric needs sharpening.

🛡️ Responsible AI Note: A severity score informs the response, and it shouldn’t make the decision for you. A person still owns the call to escalate, mitigate, or accept, and that ownership should be named in writing. The rubric exists to make human judgment better, not to stand in for it.

What does proportion actually ask of you?

Responsible AI comes down to proportion: the discipline of asking how much something actually matters before you decide what to do about it. That sits between panic and complacency, and it asks more of you than either one.

A severity scale is how you make that discipline real, repeatable, and visible to someone who wasn’t in the room. The frontier labs needed a shared vocabulary because they have to explain themselves to governments. You need one because you have to explain yourself to clients, auditors, and your own team, usually under time pressure and usually without much warning.

If you’d like a starting rubric you can drop into your own AI risk register, that’s exactly the kind of thing Violet Beacon’s Responsible AI Guidelines are built to give you.

Key References

How AI Was Used in This Post

AI helped research the June and July 2026 Fable 5 news cycle, draft this post, and edit it. Kate Waldhauser set the angle, checked every claim against a primary source, and approved publication. The header image is AI-generated.

Frequently Asked Questions

What is an AI jailbreak?
+

It's a prompt or technique that gets a model to work around the safety rules its developer built in. The term covers a wide range, from coaxing a chatbot into an off-color joke to extracting genuinely dangerous technical instructions. Because that range is so wide, the word alone tells you almost nothing about how seriously to treat a given incident.

Why score an AI incident by real-world harm instead of by whether a safeguard was bypassed?
+

Because a bypass tells you the safeguard didn't hold, and nothing about what the person on the other side gained. A dramatic-looking failure can hand someone capability they already had from a search engine, while a mundane-looking one can unlock something genuinely new. Scoring by harm keeps your team focused on consequence rather than on technical novelty.

Do we need a formal severity framework if we're not a frontier lab?
+

You need the reasoning, not the apparatus. A single severity column in your AI risk register, anchored to a handful of named questions, does the same work at your scale that a formal published scale does at a frontier lab's. The point is that two people scoring the same incident should land in roughly the same place, and be able to explain why.

How does severity scoring connect to ISO 42001 and the NIST AI RMF?
+

Both standards ask you to assess AI risk in proportion to its impact rather than treating every issue identically, and both expect that assessment to be reviewable later. Neither one prescribes a scoring method. A severity rubric is one concrete way to satisfy those requirements and leave behind evidence that you did.

What belongs in a simple AI severity rubric?
+

At minimum: the four scoring axes, the resulting score, a sentence or two of reasoning, the name of the person who scored it, the date, and the response tier that score triggers. The reasoning field matters most. A bare number tells a future reviewer what you concluded but not whether your conclusion was sound.

Explore Related Services

AI Governance
AI Governance Consulting
Learn more →
ISO 42001
ISO 42001 Planning & Consulting
Learn more →
AI Strategy
AI Strategy & Advisory
Learn more →
Kate Waldhauser
Founder of Violet Beacon. Responsible AI consultant, ISO 42001 Lead Implementer, and Certified Claris Partner with 20+ years of custom software and database expertise.

Related Posts

← Back to all posts

Want to discuss this topic?

Book a free call to talk about responsible AI, FileMaker, or anything you've read here.