Is AI Slop Real, or Did We Just Name It?
Responsible AI

Is AI Slop Real, or Did We Just Name It?

By Kate Waldhauser Aug 10, 2026 15 min read
responsible AIAI governanceISO 42001measurement
TL;DR: AI-generated content volume has clearly risen, but that doesn't prove the underlying defect rate went up, because almost nobody measured the human baseline it's being compared against. For regulated manufacturers, the fix is routing AI-assisted defects into the nonconformance, rework, and CAPA systems you already run.
Table of Contents

Everyone’s calling it AI slop now, and everyone agrees there’s more of it. What almost nobody agrees on, or even asks, is whether the rate of bad work actually rose above what it was before AI, because that’s a baseline nobody measured.

In 2022, Europol’s Innovation Lab published a report on deepfakes that included a line about the future. It speculated that as much as ninety percent of online content might be synthetically generated by 2026 (Europol).

It’s 2026. The number is about fifty percent (and it stopped climbing a year and a half ago).

I want to be fair to Europol. They said “may.” It was a speculation in a report about deepfakes, not a forecast about text, and Europol’s own updated version of the report later quietly dropped the line (after acknowledging the projection had come from an unreliable source). But that ninety percent got loose. It circulated for years in slide decks and articles and pitches as though someone had measured it, and now the year it named has arrived and we can check.

That’s the thing I’ve been trying to do with all of this. Check.

I’m Kate Waldhauser, founder of Violet Beacon. We help regulated manufacturers build AI governance that survives an audit, and measurement is where that work actually starts.

Are we talking about more content, or more defects?

I’m not going to argue about the volume, because the volume is real. AI-generated articles briefly overtook human-written ones in late 2024 and have held at roughly half of new web content since, based on AI-detection analysis of a large sample of the open web (Graphite). Merriam-Webster made “slop” its 2025 word of the year. So did the American Dialect Society. So did Macquarie Dictionary, and so did The Economist.

When three dictionaries independently reach for the same four letters, something happened.

Here’s the part that gets left out. That same research found roughly eighty-six percent of what actually ranks in Google is human-written, and about eighty-two percent of what ChatGPT and Perplexity cite (Graphite). The flood is real. It’s mostly not what anyone is reading.

That still isn’t the question I had. I kept catching myself doing a particular comparison. I’d look at something a person made a few years ago and something a person made last week with AI help, and think: this got worse.

Then at some point I asked what I was actually comparing against. Not the good version I remember. The average one. The real one.

The question isn’t whether there’s more AI-generated content. Obviously there is. The question is whether the rate of bad work went up. Those are different questions, and almost nobody was asking the second one.

So I went looking for what the work was like before.

What did quality actually look like before generative AI?

It was not good.

Newspaper accuracy research goes back to a 1936 methodology and has told the same story the whole time. When researchers went to the sources quoted in stories and asked them to check, sources found factual errors in somewhere between roughly half and six in ten stories: forty-eight percent in the US, sixty percent in Switzerland, fifty-two percent in Italy, in one cross-national study (City Research Online). The correction rate stayed around two or three percent whether or not anyone flagged the error. Most of it just stayed wrong.

Medical research is worse than you’d hope. A systematic review put the quotation-error rate in medical journal articles at 14.5 percent (PLOS ONE), and about two-thirds of those were major: the cited source didn’t support the claim, or contradicted it. Reference errors across fields run eleven to forty percent depending on who’s counting. Paper mills were industrializing fabricated research years before anyone had heard of ChatGPT. More than ten thousand papers were retracted in 2023, a record at the time, and over eight thousand of those came from one publisher, Hindawi, after Wiley found its peer review had been systematically compromised (Nature).

Software has known its own numbers for decades. Fifteen to fifty defects per thousand lines of delivered code is the ordinary range. One study found fifteen percent of Android apps contained security-relevant code copied off Stack Overflow, and ninety-eight percent of those copied snippets were insecure. Another found two-thirds of traced Java snippets were already outdated when they got reused. Copy-pasting stale, broken code is an old habit. AI just makes it faster.

Audit quality has the same problem. The PCAOB, which inspects public-company audits, reported deficiency rates climbing from 28.4 percent in 2020 to 43.7 percent in 2023 across all inspected firms, before easing somewhat in 2024 (The CPA Journal). That’s a risk-targeted sample rather than a random one, so read it carefully, but it isn’t a rounding error either.

And the web was already full of garbage. Google shipped its Panda update in 2011 specifically because content farms had gotten good enough at flooding results that about twelve percent of US queries needed fixing: Demand Media, Associated Content (which Yahoo bought for a hundred million dollars), human beings paid by the piece, producing text nobody wanted.

Theodore Sturgeon said ninety percent of everything is crud. He said it in 1957, about science fiction, and he was being defensive rather than rigorous. Nobody has measured it and probably nobody can. But the intuition is old, and it predates every technology we’ve since blamed for it.

AI isn’t off the hook here. But the bar we’re measuring it against is one we made up.

Why can’t the most-quoted slop statistic settle this?

If you’ve read anything about AI slop in the past year, you’ve met the workslop study: forty-one percent of desk workers received AI-generated work in the past month that looked finished and wasn’t, costing nearly two hours per incident to sort out, about a hundred eighty-six dollars per employee per month, or nine million dollars a year at a ten-thousand-person company. Managers get it worse than individual contributors: forty-two percent thought less of the person who sent it (HBR).

I believe all of it. It’s careful work and the finding matters.

It also cannot tell us whether the rate went up.

It’s a survey. It asks people how they feel about work other people sent them. There’s no 2019 version to compare against: nobody thought to ask, before AI, what share of the work landing on your desk was useless. The authors say as much themselves. There has always been sloppy work.

So we have a number everyone is using to prove the rate rose, and it can’t. It might have risen. It probably did in places. But the study measures the perception of a thing, and we’re treating it as measurement of the thing. Those aren’t the same, and I’d rather say so than round it off.

I sat with that a while and didn’t much like where it left me.

🛡️ Responsible AI Note: A statistic that measures how people feel about the work they received is not the same as a defect rate. Both are useful. Only one belongs in a management review.

Where does the evidence actually hold up?

There are cases where somebody was already counting. Those are the ones worth taking seriously.

Daniel Stenberg maintains curl, and curl ran a bug bounty from 2019, eventually paying out more than a hundred thousand dollars across eighty-seven confirmed vulnerabilities. In the years before 2025, north of fifteen percent of submissions turned out to be real. Last year, that fell to around five percent. He shut the program down at the end of January this year, moving vulnerability reports to GitHub instead (daniel.haxx.se). That’s a rate, measured the same way before and after.

GitClear has been tracking what happens to code across two hundred eleven million changed lines from 2020 through 2024. Copy-paste passed refactoring for the first time in 2024. Duplicated code blocks are up eightfold. Code churn, meaning code revised within two weeks of being written, went from 5.5 percent in 2020 to 7.9 percent in 2024 (GitClear). These are maintainability signals rather than confirmed defects, and I’d be careful how far to push them. But there’s a real before-number, which is more than most of this field can offer.

Notice what those two have in common. They aren’t better arguments. They’re better instrumented.

🛡️ Responsible AI Note: The only organizations that can say whether AI made their output worse are the ones that were already measuring output. Instrumentation is what lets you answer the question at all.

What matters more than whether AI made it?

Researchers ran 1,528 simulations of what happens when an AI system retrieves and cites content that AI wrote. Just under eighty percent collapsed: the answers converged until every response named the same things in the same order (Graphite). This is a different failure mode from model collapse, where models degrade after training on their own output (Nature), but the two rhyme: both describe a system losing information by feeding on itself.

The driver wasn’t that the sources were AI-generated. AI-written sources got cited about nine percent of the time, human-written ones about seven: barely a gap. What broke things were the sources the model itself had authored, cited at nearly thirty-nine percent, and that gap held up after controlling for eight separate quality measures.

Whether an independent judgment had entered the loop mattered far more than whether the source was machine-made.

That’s what Simon Willison said in 2024, before any of this had numbers attached. He defined slop as content that’s unrequested and unreviewed, not simply AI content: a claim about how the work was handled, not about what produced it.

The dictionaries dropped that distinction. Merriam-Webster’s definition is about low quality produced in quantity by AI: provenance and volume, review nowhere in it. Macquarie kept “not requested by the user,” which I think is to their credit, because once slop just means made by AI, the rate rose by definition and there’s nothing left to investigate. The question eats itself.

What’s the strongest case against this argument?

The other side is stronger than I’d like.

Francesco D’Isa argues the whole category is confused: mediocrity has always been the baseline of culture, not a defect some machine introduced (The Philosophical Salon). He’s not wrong. Clip art, WordArt, stock photography, Baudelaire complaining that photography industrialized bad taste: every cheap new medium has been accused of drowning the good stuff, and the accusation has been roughly correct every time and roughly beside the point every time.

There’s money in the panic, too. Detection vendors sell detection. Agencies sell “human-written” at a premium. Some of the best research on AI content volume comes from an SEO firm, and while I think their methodology is careful, they’re measuring a problem their industry gets paid to solve. That’s worth saying out loud, including about the numbers I’ve leaned on here.

And the detectors are a mess. One study found seven commercial detectors flagged more than half of TOEFL essays by non-native English speakers as AI-generated, some unanimously, and simplifying the language of essays written by native speakers made detectors more likely to flag those too (Patterns). Every essay in the study was written by a person. Vanderbilt turned its detector off. If your control depends on detecting AI, your control has a false-accusation rate you probably can’t defend to whoever you just accused.

Set against that, there’s decent evidence AI-assisted work can be better. A Harvard Business School and Boston Consulting Group field experiment with seven hundred fifty-eight consultants found meaningful quality gains on tasks inside AI’s capability frontier, and a nineteen percent drop in correct answers on a task outside it (Harvard Business School). Provenance doesn’t predict quality. Task fit and review do.

I don’t have a clean resolution. I think slop is partly real, partly a new name for an old defect rate, and partly a story people are selling. Which of those dominates depends on how you define the word, and we haven’t agreed on that.

What should a regulated manufacturer do about it?

Here’s where it stops being interesting and starts being useful.

If you run a regulated manufacturer, don’t build an AI slop program. Don’t buy a detector. Don’t write a policy that says “review AI output” and call yourself covered, because that’s a sentence, not a control.

None of the examples above come from a factory floor, but the tool for handling them does. First-pass yield, scrap and rework, nonconformance, CAPA: you already run that system on physical product. AI-assisted work is a new source of defects entering it, not a new category that needs its own program.

Do this instead, in three moves.

  1. Define the metric. A deliverable that goes out with fabricated citations is a nonconformance. Log it as one. The two hours somebody spent finding and fixing it is rework, and it belongs in your cost of poor quality next to every other kind of rework you already track. One number: share of AI-assisted deliverables returned for correction, and hours spent. No separate AI dashboard: you already have the system.
  2. Put a gate in the life cycle. A verification checkpoint before anything goes out, with a named person signing off. If the deliverable cites sources, verify the sources. That single control would have caught Deloitte, and it’s the exact checkpoint that was missing in a very different kind of AI governance failure. Add a disclosure line. Update the policy to say where AI drafting is allowed and where it isn’t, and name the exclusions specifically, because “use good judgment” is not an exclusion.
  3. Close the loop. Recurring failures trigger CAPA, same as anything else. If you write software, watch code churn against your pre-AI number.

The standards line up behind all of it. ISO/IEC 42001 Clause 9.1 requires you to monitor, measure, analyze, and evaluate, and that’s not satisfied by a policy document. Annex A gets the rest: A.2 for stating where AI is allowed, A.3 for naming who signs off, A.6 for verification before release, A.8 for disclosure. The NIST AI Risk Management Framework covers the same ground through its Measure and Manage functions. If you’re building this out for the first time, our ISO 42001 planning work starts exactly here.

Once you’re measuring, you get thresholds instead of arguments, the same shift that makes it possible to score an AI failure by real-world harm instead of by how alarming it sounds.

If your AI-assisted rework rate sits at or below your historical human baseline, the tool is fine. Run routine controls and stop worrying.

If it’s materially above (rework hours climbing quarter over quarter, churn well past your pre-AI number), you’re outside the range where the tool helps for that task. Tighten the gate or take AI off that work.

If you can’t measure it at all, that’s your finding. You have an unquantified risk in your process, and you’re not meeting Clause 9.1. That was true before anyone said the word slop, and it’ll be true after everyone stops.

🛡️ Responsible AI Note: Verification plus disclosure beats blanket suspicion of AI-assisted work. The goal is proving a human reviewed what mattered before release.

What does the Deloitte report actually prove?

I should finish the thing I started with, because it’s the whole argument in one object.

There’s a report on an Australian government website that cost about four hundred thirty-nine thousand dollars. Two hundred thirty-seven pages, an independent assurance review of the system that automates welfare penalties. It quoted a Federal Court judgment no judge ever wrote. It cited papers that don’t exist, attributed to real professors who were surprised to learn they’d written them. A researcher at the University of Sydney found it, because he went and looked (TechRadar; AI Incident Database).

The fix kept the model in place and changed the process. The revised version corrected the references and added a line disclosing that a GPT-4o tool chain had been used. Verification and disclosure: two controls, applied only after the researcher found the problem, after the refund, after the headlines.

The same two controls, applied before publication, would have cost somebody an afternoon.

That’s the whole thing. The real question is whether anybody read it before you did, and whether you can prove it.

If you want help building that kind of verification into your own AI-assisted work, that’s exactly what Violet Beacon’s Responsible AI Guidelines are built to give you.

Key References

How AI Was Used in This Post

AI helped research the pre-generative-AI defect-rate baselines across journalism, medicine, software, and audit, drafted sections of this post, and verified the named incidents and statistics against primary sources. Kate Waldhauser set the argument, wrote the personal framing, and gave final review before publication. The header image is still pending: see the packaging notes for the image-generation prompt.

Frequently Asked Questions

What does "nobody measured the baseline" mean?
+

It means many teams are arguing about whether AI made work worse without a reliable pre-AI quality benchmark to compare against. If nobody tracked the old defect rate, it's hard to prove how much the new one has actually changed.

What is AI slop?
+

The term is used in at least two incompatible ways. Simon Willison's original 2024 definition is about process: content that's both unrequested and unreviewed. Dictionary definitions adopted in 2025 shifted toward provenance and volume, defining slop as low-quality content produced in bulk by AI. Those two definitions lead to different conclusions about whether the problem is getting worse.

Why isn't the workslop study enough on its own to prove slop is rising?
+

Because it measures how people feel about work they received, with no pre-AI comparison point. It's valuable for showing cost and frustration in 2025, but it can't tell you whether the underlying rate of bad work is higher than it was in 2019, because nobody asked that question in 2019.

What should a regulated manufacturer measure instead of buying an AI detector?
+

Route AI-assisted defects into the quality system you already trust: log fabricated citations or unsupported claims as a nonconformance, track the correction time as rework and cost of poor quality, and send repeat failures to CAPA. Detectors are unreliable enough that output quality and documented review are usually the more defensible control.

What's the practical first step on Monday morning?
+

Pick one AI-assisted workflow, define what counts as a defect in it, name who verifies the output before release, and track correction time for a few cycles. That gives you a real number to manage instead of an argument about vibes.

Explore Related Services

AI Governance
AI Governance Consulting
Learn more →
ISO 42001
ISO 42001 Planning & Consulting
Learn more →
Kate Waldhauser
Founder of Violet Beacon. Responsible AI consultant, ISO 42001 Lead Implementer, and Certified Claris Partner with 20+ years of custom software and database expertise.

Related Posts

← Back to all posts

Want to discuss this topic?

Book a free call to talk about responsible AI, FileMaker, or anything you've read here.