Article's Content
Welcome back to What Matters This Week.
Last week we dug into Writesonic’s research on how quickly citations are won and lost and how precarious AI visibility really is. This week, the story shows up one level deeper, inside Google’s own research, and it points at what separates the content that holds its place from the content that gets flushed out.
Google Published How It Catches AI Content at Scale. The Data Shows What Happens When It Does.
Here’s the TL;DR
- Google researchers published a paper describing S-CTS (Scalable Cluster Termination System), a system that catches AI spam by finding coordinated networks of accounts rather than grading content one piece at a time. Over six months, it terminated 50,000 clusters covering 130,000 channels.
- The paper is about video, but it names Sentence-BERT as a way to catch AI-generated text, confirming that scaled AI writing leaves a detectable mathematical footprint.
- You don’t have to take the theory on faith. Lily Ray’s analysis of 220+ sites running scaled AI content found more than half lost 30% or more of their peak organic traffic. The mechanism and the consequences now line up.
What’s happening
The most important shift in how Google fights AI spam is where the judgment happens. It’s moving from the individual page to the network of accounts behind it. A new Google research paper spells out how.
The paper, Scalable Detection of Adversarial Synthetic Slop and Coordinated Media Abuse, describes a live system deployed at a major online video platform (you can probably guess the one), built to catch the flood of AI spam that overwhelms traditional quality filters.
Older approaches grade content one piece at a time, asking whether a single page or video looks low quality. That breaks down when spammers use AI to generate endless unique variations of what the paper calls “functionally identical” content. Each piece looks a little different, so each piece slips under the filter.
This is where Google’s Scalable Cluster Termination System (S-CTS) comes into play.
Instead of judging content in isolation, it looks for the fingerprint of coordinated production: networks of accounts that publish on the same automated cadence, draw from the same templates, and share the same infrastructure signals. Two checks have to line up before anything happens:
- A bot-net detector decides whether accounts form a coordinated cluster.
- A content classifier scores whether the material is synthetic.
Only content that is both part of a cluster and flagged as synthetic gets pushed toward enforcement.

The results Google reports are not small. Across six months, the system terminated 50,000 clusters made up of 130,000 channels, while cutting the human review load by roughly half.
The paper isn’t only about video — it describes text-based detection too. To catch text, the paper points to Sentence-BERT, which converts sentences into embeddings and measures how semantically similar they are. Scripted, templated AI writing clusters together mathematically, which gives Google a way to spot it by its footprint rather than by reading it line by line.
Why it matters
- Detection moved from the page to the network.
For a long time the game was to make each page look good enough to pass. This research describes a system that doesn’t care how good any single page looks. It’s hunting the pattern across accounts: same template, same cadence, same origin. That’s the same move Reddit’s communities made when they stopped flagging individual posts and started tracking brands across subreddits. When the unit of judgment becomes the network instead of the page, “just make it slightly different each time” stops working, because the similarity is the signal.
- This is already happening in web search, and there’s data on it.
The video paper tells you how cluster-level detection works. Lily Ray’s recent analysis tells you what it looks like when scaled content meets Google’s ranking systems in the wild.
Ray monitored more than 220 sites publicly identified as customers of AI content and automation platforms. The pattern she documents is consistent: rapid growth in pages over six to twelve months, an organic traffic peak, then a steep drop that erases most of the gain and often falls below where the site started. Across the group, 54% lost 30% or more of their peak organic traffic, 39% lost 50% or more, and 22% lost 75% or more. Glenn Gabe calls the shape “Mount AI”: a sharp climb followed by a matching collapse once Google’s systems gather enough signal.
- The moat is the stuff that can’t be mass-produced.
AI tools aren’t the problem. It’s how we implement them that earns the ire of the algorithms. Used for research, briefs, synthesis, and pulling proprietary data into the workflow with a human expert in the loop, they’re genuinely useful. The trouble starts when the goal becomes volume and the people closest to the content stop reviewing what ships.
Google’s system is built around the same distinction. The cluster requirement exists so the platform targets coordinated, mass-produced behavior and leaves individual creators experimenting with AI alone. Original research, named experts, first-party data, a real point of view: these are the literal opposite of a detectable cluster. They’re hard to fake, hard to scale, and hard to replicate with the same prompt a competitor is using. Every time a platform gets pickier, that kind of content gains ground.
Google has even given this a name. Its guidance points to “non-commodity content,” the unique, specific, and firsthand material that AI can’t mass-produce, as the type that wins visibility now. That’s the same conclusion from the other direction: the content that survives detection and the content Google actively rewards are one and the same.
What to do about it
- Run Google’s Helpful Content questions on your own subfolder. Could a competitor publish a near-identical version of this page tomorrow using the same prompt? Is there any first-party data, expertise, or original perspective here that isn’t already sitting in the top ten results? If the answer is no, the page is a liability waiting for an update.
- Reframe AI internally as the assistant, not the author. Research, outlines, briefs, data synthesis, faster workflows, all fine. Autopublishing at scale is the part that backfires. Put a human expert on the output before it ships.
- Watch the pattern move from video to web. Treat this as a reasonable bet, not a confirmed fact: the detection logic Google is publishing for video is the logic that eventually reaches text. Building real authority now is cheaper than digging out of a decline later.
Go Deeper on AI Spam and Scaled Content Abuse
→ Scalable Detection of Adversarial Synthetic Slop and Coordinated Media Abuse — The primary source. Google’s full research paper on the S-CTS system, the two-stage LLM classifier, and the LoRA/APO approach that lets them adapt to new generative models fast. Google Research
→ It Works Until It Doesn’t: AI Content Strategies That Backfire — Lily Ray’s analysis of 220+ sites that scaled AI content, the eight risky templates, and the boom-bust trajectory that keeps repeating. The evidence layer under everything in this issue. Lily Ray
→ Understanding Non-Commodity Content (With B2B Examples) — Google’s own term for the content that wins AI search: unique, specific, authentic. We break down what that means with examples from 6sense, Clarify, and Beehiiv. The build-side companion to this issue’s “don’t get caught” warning. Foundation Labs
That’s it for this week.
If something landed, tell us. If something felt off, tell us that too. Reply to this email or DM me on LinkedIn.
Have a great weekend,
Ethan Crump