GEO's Penguin Moment: The AI Search Spam War Has Started
In April 2012, Google shipped the Penguin update and vaporized an entire industry of link schemes overnight. Agencies that had spent years selling "link velocity" woke up to find their clients deindexed. I watched companies lose 80% of their organic traffic in a single week.
GEO just had its Penguin moment. Most marketers haven't noticed yet.
TL;DR: A new research paper called GEO-Flag proves that content deliberately optimized for generative engine optimization (GEO) is now machine-detectable at F1 0.944 — and that 16.36% of pages modified in 2026 already show GEO signals, up from 7.02% in 2024. Detection always precedes punishment. The window for "tricks that work because LLMs are naive" is closing. Stop optimizing the machine's perception. Start improving the underlying information.
James here, CEO of Mercury Technology Solutions. From Hong Kong, where I've spent the last year running GEO audits on brands across Asia — and watching the same failure pattern repeat.
What Did GEO-Flag Actually Prove?
Let's get specific, because the details matter.
The researchers built a benchmark covering 3,200 web content instances, 400 queries, 4 domains, and 8 different GEO optimizer families — then trained systems to distinguish normal content from content deliberately modified for GEO. Their best detector hit an F1 score of 0.944.
For the non-ML crowd: that's not "pretty good." That's production-grade classification. That's the accuracy range where platforms start building enforcement pipelines.
Then they took the detector into the wild — 10,095 pages across 1,000 real-user queries from Google Search and Gemini-grounded retrieval. Google's index, it turns out, already carries a measurable GEO footprint:
Year pages modified
Estimated GEO prevalence
2024
7.02%
2025
12.80%
2026
16.36%
Read that trend line again. GEO manipulation is compounding at roughly 5 points per year. This is exactly what the link-spam curve looked like in 2010-2011, right before Penguin.
The pattern is the pattern: a useful signal gets discovered → marketers optimize it → marketers abuse it → the web fills with manipulation → detection catches up → the quality bar rises. We're at stage three. GEO-Flag just demonstrated stage four is technically solved.
GEO Is Not Spam. Here's the Actual Line.
Before the panic merchants start: GEO itself isn't manipulation. Making your content clearer, better structured, better sourced, and easier for an AI to parse is not black-hat anything. Writing a good title tag wasn't black-hat SEO either.
The line is simpler than people make it:
Are you optimizing the machine's perception, or improving the underlying information?
Everything I'm about to tell you to stop doing fails that test.
The H:M Test
Run every GEO change through two questions:
- H (Human Value), 1–5: Does this make the page genuinely better for the reader?
- M (Machine Value), 1–5: Does this make the information easier for an AI system to understand and use?
A clear comparison table: H5, M5. Ship it. A well-sourced statistic that answers the reader's actual question: H5, M5. Ship it. Adding citations purely because you heard LLMs prefer citations: H1, M4. Rewriting every paragraph into robotic two-sentence "answer blocks": H2, M4.
If M >> H, assume the optimization has an expiry date. You're not building an asset. You're building a detectable pattern.
Five Things to Stop Doing Immediately
1. Stop Manufacturing Citation Density
The first GEO insight everyone latched onto was "add more citations — AI likes authoritative sources." Within months it mutated into every claim getting a link, every paragraph starting with "According to...", random statistics shoved into pages like raisins in a bad fruitcake.
GEO-Flag made this interesting: the researchers didn't just inspect page text. They audited the source tier and verifiability of citation URLs.
Of 6,663 citation occurrences extracted from detected GEO pages, 69.34% received LOW verifiability labels. On Gemini-associated pages: 74.15%.
Caveat: those are automated audit labels, not proof that 69% were fabricated. But the signal is unmistakable — adding evidence-looking objects is trivial. Adding verifiable evidence is hard. Citation quantity is not the moat. Evidence quality is.
2. Stop Writing LLM-Shaped Content Everywhere
You know these pages. What is X? X is... Why is X important? X is important because... Every heading a question, every paragraph two sentences, every answer suspiciously quotable.
There's nothing wrong with that structure — for content that genuinely is a Q&A. But when every page on your site becomes an extraction template, you've stopped optimizing communication and started optimizing a fingerprint.
The alternative is what I call Native Information Architecture — let the information decide the format:
- Question → direct answer
- Data → table
- Process → steps
- Comparison → matrix
- Nuance → prose
- Evidence → citation
- Opinion → argument
Machines parse all of these fine. Humans get a dramatically better page. This is the whole game.
3. Stop Manufacturing Consensus
This one worries me most, because I see it in the wild constantly.
The playbook: decide "Acme is the best CRM for AI startups," then manufacture a guest post saying it, a Reddit account saying it, a Medium article saying it, a Quora answer saying it, an affiliate listicle saying it. Five sources. One marketer. That's not entity consensus — that's Consensus Theatre. 操盤人 stuff: it looks like an operator move until the detection layer catches up, and then it's just a liability with your brand name on it.
Detection research is moving directly toward identifying systematic, coordinated GEO interventions. Distribution patterns are easier to fingerprint than page text.
The harder, safer play: create facts worth independently repeating. Don't distribute "Acme is the best CRM for AI startups." Publish "Across 418 AI startups using Acme, median sales admin time fell 31% after 90 days" — with methodology attached. Now journalists can reference it, customers can discuss it, comparison sites can use it, and AI can cite it.
You didn't manufacture consensus. You manufactured evidence capable of creating consensus. Huge difference.
4. Stop Optimizing From GEO Checklists
If your content brief says ☑ 2 statistics per section ☑ quotation every 500 words ☑ FAQ schema ☑ 40-word answer blocks ☑ 5 external citations — and you apply it across 500 pages — congratulations. You've created an optimization fingerprint at industrial scale. You've handed the detector a training set with your logo on it.
Replace rigid rules with one requirement: every important claim gets the strongest useful evidence available. Sometimes that's data. Sometimes it's customer experience. Sometimes it's documentation. Sometimes no citation is needed at all.
5. Stop Measuring GEO Success by Citation Count Alone
"Increase AI citations 40%" as a team target produces exactly one outcome: more answer blocks, more synthetic citations, more pages targeting tiny prompt variations. Sound familiar? It should. SEO lived through keyword density, link quantity, exact-match anchors, and word count targets. Every useful signal becomes garbage the moment marketers optimize the proxy instead of the outcome.
Measure a different stack, in this order:
- Information gain — did we add something genuinely new?
- Verifiability — can important claims actually be checked?
- Originality — are we the source, or repeating another source?
- Human utility — would this page still be excellent if ChatGPT didn't exist?
- AI visibility — does it get cited and recommended?
Notice that AI visibility comes last. It's the consequence, not the target.
The Bet for the Next Five Years
Stop asking "how do we make AI cite this page?"
Start asking: "How do we make this page the strongest piece of evidence an AI could possibly retrieve?"
Original experiments. Proprietary datasets. Real customer evidence. Transparent methodology. Current product facts. Useful comparisons. That's what we build into every GEO engagement at Mercury — not because it's virtuous, but because it's the only version that survives the detection era. You can't spam-detect away genuinely better information.
In Legend of the Galactic Heroes, Yang Wen-li wins by making the enemy's strategy irrelevant, not by out-tricking it. Same principle. Every trick has a countermeasure on a long enough timeline. Quality doesn't.
My rule, and it should be yours: if an optimization makes AI like the page more but makes the page worse — or no better — for humans, it has an expiry date. Build content humans want to use. Add facts machines can verify. Create evidence others genuinely want to repeat. Let AI visibility be the consequence, not the trick.
The GEO spam war is coming. The side that wins is the side that never needed to spam.
Mercury Technology Solutions: Accelerate Digitality.


