TL;DR: The fastest way to vet an AI SEO agency is not "Do you offer GEO?" It's three questions: How do you measure my AI visibility today? What is the baseline? What exactly will you improve next? An agency that can't answer those in the pitch can't prove it caused anything later. Below: the seven questions that separate 操盤人 (real operators) from slideware, and two proof points — a citation screenshot and a 20-tool stack — that tell you almost nothing.
I am Akira, GEO Specialist of Mercury Technology Solutions, writing from Hong Kong — a market where AI SEO, GEO, and AEO get pitched in the same meeting as if they were the same product. They are not. And in 2026, the agency selection problem has inverted: everyone claims the capability, almost nobody claims the measurement. And because this budget usually sits inside a broader digital transformation programme, the measurement question is the one your CFO will eventually ask.
Why AI SEO Agencies Are Suddenly Impossible to Compare
Traditional search engine optimization (SEO) had a shared scoreboard: impressions, rankings, clicks, conversions. AI search shattered it. Two agencies can both report a "35% AI Visibility Score" and be measuring completely different things — different platforms, prompt sets, markets, languages, testing windows, different definitions of what even counts as a citation versus a mention.
A score is only comparable when the methodology underneath it is comparable. So never ask "What's the score?" Ask "How was this score measured?" Google's own 2026 guidance on third-party SEO tools says it plainly: no external tool has Google's internal ranking data, and proprietary scores are interpretations, not gospel. Methodology is the product. Everything else is packaging.
The 7 Questions, in the Order Worth Asking
1. What Is Our AI Visibility Baseline Today — and How Do You Measure It?
The first question is never "How much can you raise it?" It's "Where are we now?"
A real baseline names its conditions: which prompts, which AI platforms, which market, which language (English vs Traditional Chinese is not a footnote in Hong Kong), what dates, how many runs per prompt, and what counts as a citation versus a brand mention. A confident answer sounds like: "Measured 1–7 August, 40 fixed commercial prompts, Hong Kong market, each platform separately, three runs per prompt, citations and mentions tracked independently."
A concerning answer sounds like: "We'll figure out measurement after you sign." Run.
This is exactly why our engagements start with a GEO Audit — a 48-hour baseline, not a retainer pitch. If you want the thinner, free cut first, run the GEO Scorecard and see how AI cites you today.
2. Which AI Platforms Do You Actually Work On — and What's Different About Them?
ChatGPT, Perplexity, Gemini, Claude, and Google's AI Overviews are not one surface. They don't even retrieve the web the same way:
- ChatGPT — OAI-SearchBot surfaces sites in ChatGPT Search; GPTBot is training crawling (a separate policy decision); ChatGPT-User is user-triggered fetches.
- Perplexity — PerplexityBot builds its own index; Perplexity-User fetches on request and largely ignores robots.txt.
- Claude — ClaudeBot (training), Claude-User (user fetches), Claude-SearchBot (search retrieval) are three different gates.
- AI Overviews / AI Mode — features inside Google Search, riding Googlebot and Google's index. Google-Extended does not control them; nosnippet and friends do.
If an agency says "ChatGPT, Gemini, and Perplexity — it's all GEO, the approach is basically the same," you're listening to a horoscope, not a methodology. Ask which parts of their claim come from platform documentation and which from their own testing.
3. Show Me One Before-and-After Content Rewrite — and Explain Every Change
Not a case-study PDF. One page. Before → After → Why.
Legitimate reasons: the copy didn't answer the question directly, claims had no sources, conditions were missing, one paragraph mixed three intents, the heading promised what the paragraph didn't deliver. Illegitimate reason: "We added more AI keywords so LLMs will like it." That's not optimization; that's seasoning a steak with salt you found on the floor.
Google's 2026 generative-AI guidance doesn't require a special "AI writing format." So the strong question is: what did you change, what problem did it solve, and how would you test whether it worked?
4. What Original Research Will You Help Us Produce?
If every company's GEO plan is "What Is GEO?", "5 Benefits of AI SEO," and "10 ChatGPT Tips," everyone converges on the same commodity content — and commodity content is what AI engines summarize without citing anyone.
The agency's job in any serious AI SEO or GEO engagement is to help you manufacture evidence competitors can't copy: first-party data, small market studies, benchmarks with stated methodology, expert knowledge turned into Answer Assets. First-party data doesn't guarantee a citation — it guarantees you're not rewording the same search-results page as everyone else. This is the core of what we call LLM-SEO: citations are won with evidence, not adjectives.
5. How Do You Handle English and Traditional Chinese Separately?
In Hong Kong this question eliminates half the market. The failure mode is: English article → AI translation → publish. A bilingual agency should explain URL structure, separate keyword research per language, hreflang reciprocity, same-language canonicals — and, more importantly, whether Hong Kong Traditional Chinese searchers use terms that aren't translations of English at all. In our own SERP sampling, Chinese-language queries triggered AI Overviews at a visibly higher rate than English ones — topic-confounded, so we treat it as an observation, not a law. The point: an agency that can't discuss the difference isn't measuring it.
6. What Will Reporting Show in Months One Through Six?
Refuse "SEO takes six months, let's talk then." Leading indicators exist from week one: crawler access status, indexing fixes, baseline citation and mention counts, prompt coverage, completed content changes — and by now, first-party data from Google Search Console's generative AI reporting and Bing Webmaster Tools' AI Performance reports. None of these metrics are interchangeable — a Bing citation is not a Google AI Overview impression — and a serious report says so. If you want to know how the KPIs fit together, that's the LLM KPI framework conversation.
7. If AI Overviews Change Again in Six Months, What Happens to This Work?
AI search is young. Any agency presenting a fixed playbook is describing a surface that already moved. Durable SEO and GEO work — crawlability, indexability, genuinely useful content, original evidence, clean entity information, repeatable measurement — survives interface changes. Tactics that exploit one temporary platform quirk do not.
My favorite kill-shot question: "What has your agency changed its mind about in the last six months?" An agency that has never revised a benchmark hasn't been consistent. It hasn't been measuring.
Two Proof Points That Prove Almost Nothing
The single citation screenshot. It proves one appearance, one prompt, one moment. It does not prove consistency, share of voice, or revenue. A screenshot is an occurrence, not a measurement system.
The 20-tool stack. A tool list tells you the agency's software budget, not its methodology. One tool run under a written, repeatable protocol beats twenty tools used ad hoc. Every time.
The One-Line Test
Stop asking "Do you do GEO?" Start asking "Show me the baseline protocol." Whether the proposal says AI SEO, GEO, or LLM SEO, the protocol is the product. If you're evaluating agencies in this market right now, our best GEO agency in Hong Kong comparison page lays out the criteria publicly — including the ones we'd lose on. And if you want proof rather than pitch, the Hong Kong GEO + AIO case studies show what measured citation work looks like in a bilingual market.
Measure first. Sign second. That's what a real AI SEO and GEO engagement looks like.
Mercury Technology Solutions: Accelerate Digitality.
