The number of agencies offering GEO grew several times over in under a year. Most did not build a new capability; they added a section to the proposal. From the buying side those two situations look identical, and the difference shows up around month four.
Some of the overlap is legitimate. Technical hygiene, structured content, and crawlability genuinely matter for both, and an agency doing that work is not being dishonest by charging for it.
The problem is what gets left out. A programme built entirely from SEO deliverables addresses the smaller half of what determines whether a model recommends you, and the reporting will not reveal the gap.
Key Takeaways
- Check the scope for earned work: a programme with no third-party component is an SEO programme.
- Ask how the prompt set was built: a generic list measures someone else's buyers.
- Insist on a baseline before work starts: without one, no later claim is checkable.
- Treat precise percentages as a warning: the method rarely supports the decimal.
- Timelines in weeks are a red flag: the mechanisms involved move in quarters.
What Legitimately Overlaps
Start here, because a fair test has to separate genuine overlap from repackaging.
Crawler access, structured markup, clean semantic HTML, page speed, and extractable content structure all matter for AI retrieval and are all standard SEO work. An agency doing them well is doing something useful, and a proposal containing them is not automatically suspect.
The technical layer is a prerequisite. Blocking AI crawlers removes you from consideration entirely, which is why crawler configuration belongs in any credible scope. What it is not is the whole programme.

The Missing Half
Analysis of 25 million cited links found earned media accounts for the large majority of AI citations, while paid and advertorial placements account for a fraction of a percent. A programme working only on your own domain is optimising the smaller input.
Three components separate a GEO programme from an SEO one with new labels.
Third-party presence. Editorial coverage, review platform positioning, and community presence. This is the work that produces corroboration, and it is slower and harder to systematise than content production, which is exactly why it gets omitted.
Source-level correction. Identifying which specific external sources produce the answers about you, and fixing or displacing the ones that are wrong. This requires looking at outputs rather than at your site.
Entity consistency across sources. Making your category, audience, and positioning description match everywhere it appears, not only on pages you control.
If none of those appear in a scope of work, the programme cannot move the inputs that carry most of the weight, however well it executes everything else.
Questions That Separate the Two
Five questions, and what the answers tell you.
How was the prompt set built?
A credible answer describes pulling phrasings from your sales calls, support tickets, and buyer research. A generic list of category questions measures a hypothetical buyer rather than yours, and human phrasings of the same intent diverge far more than most teams assume.
What did the baseline show?
Any programme should start by recording what models currently say, across platforms, with sources. If work began without one, no subsequent improvement claim can be checked, and the method that makes this credible is set out in tracking AI search visibility.

Which sources are producing our answers now?
This is the sharpest single question. An agency that has done the diagnostic can name specific pages, threads, and profiles. One who cannot has looked at your site rather than at the outputs.
What is in scope that we cannot do ourselves?
Content production is something most companies can do in-house. Earned coverage, source correction, and multi-platform positioning generally are not. If the entire scope is work your team could run internally, the pricing should reflect that.
What would make you tell us to stop?
A serious answer exists: positioning still changing, no third-party footprint to build on, a category citation set locked by incumbents, or a horizon shorter than two quarters. An agency that cannot describe the conditions under which this is a bad investment has not thought about it carefully.
Signals in the Reporting
| What you see | What it usually means |
|---|---|
| Only owned-content metrics | No third-party work in scope |
| A single visibility percentage | Method not stated; likely not segmented |
| Precise figures from small samples | Implied precision the data cannot carry |
| Citation counts without sources | Outputs not actually being read |
| No mention of description accuracy | Presence measured, framing ignored |
| Improvement claimed within weeks | Either technical-only, or noise |
The last row deserves particular attention. Technical fixes can produce movement quickly, and that movement is real but limited. Corroboration-driven change runs on the timeline of third-party publishing, which is in quarters.
The False Precision Problem
Reported figures in this field frequently carry more decimal places than their methods support, and buyers rarely ask what sits behind them.
The questions that resolve it are simple: how many prompts, run how many times, on which platforms, from which locations, and on what dates. A visibility figure without those is not verifiable and often not reproducible, which is a recurring pattern in the statistics the industry cites most.

This is not a reason to avoid measurement. It is a reason to expect ranges, trends, and stated methods rather than a headline number that moves conveniently in one direction.
What Honest Uncertainty Sounds Like
Counterintuitively, the clearest signal of a serious programme is how comfortably it discusses what it cannot prove.
Attribution here is genuinely hard. There is no rank to observe and no way to split-test a retrieval pipeline. Model updates, competitor activity, and seasonal shifts in phrasing all move the same numbers your work moves.
An agency that acknowledges this and describes contribution rather than causation is being accurate. One that presents clean causal claims is either not measuring carefully or not telling you what the measurement can support.
Where the Real Difficulty Sits
A fair note on the other side, because this is not simply a matter of buyer vigilance.
Third-party work is genuinely harder to sell and deliver than content production. It has less predictable output, no guaranteed monthly deliverable, and outcomes that depend partly on editors and communities. Content calendars exist because they are legible, schedulable, and reportable.
That is why programmes drift toward publishing even when the people running them know better. Recognising the pressure makes the diagnosis fairer, and it does not change what the buyer should be checking. The distance between doing visible work and moving the outcome is what our AI Recommendation Gap report measures.

A Short Checklist
Before signing, confirm the following appear somewhere in the scope or the answers.
- A baseline measurement performed before work begins.
- A prompt set built from your buyers rather than a template.
- Named sources currently producing your answers.
- Third-party work that is not content published on your own domain.
- Reporting that includes description accuracy, not only presence.
- A stated method behind any figure.
- A timeline measured in quarters.
- Conditions under which they would advise against proceeding.
An agency missing two or three of these may still be worth working with. One missing most of them is selling something else under a current name.
Run the Checklist on Us
Publishing this only works if we pass it, so the fair thing is to invite the test rather than claim the result.
Ask us which sources are producing your current answers. Ask what our baseline showed and what method sits behind any figure we quote. Ask what is in the scope that your team could not run internally, and ask under what conditions we would tell you to wait. Those are the questions above, and they are the ones worth an hour of your time whoever you end up hiring.
If you want the diagnostic itself rather than a pitch, that is what Lureon starts with: a baseline across platforms, the named sources behind your answers, and an honest read on whether the gap is worth closing now. Our crypto payroll case study covers what followed from one of those: +288% organic click growth and +575% ChatGPT session growth across 100+ countries over twelve months.

FAQs
1. What is the difference between GEO and SEO?
They overlap on technical foundations such as crawlability, structured markup, and extractable content. They diverge on what drives the outcome: SEO is largely won on your own domain, while AI citations are driven predominantly by third-party sources, which requires earned coverage and source-level correction rather than publishing volume.
2. How can I tell if an agency is just relabelling SEO?
Check whether the scope contains work you could not do in-house. If every deliverable is content, schema, and on-site optimisation, with no third-party component and no source-level diagnostic, it is an SEO programme with a new name.
3. How long should GEO results take?
Technical improvements can show movement within weeks, but corroboration-driven change runs on third-party publishing timelines, meaning two to three quarters for readable results and longer in regulated categories or against entrenched incumbents.
4. Should I trust a specific AI visibility percentage?
Only alongside its method. Ask how many prompts were run, how often, on which platforms, from which locations, and on what dates. Figures presented without that context are frequently not reproducible, and precise decimals often come from samples too small to support them.
5. What is a reasonable thing for an agency to guarantee?
Process rather than outcome. Guaranteed citations or fixed visibility percentages are not deliverable, since outputs are non-deterministic and influenced by model updates outside anyone's control. Baseline measurement, defined deliverables, and reporting cadence are reasonable commitments.
This describes evaluation criteria rather than a measured study. Programme scope appropriately varies by category, starting position, and budget, and the absence of any single element is not by itself disqualifying.