Skip to main content
geo

How to Run a Citation Study Without a Research Team

Industry benchmarks describe an average your category isn't. A practical method for running a defensible AI citation study in a day, with a spreadsheet.

ยท By Veljko Plavsic ยท 7 min read

Every GEO statistic you have read came from someone running prompts and counting results. The method is not complicated and the barrier is not technical. Most of the widely-cited findings in this field rest on a few hundred queries, a spreadsheet, and a clear question.

Industry benchmarks describe an average. Your category, your buyers, and your competitors are not the average.

Running your own study takes about a day and produces something no benchmark can: a measurement of the thing you actually care about, with a method you can defend and repeat.

Key Takeaways

  • A useful study needs one clear question: broad exploration produces data you cannot interpret.
  • Repetition matters more than volume: three runs of forty prompts beats one run of a hundred.
  • Record sources and full answers: counts alone discard the most valuable output.
  • Fix the variables you are not testing: platform, location, date, and phrasing all move results.
  • Publish the method alongside the finding: it is what makes the result citable by others.

Start With One Question You Can Answer

The most common failure is starting with a topic rather than a question. "How visible are we in AI search" produces a number with no comparison and no conclusion. "Do citations change when the same question is phrased ten different ways" produces a finding.

Good questions share a shape. They compare two conditions, they have a plausible answer in either direction, and the result changes what you would do next. That last criterion rules out most of what teams instinctively want to measure, and it is worth applying strictly given how easily AI visibility figures get over-read.

How to Track AI Search Visibility Across ChatGPT, Perplexity, Gemini, and Claude in 2026

Questions worth testing

  • Phrasing sensitivity. How much does the cited source set change across ten phrasings of one buying question?
  • Platform divergence. How many cited sources are shared between two engines for identical prompts?
  • Geographic variance. Does the same prompt from two countries return different brands?
  • Source age. How old is the median cited source for questions about your category?
  • Description drift. Does the language used about your brand differ by query type?

Design the Sample

A defensible study fixes everything except the variable under test. Decide in advance which one thing varies, then hold platform, location, date range, and account state constant.

How many prompts

Fewer than you would guess; run more times than you would guess. Forty prompts run three times produces more interpretable data than a hundred and twenty run once, because AI outputs vary between runs and single results cannot be distinguished from noise.

A practical baseline for a one-day study: ten to fifteen prompts, three to four platforms, three runs each. That is 90 to 180 queries, which is a few hours of work and enough to see whether an effect exists.

Where prompts should come from

Pull phrasings from sales call recordings, support tickets, and search query data rather than writing them yourself. Prompts written by marketers reliably sound like marketing, and buyers describe problems in their own vocabulary.

This is the step most studies skip and the one that determines whether the finding generalises to real behaviour.

What to Record

FieldWhy it matters
Exact prompt textReproducibility; small wording changes move results
Platform and modelFindings rarely generalise across engines
Date and timeModel updates make undated results uninterpretable
LocationGeography shifts results independently of language
Full answer textThe qualitative finding is usually the valuable one
Cited sources and datesEnables source-age and source-type analysis
Brands named, in orderPosition and presence are separate measures

Storing full answers costs nothing and is what separates a study you can revisit from one you have to re-run. Counts answer one question; the text answers questions you have not thought of yet.

Analyse Honestly

The analysis step is where most self-run studies overreach. Three disciplines keep a finding defensible.

Report the spread, not just the average

If a brand appears in 60% of runs, say so, and say how that varied across prompts. An average that hides a range from 10% to 95% is worse than no number, because it invites a conclusion the data does not support.

Name your confounders before someone else does

List what else could explain the result. A model update mid-study, a competitor's launch, a seasonal shift in phrasing. Stating these makes the finding more credible, not less, and it is the standard the field mostly fails to meet, as we covered in our audit of the GEO statistics everyone cites.

The GEO Statistics Everyone Cites, Checked

Distinguish what you measured from what you infer

You measured citation frequency across a prompt set. You did not measure buyer exposure, and you cannot claim causation from an observational design. Saying so plainly costs a sentence and protects the entire result.

Turn It Into Something Others Cite

A study becomes an asset when other people reference it, and that depends on presentation more than scale.

Lead with the number

Put the single most surprising figure in the first line, with its scope attached. "Across 180 queries in four engines, only 14% of cited sources were shared between platforms" is citable. A paragraph of context before the number is not.

Publish the method in full

Prompt count, platforms, dates, locations, run frequency, and how you classified results. This is what allows a journalist or another practitioner to cite you with confidence, and it is why earned coverage follows original data more reliably than it follows opinion.

State the limitations yourself

One model, one phrasing family, one vertical, one point in time. Naming the boundaries is what makes the claim inside them credible, and it pre-empts the obvious objection.

Make the finding structurally quotable

A self-contained sentence carrying the number, the scope, and the conclusion travels further than the same information spread across a paragraph. Since platforms draw on different source sets, a finding that is easy to extract has more chances to be picked up.

Which Domains AI Engines Trust Most: 86% of Top Sources Are Not Shared Across Platforms

Repeat It

A single study is a snapshot. The same study run quarterly, with the same prompt set and method, becomes a time series, and time series are considerably more valuable than one-off findings.

It also solves the measurement problem that industry benchmarks cannot. You stop asking whether your visibility is good in absolute terms and start asking whether it moved, on a method you control. That is a more useful input than any benchmark, and it complements tooling rather than replacing it.

Why This Is Worth a Day

Original data is the one content asset in this category that competitors cannot replicate and that compounds rather than decays. Analysis of 25 million cited links found earned media accounts for the large majority of AI citations, and original research is the most reliable way to earn it.

Everything else in a content programme is an interpretation of someone else's numbers. A study, however small, produces a number that belongs to you.

What the Study Will Not Tell You

Here is what usually happens on the first run.

You open the spreadsheet expecting a visibility number. What you find instead is a competitor you had not thought about appearing in half the answers. Or your own brand described in language your positioning abandoned two years ago. Or a single forum thread showing up as the source behind four different questions.

The study is very good at surfacing that. It is not designed to tell you what to do about it, and that is where most self-run studies stop. The finding sits in a slide, everyone agrees it is concerning, and nothing moves.

Closing a citation gap is a different discipline from measuring one. It means working out which sources are producing the answer, whether they can be corrected, displaced, or only diluted, and which of those is worth the quarter it will take. That work is slower than the study and considerably less satisfying, which is why it is usually the part that gets skipped.

If You Would Rather Skip to the Answer

Running the study yourself is genuinely worth a day, and we would rather you did it than took anyone's benchmark on faith.

But if the result comes back messier than expected, or you would rather start from a diagnosis than build one, that is what we do. Lureon runs this measurement across engines and markets, then works the gap it exposes: entity clarity, third-party presence, and the source-level work that actually shifts what models say.

Our crypto payroll case study shows what that looked like over twelve months: +288% organic click growth and +575% ChatGPT session growth across 100+ countries.


FAQs

1. How many prompts do I need for a useful AI citation study?

Ten to fifteen prompts run three times across three or four platforms is enough to detect whether an effect exists, producing 90 to 180 queries. Repetition matters more than prompt count, because AI outputs vary between runs and single results cannot be separated from noise.

2. Do I need special tools to run one?

No. A spreadsheet and manual runs are sufficient for a first study. Visibility tools make repetition easier at scale, but the method does not depend on them and the analysis discipline matters more than the tooling.

3. What should I record for each query?

The exact prompt, platform, date, location, full answer text, cited sources with their publication dates, and the brands named in order. Storing full answers is what lets you revisit the data for questions you had not considered.

4. How do I make my study credible?

Publish the complete method, report the spread rather than only the average, name plausible confounders, and state the limitations yourself. Explicit boundaries make the claim inside them more defensible, not less.

5. How often should I repeat the study?

Quarterly, using an identical prompt set and method. A single study is a snapshot; the same study repeated becomes a time series that shows direction, which is more useful for decisions than any absolute figure.


This describes a practical method for small-scale observational studies. Observational designs show association rather than causation, and results are specific to the prompts, platforms, locations, and dates tested.

Updated on Aug 11, 2026