Generative Engine Optimization Tool
How a generative engine optimization tool actually works
Under the hood: how prompts are built, why the same question gives different answers, how a mention becomes a number, and where the cited sources come from.
No site access needed. Nothing to install.
What is a generative engine optimization tool?
A generative engine optimization tool is a measurement instrument. It sends a defined set of questions to generative AI engines, captures the answers those engines return, and extracts two things from each answer: which brands were named, and which sources the answer drew on. Aggregated across hundreds of answers, those two extractions become a visibility rate and a source map.
Everything else such a tool does — scoring, competitor comparison, fix planning — is derived from that raw capture. Understanding the capture is therefore the fastest way to judge whether any given tool's numbers deserve your trust.
This page walks the pipeline in order, and is explicit about the places where methodology choices change the result.
Step one: which engines get asked
Four engines, four different retrieval behaviours. Measuring one and generalising is the field's most common mistake.
ChatGPT
The highest-volume assistant, and often the one a buyer reaches for first. Blends what the model already knows with retrieved sources, so brand presence in training-era content and in current sources both matter.
Perplexity
Built around visible citation. Answers carry their sources openly, which makes it the clearest window into which domains are winning your category.
Gemini
Google's assistant, drawing on Google's index and its own judgement about authority. Frequently the engine where established, well-structured sites do best.
Copilot
Microsoft's assistant, whose citations are fed by the Bing index — which is why Bing indexation is worth checking separately when you are absent here but present elsewhere.
Every prompt in an analysis goes to all four, which is what makes divergence visible. A brand strong on Gemini and absent on Perplexity has a source problem, not a content problem — and you can only see that by comparing.
Step two: building and sampling the prompt set
This is where most of the methodological weight sits.
Prompts are questions, not keywords
People do not type “crm software” into an assistant; they type “what CRM should a 12-person agency use if we already run Slack and Xero?”. A prompt set built from keywords measures a behaviour nobody exhibits. The set has to span problem-shaped questions asked before vendors are known, direct comparisons, alternative-seeking, and constrained recommendations.
Coverage is spread deliberately across the funnel
Concentrating prompts at the bottom of the funnel flatters you, because “is BlueJar any good?” will always mention BlueJar. The informative prompts are the ones that do not contain your name, because those are the ones where an engine chooses who to recommend.
Each prompt is sampled, not asked once
Generative answers are non-deterministic: identical inputs produce different outputs. A single response is an anecdote. Presence is therefore recorded as a rate across repeated runs, and the sample size is what determines whether a change between two audits is signal or noise.
The engines are queried live
Some products infer AI visibility from classic search features instead of querying the assistants. That is a proxy, and it inherits exactly the blind spot the category exists to fix — because rankings and citations correlate only loosely.
Step three: turning answers into a number
Three extractions per answer, and the judgement calls inside each.
Was the brand named?
Harder than string matching. Brand names collide with common words, appear misspelled, and get referred to obliquely. A mention also has to be read in context — being named as the cautionary example is not the same as being recommended, and counting both identically inflates the score.
Who else was named?
The competitor set is captured from the answers rather than supplied by you. That matters, because the competitors engines name are routinely not the ones a company lists in its own deck — and the engines' list is the one buyers are being shown.
What was it built from?
Cited links and referenced domains, pulled per answer and aggregated across the set. This is the input to source-level strategy, and the reason the output can name specific destinations rather than generic advice.
Aggregate the first extraction and you have a visibility matrix — presence per prompt, per engine. Aggregate the second beside it and you have share of voice. Aggregate the third and you have the map of who the engines treat as authoritative in your category.
Step four: the sources that decide the answer
When you aggregate cited domains across hundreds of answers in one category, the distribution is rarely even. A small number of domains turn up again and again — a couple of review platforms, a comparison site, an industry publication, sometimes a single well-placed listicle. Those are the sources doing the deciding.
This is the most actionable output of the whole pipeline, and the least intuitive. The instinct after a bad result is to write more content on your own site. But if eleven answers you were missing from all drew on the same three domains, the highest-leverage work is getting accurately represented on those three — which is a listings, PR and partnerships job more than a publishing one.
- Which domains the engines keep returning to in your category
- Which of them already mention your competitors and not you
- Which are realistically reachable — open to listings, reviews, contributions or corrections
- Which prompts would plausibly change if you appeared there
How it works — FAQs
Why do I get different results if I ask the same question twice?
Generative models sample from a distribution rather than looking up a fixed answer, so identical prompts legitimately produce different wording and sometimes different brand recommendations. This is a property of the engines, not a fault in the tool. It is also the reason visibility has to be reported as a rate over many samples rather than as a yes or no, and why small prompt sets produce unstable numbers.
How is an AI visibility score calculated?
At its core it is the proportion of sampled answers in which your brand is named, broken down by prompt and by engine. Beyond that, implementations differ in ways worth asking about: whether a mention is weighted by prominence in the answer, whether negative or cautionary mentions are separated from recommendations, and whether prompts are weighted by commercial value. Two tools can report different scores for the same brand entirely because of these choices.
How does the tool know a mention refers to my brand?
Matching is contextual rather than literal. Brand names collide with ordinary words, appear misspelled, and get referenced without being spelled out, so the answer text has to be interpreted rather than scanned. The mention is also classified by how the brand was presented, because being cited as a poor fit and being recommended should not contribute to a visibility score in the same direction.
What are kingmaker sources?
Kingmaker sources are the domains that disproportionately shape answers in your category — the handful of sites engines keep drawing on when they decide who to recommend. They are identified by aggregating cited domains across the whole prompt set and looking at which recur. They matter because appearing on one of them can shift several prompts at once, which is usually cheaper than trying to win those prompts through your own content.
Does the tool need access to my site to work?
No. Everything measured is public engine behaviour observed from the outside, so there is nothing to install, verify or connect — you supply a domain and a category. A useful side effect is that you can point it at any domain, including a competitor's, and get an equally valid reading.
Related reading
See the pipeline run on your own category
The free tier runs 60 prompts and lets you read every raw answer it collected.