Guide
Retrieval vs training data: why on-page SEO alone won't get you cited by AI
An AI answer comes from two different places: retrieval (the model searches the live web and cites sources) and training data (what the model already knows without searching). On-page SEO mostly moves the retrieval side. It does almost nothing for the training-data side, which is an entity-presence problem built over time through mentions, reviews, and consensus across independent sources. That is why "GEO is just SEO" is half right and half wrong.
Spend ten minutes in any SEO or GEO community and you will find the same argument going in circles: is generative engine optimisation actually a new discipline, or is it just SEO with a fresh coat of paint? Both camps sound convincing, and both are partly right. The reason the debate never resolves is that people are conflating two completely different mechanisms behind an AI answer.
Once you separate them, the argument mostly dissolves - and, more usefully, you can tell which mechanism is failing you and stop wasting effort on the wrong fix.
The two sources of an AI answer
When ChatGPT, Perplexity, Gemini, or Google's AI Overviews produce a recommendation, the content in that answer comes from one of two places (and sometimes a blend of both).
Source 1
Retrieval - the model searches the live web
The model fires off a live search, pulls back current pages, and cites sources it just read. This is mostly classic SEO territory. If your pages are crawlable, clearly written, and corroborated by third parties who also rank, you have a genuine shot at being retrieved and cited. Being cited elsewhere on the web matters as much as your own page here - the model is looking for consensus, not just a single claim.
Source 2
Training data - what the model already knows
The model answers from knowledge baked in during training, with no live search at all. If the model has simply never encountered your brand across the text it learned from, it will not go looking for you - and no amount of on-page tweaking changes that. This is an entity-presence and reputation problem, accumulated slowly through mentions, reviews, directories, and broad agreement across many independent sources.
Here is the uncomfortable part for anyone selling "GEO is just SEO": on-page work moves Source 1. It barely touches Source 2. You can have a technically flawless page and still be absent from the model's knowledge, because that knowledge was formed from how often and how consistently the wider web talked about you - long before you optimised a single heading.
Why the distinction actually matters
These two sources respond to different work, on different timescales, so treating them as one thing leads you to the wrong remedy.
- The retrieval side is responsive. Fix crawlability, publish answer-ready content, earn a few corroborating citations, and you can see movement in weeks. This is where SEO fundamentals earn their keep, and where the "GEO is just SEO" crowd have a real point.
- The training-data side is slow and cumulative. It is built the way a reputation is built: lots of independent sources describing you consistently over a long stretch. You cannot on-page your way into it. You earn your way in - mentions, reviews, coverage, being referenced by others - and it lags by months or longer.
So when someone says on-page SEO "did nothing" for their AI visibility, they may be completely right - for the training-data layer. And when someone else swears SEO fixes moved their citations, they may also be right - for the retrieval layer. Both experiences are real. They are just describing different machines.
Not every prompt even triggers a search
This is the detail that catches people out. It is tempting to assume the model always searches the live web, so good SEO always has a chance to win. It does not. Plenty of prompts are answered purely from training data with no retrieval at all - especially broad, category-level questions the model feels it already "knows" the answer to.
If the prompts your buyers actually type tend to be answered without a search, then your crawlable, well-optimised page never enters the running. The model is drawing entirely on what it learned in training - and if you were not part of that, on-page work cannot rescue you on that prompt. This is one of the reasons AI answers and Google rankings so often fail to line up, which we cover in why ChatGPT recommends you but Gemini doesn't.
How to check which mechanism is failing you
You do not have to guess. There is a simple, repeatable test that separates the two layers.
The test
Ask with web browsing turned off
Ask a model about your brand and your competitors with browsing or search disabled. Now the model can only draw on training data, so the answer reveals the knowledge layer directly. If it describes your competitors in detail and has never heard of you, that is a training-data gap - not something a meta description will fix. If it knows you but gets you wrong, that is an entity-clarity signal worth acting on.
The useful property here is stability. The training-data layer does not drift day to day the way live retrieval does. It stays broadly the same for months, until the model's knowledge cutoff moves and a newer version learns from more recent web text. That stability is what makes it measurable - you can take a reading now, act, and take another reading later without random noise swamping the result. It contrasts sharply with retrieval-based answers, which can change run to run for reasons that have nothing to do with you. We go deeper on separating signal from noise in how to measure AI visibility without fooling yourself.
Run both readings - browsing on and browsing off - and the picture sharpens. Browsing on tells you whether your retrieval foundation is working. Browsing off tells you whether you exist in the model's knowledge at all. A weak result in each points to a different fix.
What to actually do about it
The honest answer is that you work both layers, because they are not substitutes for each other.
- Fix the retrieval foundation. This is SEO fundamentals, and they still matter: let AI crawlers in, publish clear answer-ready content and FAQs, make your key facts (pricing, location, what you do) plainly visible as on-page text, and earn corroboration from sources that already rank. This is what gives you a chance whenever a prompt does trigger a live search. Our explainer on AI SEO vs traditional SEO lays out where the two overlap and where they part ways.
- Build entity presence over time. Get mentioned, reviewed, and referenced across many independent sources so the wider web describes you consistently. This is what feeds the training-data layer on the next model update. It is slower and less controllable than on-page work, and there are no shortcuts - but it is the only thing that addresses Source 2.
None of this is a guarantee that AI will recommend you. Nobody can promise that, and you should be wary of anyone who does. What you can do is measure your readiness on both layers, fix the specific gaps, and re-check whether the signals moved. That is a directional, honest way to work - not a crystal ball.
See where you stand on both layers
Free audit. A deterministic readiness score out of 100, your top fixes in plain English, no guarantees and no credit card.
Run free AI visibility auditFind GetVisus useful? Make us a preferred source in Google: