Getting cited by ChatGPT, Perplexity, Gemini or Claude is a different problem from ranking on Google, and the differences are mechanical rather than stylistic. This guide covers the citation mechanics specifically, alongside our complete guide to AI search optimisation.
How many sources each engine names
Watch: the short version
A minute on what assistants quote, and the check worth doing before anything else.
| Engine | Typical citations | Notable behaviour |
|---|---|---|
| Perplexity | 5 to 10 | Most generous and most transparent, strong recency preference |
| ChatGPT search | 3 to 5 | Leans on Bing’s index, Wikipedia heavily represented |
| Google AI Overviews | 3 to 4 | Built on Google’s ranking stack, local relevance weighted |
| Claude | Varies | Synthesises more than it quotes, favours logical structure |
The arithmetic is worth sitting with. Page one of Google offered ten organic positions. An AI answer offers three or four, and there is no page two.
The technical prerequisites
Before anything else, confirm the crawlers can reach you. The agents that matter are GPTBot and ChatGPT-User, ClaudeBot, PerplexityBot, Google-Extended and Bingbot.
Two traps catch people repeatedly. An old robots.txt with a broad disallow that now blocks agents nobody anticipated. And Cloudflare, which changed its default posture to block AI crawlers, meaning a site can be excluded without its owner changing a thing.
Check your server logs for those user agents rather than reading your configuration. Absence of hits is the answer.
Then check rendering. AI crawlers execute JavaScript far less reliably than Googlebot. Disable JavaScript in your browser and reload your most valuable page. What survives is roughly what the machine sees.
Bing is ChatGPT infrastructure
ChatGPT’s retrieval leans on Bing’s index, and so does Copilot. That makes Bing Webmaster Tools direct infrastructure for the largest AI search platform rather than optional housekeeping, and it is routinely neglected because Bing’s own traffic looks small.
Verify your site is indexed there, submit your sitemap, and check for crawl errors. It is an afternoon of work with leverage well beyond the effort.
Earning citation beyond your own site
Models form a view of you from the whole web. Unlinked brand mentions appear to contribute, unlike classical link equity, which makes digital PR directly relevant.
The fastest route in is getting represented on pages the engines already cite for your target questions. Being added to an existing frequently cited comparison article beats building a competing page from nothing. Reddit, YouTube and specialist forums also appear disproportionately in AI answers, and genuine participation helps while astroturfing is detectable.
When the citation is wrong
Being cited inaccurately can cost more than being absent. Ask each engine directly what your company does, what it costs and who runs it. Most businesses have never done this once.
Where a claim is wrong, engines that cite will usually show you the source. It is normally one specific stale page, sometimes your own. Fix it there, publish a clear current version, and re-check on a schedule, because corrections propagate slowly.