Brain+AI

SEO and AI visibility 7 min

How to get cited by ChatGPT and AI search engines

To get cited by ChatGPT and other AI engines, their crawlers first have to be able to read your site, which many setups quietly prevent. Then each page should answer one precise question in its opening lines, with sourced and dated facts. Nobody can guarantee a citation: you can only remove the obstacles and check, query by query.

The first reason you are not cited: the bots cannot read your site

There is a lot of talk about “GEO” (Generative Engine Optimization) as a brand new discipline. In our audits, the first problem is not new at all: the site simply cannot be read by the assistants’ crawlers. Check this before rewriting a single page, because no amount of editorial work makes up for a locked door.

Each vendor publishes its list of crawlers, and they do different jobs. At OpenAI, the official documentation separates GPTBot, used to improve its models, from OAI-SearchBot, used to surface websites in ChatGPT’s search features. The same page states that a site blocking OAI-SearchBot will not be shown in ChatGPT search answers, and that a robots.txt change takes roughly 24 hours to be picked up. At Anthropic, the dedicated help page describes three bots: ClaudeBot (collection for training), Claude-User (when a user asks a question that requires visiting a page) and Claude-SearchBot (search result quality). All of them honour robots.txt.

In practice, you can refuse training and accept search, or the other way round. That is a legitimate editorial choice. What is not legitimate is having the choice made for you without knowing it.

What we did: the Palmora case

Palmora Property is our own real estate agency in Phuket, with a site in six languages. The agreed strategy was clear: stay readable by AI assistants. The robots.txt file in the repository was clean and allowed everyone except private areas.

Yet the SEO and GEO audit of 22 July 2026 found something else in production. The robots.txt actually served started with a block added by Cloudflare, its “managed robots.txt”, which did not exist in the code. It contained Disallow: / for GPTBot, ClaudeBot, CCBot, Google-Extended, Amazonbot, Applebot-Extended, Bytespider, meta-externalagent and Cloudflare’s own rendering crawler. Nobody had asked for it. We switched it off the same day.

Three lessons from that episode:

  • Check what is served, not what is written. One curl on the production URL showed the block. Reading the repository file would never have revealed it.
  • The network layer can rewrite your files. A CDN, a web application firewall or an “AI bot protection” option enabled by default all act before your site does.
  • The block had side effects. The non-standard lines of the managed block made Lighthouse report the robots.txt as invalid, which capped the SEO score at 92.

An honest caveat: the blocked list mostly targeted training crawlers, and OAI-SearchBot was not on it. So the site was not necessarily missing from every ChatGPT answer, but it was shut out for several vendors, against our will. Whenever the topic comes back, the first move is now to recheck that Cloudflare setting.

Answer straight away, and answer one question

Once the door is open, page shape matters. An assistant composing an answer extracts passages. A passage that makes sense on its own, lifted out of its page, has a better chance of being reused than a paragraph that assumes you read the three before it.

We apply two simple rules:

  • The answer first. On this blog, every article opens with a direct answer of two to four sentences, displayed under the title. That is the fragment we want extracted. The rest of the article backs it up.
  • One page, one question. That is the principle behind the knowledge base of Challenge Us, our group challenge app: 45 pages, each titled with a question in users’ own words, with a two or three sentence answer that has to stand on its own. The site refuses to build if two pages serve the same answer or if a title is not a question.

The same discipline helps people too: a reader in a hurry reads three lines and leaves with the answer.

Sourced and dated facts

The only serious academic study we know on the subject is “GEO: Generative Engine Optimization” (Aggarwal et al., presented at KDD 2024). The authors built a benchmark of 10,000 queries and tested nine ways of rewriting content. The three that worked best: adding quotations, adding statistics, citing sources. Keyword stuffing, on the other hand, brought little to no improvement. They report a visibility gain of up to 40% in generated answers, and up to 37% in a test on Perplexity.

Read those figures for what they are. The visibility measured is the share of the answer’s text that comes from a source, weighted by position: not traffic, not revenue. The main experiments rely on a simulated engine (GPT-3.5-turbo fed with the top five Google results), and the authors themselves warn that the methods will need to adapt as engines evolve. It is solid evidence about direction, not a guaranteed recipe.

That direction matches what we do for other reasons anyway: on our sites, no figure is published without a link to its primary source, actually opened. An assistant choosing between two pages has good reasons to prefer the one that says where its claims come from. And dating every page (publication date, update date, and for Challenge Us the app version in which the answer was checked) tells the reader, human or machine, what it is talking about.

Structured data, only for what is displayed

Structured data (JSON-LD) helps engines understand what a page is: an article, an organisation, a question and its answer. It is not a box to tick for AI engines. Google says so plainly in its documentation on AI features: there are no additional requirements to appear in AI Overviews, and no special file or markup to create.

The rule that matters is elsewhere. Google’s general guidelines forbid marking up content that readers cannot see. We turned that into a technical constraint: on the Challenge Us site, the displayed FAQ and the marked-up FAQ read the same list, so they cannot drift apart. A marked-up FAQ that does not exist on screen is a penalty you build yourself.

How to measure without fooling yourself

This is the part “GEO” offers skim over the most. There is no Search Console for assistants. Here is what we do, and nothing more:

  1. A list of target queries, typed by hand into ChatGPT, Claude and Perplexity on a fixed schedule, noting whether the site is cited and which page. It is manual, answers vary from one session to the next, and that is precisely why we repeat the measurement instead of drawing conclusions from one try.
  2. Server or CDN logs, filtered on the bot names (OAI-SearchBot, GPTBot, ChatGPT-User, ClaudeBot, Claude-User, Claude-SearchBot). They tell you whether the bots come and which pages they read. They do not tell you whether you are cited.
  3. Visits from assistants in your analytics, when the referrer is passed. That is a lower bound: many citations produce no click at all.

What we do not do: promise a number of citations. Nobody controls what a model decides to cite. An agency that guarantees it is selling an impression, not a result.

Where to start

In this order: check the served robots.txt and your CDN settings, put the answer at the top of the pages that matter, source and date, then mark up what is displayed. The llms.txt file comes after that, and we have detailed what you can really expect from it. On a multilingual site, add a hreflang check: an assistant does not cite a page it cannot find in the right language, and the most common mistakes are easy to spot.

These checks are part of our SEO audit, and the ongoing work of our SEO and GEO engine. The overall method is laid out in what an SEO agency should deliver in the age of AI.

Frequently asked questions

Should I block GPTBot to protect my content? It is an editorial decision. Blocking GPTBot refuses training of OpenAI’s models; OAI-SearchBot is the one that decides whether you appear in ChatGPT search. Decide bot by bot, reading each vendor’s documentation, and check what is actually served in production, CDN included.

Does GEO replace SEO? No. Generative engines largely rely on pages they can find and read, and the reference study was run on an engine fed by Google results. A technically sound, fast, well structured site remains the foundation for both. GEO adds a way of writing on top: the answer first, every fact sourced.

How long until we get cited? Nobody can honestly say. OpenAI’s documentation says a robots.txt change is picked up in about a day, but being chosen in an answer depends on the query, the competition and the model. Measure query by query, on a fixed schedule, and adjust.

Jim · founder of Brain Plus AI

Also in seo and ai visibility

  • Understand

    llms.txt: what it is really for

    The llms.txt file promises to help AI read your site. What the proposal says, what can be verified about real usage, what Google, OpenAI and Anthropic say about it, and why we publish one anyway.

    6min
  • Choose and priceThe topic guide

    AI SEO agency: what it should actually deliver

    "AI" has become every SEO agency's sales pitch. Here is what an AI SEO agency should concretely deliver, the questions that expose the others, and why human review remains non-negotiable.

    5min
  • Practical guide

    Hreflang and multilingual SEO: the costly mistakes

    Hreflang pointing to pages that do not exist, a poorly chosen x-default, 404s in the wrong language, untranslated titles, a language routed but forgotten elsewhere: the mistakes we found on our own multilingual sites, and how to avoid them.

    6min
WhatsApp