SEO

Getting cited by ChatGPT: AI search optimization, shown on our own site

August 13, 2026
Xavier PeichBy Xavier Peich

How an SMB gets cited by ChatGPT, Perplexity and Claude. The method we apply to peich.xyz, with no magic promises.

Getting cited by ChatGPT: AI search optimization, shown on our own site

More and more potential clients no longer type their question into Google: they ask ChatGPT, Perplexity, or Claude, and read the answer without ever clicking a link. For an SMB, the question becomes concrete: when an AI assistant sums up your industry, is your business part of the answer? This is the new front of SEO, and it plays by different rules than classic search. People sometimes call it GEO (generative engine optimization), but the name matters little.

This article isn't theory. It's what we did to our own site, peich.xyz, to be retrievable and quotable by the assistants, and what we refuse to do because nobody can guarantee it. Our business model (websites and AI agents on subscription) forces us to be visible where our clients look. So we treated our own site as the test bench.

The good news for a Quebec SMB: in French, competition on this ground is still thin. The basics, done well, are often enough to stand out.

The short answer, for the busy

To get cited by ChatGPT, Perplexity, or Claude, an SMB has to be retrievable first, then quotable. Retrievable means the AI search bots can read your pages: OpenAI uses OAI-SearchBot for its search answers, Anthropic a Claude-SearchBot, and Perplexity its PerplexityBot. Allow them explicitly in your robots.txt, because some hosts block AI bots by default. Quotable means each page carries a short, self-contained answer, written to be lifted as-is, with facts, numbers, and names that stay consistent everywhere your business appears. The llms.txt file can help, but it's still emerging: no major engine has confirmed it influences citations. The most underrated factor is off-site: assistants pull from public discussion (reviews, forums, directories), so a business talked about in credible places gets cited more often. And nobody guarantees a citation: be wary of anyone who promises one. Set up the conditions, then measure by asking the assistants directly.

Two paths into an AI answer

You have to separate two mechanisms, because you don't work them the same way.

The first is training memory. The model read part of the web during its training phase, months ago. If your business was already widely mentioned there, it "knows" something about you. You don't directly control this path: it's slow, frozen at training time, and it favours already-established brands. For an SMB, this is not where the game is won.

The second is live retrieval. When you ask a timely or local question, the assistant runs a web search at answer time, reads a few fresh pages, and cites its sources. This is the path within your reach. It doesn't reward age; it rewards a page the bot can read today and quote effortlessly. Everything below targets this second mechanism.

Retrievable and quotable beats "optimized"

The old SEO reflex is to "optimize": keyword density, tags, tricks. AI assistants don't care. What they want is a passage they can extract and present as an answer, with confidence.

In practice, a quotable page answers a real question up front, in a few sentences that stand on their own, without forcing the reader (or the machine) to reconstruct meaning from ten paragraphs. That is exactly the logic of the "short answer" block you just read, which we put at the top of every blog article. It isn't a gimmick: it's the format a generative search engine prefers to cite.

The other pillar is entity consistency. Your business name, your industry, your city, your services and your numbers have to be identical everywhere: site, Google profile, directories, social. An assistant that sees three different versions of your offering hesitates to cite you. An assistant that sees the same information, consistent and verifiable, treats you as a reliable source.

What we did to peich.xyz

Here's the exact list, reproducible on any site.

First, a robots.txt that explicitly allows the AI bots. By default our site lets everything through, but we still name GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and Google-Extended with an "Allow" directive. The reason is defensive: several hosts and firewalls (CDNs) have started blocking AI bots by automatic setting. A silent block makes you invisible in ChatGPT's answers without you knowing. A detail that matters: at OpenAI, OAI-SearchBot is the bot that governs appearance in ChatGPT search, separate from GPTBot, which is for training. Blocking one is not blocking the other.

Second, an llms.txt file at the root. It's a structured summary of who we are, what we do, our pricing, our key pages, written to be read by a machine. More on it below, because its status deserves nuance.

Third, the self-contained answer block at the top of every article, and entity consistency kept by hand across the site, the Google profile, and directories. Nothing clever. Just the basics, done seriously.

llms.txt: useful, emerging, not magic

Let's be honest, because plenty of agencies sell this file as a magic wand. The llms.txt is a proposed standard, not a citation guarantee.

The numbers say so. A May 2026 scan of the 10,000 most popular sites (Tranco ranking) found a valid llms.txt on about 5.9% of them: adoption is rising, but stays a minority, concentrated in tool and documentation sites. More importantly, to date no major provider (OpenAI, Google, Anthropic) has publicly confirmed that this file influences how its assistants rank or cite content. Its real current users are mostly coding assistants, which use it to find the right page without wasting context.

Our position: we keep the llms.txt, because it takes a few minutes to write and can't hurt. But we don't charge anyone for "llms.txt magic", and you shouldn't pay much for it either. It's hygiene, not a lever.

Off-site weighs more than you think

Here's the counter-intuitive part. For retrieval answers, what's said about you elsewhere often weighs more than your own site. Assistants readily pull from public discussion: Google reviews, forums, Reddit threads, industry directories, local press. A business talked about in several credible places becomes a "safe" source to cite.

The practical consequence overlaps with work you're already doing. Your Google reviews and your local SEO presence no longer serve only Google: they also feed what the AIs know about you. Being listed in a Quebec business directory, answering in a forum in your industry, getting local coverage: all of it now counts twice.

Nobody guarantees a citation

It has to be said plainly: nobody can promise you'll be cited by ChatGPT. Models change, search mechanisms evolve, and the providers don't publish their recipe. Anyone selling you an "AI citation guarantee" is selling air. The right posture is that of honest SEO: you set up the conditions that raise your odds, and you measure.

Measuring, precisely, is within reach. Ask ChatGPT and Perplexity, with search on, the way a client would: "best web agency for SMBs in Quebec", "AI agent provider in Montreal". Look at who's cited, and whether you're there. Then watch your traffic: in your analytics, visits coming from chatgpt.com or perplexity.ai are a concrete signal that the assistants are sending you people. They're imperfect indicators, but real ones.

Where to start

Three moves, in order. Check that your robots.txt doesn't block the AI bots (and name them explicitly if your host has a hair trigger). Add, at the top of your important pages, a short, self-contained answer to the visitor's real question. Then strengthen your off-site presence: reviews, directories, mentions, consistency everywhere.

That's exactly what we check in an audit. We look at whether the assistants can read you, whether your pages are quotable, and where your off-site presence has holes. 30 minutes, no commitment.

→ Request a visibility audit

Xavier Peich

Written by

Xavier Peich