← All posts
Frameworks··9 min read

How to Get Your Business Cited by AI Answer Engines

Ask ChatGPT or Perplexity a question and it answers directly, citing a few named sources. Here is the concrete, mostly technical checklist that makes your business one of them: an llms.txt file, Speakable schema, and an AI crawler allowlist you control.

A single clean sheet of paper lifted out of a dark stack of documents, caught in a warm gold beam of light on a wooden desk, representing a page chosen as the source an AI assistant quotes
Answer

To get your business cited by AI answer engines, make your site easy to read and safe to quote. Publish a clean llms.txt file, add Schema.org structured data including Speakable, format answers as clear questions and facts, and allow the AI crawlers such as GPTBot and PerplexityBot in your robots.txt.

Ask ChatGPT, Perplexity, or Google's AI Overviews a real question and you get one synthesised answer with a few named sources under it. This guide is about how to get your business cited by AI: not ranked as a blue link, but quoted as the source the assistant trusts. The work is concrete and mostly technical, and most sites in your market have not done it yet. We build this readiness layer for clients as an audit-first consultancy, so what follows is the checklist we actually ship, not theory.

Getting cited is a different job from ranking

Classic search optimisation aims to rank a page so a person clicks the link. Answer engine optimisation, sometimes called AEO or generative engine optimisation, aims for something else: to be the quoted source inside the answer the machine writes. A page can rank well and never get cited. A page can get cited without being the top link. The engine is not rewarding position. It is rewarding copy that is easy to read and safe to quote.

Both still matter. You keep the search foundation, and you add the answer layer on top of it. We wrote the strategy side of this in the answer engine optimization guide and the wider version in the generative engine optimization guide. This post is the build side: the specific files and formats that make a site citable.

The five things that make a site citable by AI

Everything below is checkable, and none of it needs a new platform or a large budget. It needs someone to do the unglamorous parts that most sites skip. Do these five and you have the readiness layer answer engines look for.

1. A clean llms.txt file

An llms.txt file is a plain-text file at the root of your domain, at /llms.txt, that points assistants at your most important pages in a form they can read without wading through navigation and scripts. The format is a small emerging standard, documented at llmstxt.org. It does not replace your robots.txt or your sitemap. It is a short, curated map that says, in plain words, here is what this business is and here are the pages that answer the common questions. It costs an afternoon and almost nobody in your market has one.

2. Schema.org structured data, including Speakable

Structured data tells the engine what each block on your page actually is, so it does not have to guess. Use the vocabulary from Schema.org: Organization for the business, Article for a post, FAQPage for a question block, Product or Service where they fit. Then add Speakable, the property that marks the specific sentences meant to be read aloud or lifted as a short answer. Speakable is a direct hint about which lines to quote. Most sites ship none of this, which is why the ones that do get read cleanly.

3. Question-and-answer formatting, answer first

Write the way people ask. Lead each page with a direct 40 to 60 word answer in plain language, then expand below it. Use question-shaped headings, short FAQ blocks, and tables where a comparison fits. These are pre-structured, so the engine does not have to interpret a wall of prose to find the point. Our SEO and AEO content engine ships every page this way by default, because answer-first copy is the single highest-return change for citation.

4. Factual, well-sourced copy

An engine quotes what it can verify. Name specifics, attribute your facts, and link the source when you cite a figure or a standard. Vague, unsourced copy is risky for a machine to repeat, so it skips it in favour of a competitor who was precise. This is also where schema and clean writing meet: the pattern is spelled out in our note on schema markup for AI search citations. Precision is not a style choice here. It is what makes you safe to cite.

5. Let the AI crawlers in

None of the above matters if your robots.txt blocks the bots that read the web for the answer engines. Check that these named user-agents are allowed, not disallowed by an old copied rule:

  • GPTBot and OAI-SearchBot, from OpenAI, for ChatGPT and its search.
  • ClaudeBot and anthropic-ai, from Anthropic, for Claude.
  • PerplexityBot, from Perplexity.
  • Google-Extended, Google's token for AI Overviews and Gemini.
  • Bingbot, for Bing and Copilot.

Each provider documents its crawler and the exact user-agent string: OpenAI, Perplexity, and Google all publish theirs. The robots.txt standard is how you allow or block them. Blocking a bot is a valid choice for some businesses. The point is to do it on purpose, after a decision, not because a default config quietly shut the door.

How to check whether it is working

You do not need a paid tool to get a first read. Ask the assistants the questions your buyers ask. Open ChatGPT, Perplexity, and Google's AI Overviews, type the real queries in your category, and see whether your domain is named in the answer. Do this before you change anything, so you have a baseline, and again a few weeks after you ship the readiness layer. The move from unnamed to named is the result you are buying.

Then confirm the crawlers are actually reading you. Your server access logs record the user-agents that fetched your pages, so search them for GPTBot, ClaudeBot, PerplexityBot, and the rest. If a bot never appears, it is either blocked or has not found the pages your llms.txt points to. That log check is the difference between assuming you are readable and knowing it. Watch your branded searches too: when people see you cited, they tend to search your name next.

Two mistakes that keep you uncited

Two errors show up again and again, and both are invisible to a human visitor. The first is content that only appears after JavaScript runs. If your key answer is loaded by a script after the page opens, a crawler may never see it, so keep the important answer in the initial HTML. The second is a thin page: schema cannot give a page substance it does not have, so if the page answers no real question, mark-up will not save it. Fix the content first, then structure it. Both are cheap to correct once you know to look, which is exactly what a short audit surfaces.

What we have actually built

Receipts beat promises, so here is real work and no invented numbers. For a transport company in Tartu we prepared the site for AI search. We added an llms.txt file, Speakable schema, and an explicit allowlist for the AI bots, so assistants like ChatGPT and Perplexity can read and cite it. On the same rebuild the valid structured-data blocks went from 4 to 75, and the hero image dropped from 60 KB to 14 KB. That is the readiness layer described above, shipped on a real site. We do not dress it up as a citation count or an ROI figure we cannot show you, because a number we cannot prove is worth less than the honest version.

Who should do this now, and who should wait

Do it now if your buyers research with AI before they contact you. If people ask ChatGPT or Perplexity for a shortlist in your category, the sites that are readable and citable get named and the rest are invisible in a way no rank report will show. Being cited also reads like a referral: the assistant vouches for you inside its own answer.

Wait if your real leak is elsewhere. If you sell locally on referral and repeat business, or leads come in and fall through a broken booking flow, fix that first. More citation just fills a leakier bucket. The free AI audit tells you which leak is top of the list, and sometimes the honest answer is that it is not this one. When the gap is operational rather than visibility, we point you at automation instead.

How we build and keep the readiness

Our method is audit first. A free audit checks whether your content is readable and citable by AI search, in plain terms you can act on. Then we build the readiness: the llms.txt file, the structured data including Speakable, and the crawler allowlist. Then we keep it in tune as your site changes and the crawlers evolve, because a config that was right last year can quietly drift. The code and the accounts stay in your name, so nothing stops working the day an engagement ends. You can read how we work as an AI consultancy or book the free audit.

How do I get my business cited by AI answer engines?

Make your site easy to read and safe to quote. Publish a clean llms.txt file, add Schema.org structured data including Speakable, format your content as clear questions and answers, keep the copy factual and well sourced, and confirm your robots.txt allows the AI crawlers such as GPTBot, ClaudeBot, PerplexityBot, and Google-Extended. That is the readiness layer engines look for.

What is an llms.txt file?

An llms.txt file is a plain-text file at the root of your domain, at /llms.txt, that points AI assistants at your most important pages in a clean form they can read without wading through menus and scripts. The format is defined at llmstxt.org. It does not replace robots.txt or a sitemap. It is a short, human-readable map of what matters on your site.

Which AI crawlers should I allow in robots.txt?

The main named user-agents are GPTBot and OAI-SearchBot from OpenAI, ClaudeBot and anthropic-ai from Anthropic, PerplexityBot from Perplexity, Google-Extended from Google, and Bingbot from Microsoft. Check that your robots.txt does not disallow them by an old copied rule. Blocking a bot is a valid choice, but make it a decision, not an accident.

Is answer engine optimization different from SEO?

Yes, and they work together. Classic SEO optimizes to rank a link so a person clicks it. Answer engine optimization, also called AEO or GEO, optimizes to be the quoted source inside an AI answer. You still need the search foundation. AEO is the newer layer on top of it, and most sites in a given market have not built it yet.

Can a small business get cited by ChatGPT?

Yes. Citation rewards clarity and structure more than size or budget. A small business with a clean llms.txt file, Speakable schema, factual answer-first pages, and an open crawler allowlist is often easier to quote than a large site buried in scripts. The work is unglamorous rather than expensive, which is exactly why so few competitors do it.

If you want the version scoped to your business instead of a generic checklist, start where we always start: book the free audit through contact, or see the SEO and AEO service for how the build works. Thirty minutes, your site, and a straight answer on whether AI search is worth your next move.

Next move

Find your leak. Book the audit.

The free AI audit maps your inbound, qualification, booking, and follow-up. We rank exactly where the leak is before you spend a dollar.

AI consultancyShip in daysGlobalNow booking July
kratt

The AI consultancy that finds the money your business is losing, then builds, hosts, and runs the AI to get it back. Shipped in days, not months.

★ Now bookingEU + APAC
The newsletter

Occasional notes on
what’s actually working.

No spam. Cancel anytime. Occasional notes only.
DOC · KRATT-FOOT-001 · © 2026 Kratt · All rights reserved
Book your free AI audit