How do you get a small business found in ChatGPT and other AI assistants?
People are starting to ask instead of search. They open ChatGPT, Perplexity, Claude or Google’s AI Mode, describe what they need (“a physiotherapist in Wrocław who speaks English”), and contact whoever gets mentioned. This guide explains how a small business ends up in those answers. Everything here about how the crawlers behave comes from the companies’ own documentation, linked at the bottom.
Updated:
In short
- AI assistants can only quote pages they can fetch and read. Text that appears only after JavaScript runs is usually invisible to them.
- Each AI company runs several crawlers with different jobs: blocking the training one doesn’t block the search one, and vice versa.
- Google states plainly that no special markup or “AI file” is needed; a page has to be indexed and snippet-eligible to qualify as a supporting link.
- Assistants quote answers, not brochures. A page that answers one real question, with specifics, gets quoted; a page that praises a company doesn’t.
- Most recommendations of a business happen off your website — in directories, profiles and forums. If you exist nowhere else, you’re hard to recommend.
- AI visibility is slow: weeks for a page to be picked up, months before a new name shows up in competitive recommendations. Anyone promising a week is selling something.
Why are customers finding businesses through ChatGPT now?
The question changed shape. Instead of typing two keywords and scanning ten blue links, people describe their situation in a sentence and get one answer with a handful of sources. A business is no longer competing for a position on a page. It’s competing to be the source an assistant decides to cite.
The practical consequence is uncomfortable for a lot of small-business websites. Take the classic homepage: a warm headline, three service cards, a contact form. It gives an assistant almost nothing to quote. It says who you are, but it answers no question a customer ever asked. An assistant summarising “how much does a small business website cost in Poland” has nothing to take from it, so it takes something from somebody else’s page instead. That is the whole mechanism, and it explains why sites that look expensive and modern can be completely absent from AI answers while a plain page from a competitor gets quoted twice.
The good news is that the bar here is technical and editorial, not financial. You don’t need an agency retainer or a subscription to a visibility tool. You need pages that answer real questions, written by someone who knows the subject, and a website a crawler can read without running a browser. Most of what follows costs nothing but attention. The free checks in the next section take an afternoon, and they decide whether anything you write later can be seen at all.
How do AI assistants find your pages in the first place?
AI assistants find pages through crawlers, and every company runs several of them, each with its own job. Some collect training data, some build the index used to answer questions, and some fetch a page live while a person waits. Each one is controlled separately in robots.txt, so a decision about one doesn’t automatically apply to the others.
Two things follow from how these crawlers are documented. First, “blocking AI” isn’t one switch: it’s a set of decisions, and the most common accident is blocking the search crawler while meaning to block the training one. A business that opted out of AI training in 2024 with a single sweeping rule may have opted out of ChatGPT’s search answers at the same time, without ever noticing. Second, the assistant answering a live question may fetch your page at that very moment. A page that assembles its content in the browser, not in the HTML, can therefore be invisible exactly when it matters most.
Google works differently: its AI features draw on the ordinary search index. Google’s own documentation says there are “no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary”. It adds that you don’t need to create new machine-readable files, AI text files or markup. What Google does specify is the floor. To be eligible to be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet. In practice that means the old work is still the work: being indexed, being useful.
| Crawler | Run by | What it does | If you disallow it |
|---|---|---|---|
| OAI-SearchBot | OpenAI | Surfaces websites in ChatGPT’s search features | The site “will not be shown in ChatGPT search answers”, though it can still appear as a navigational link |
| GPTBot | OpenAI | Collects web content used to improve OpenAI’s foundation models | Signals the content shouldn’t be used for training. Search visibility is a separate setting, controlled by OAI-SearchBot |
| ChatGPT-User | OpenAI | Fetches a page during a user’s action inside ChatGPT | OpenAI notes robots.txt rules may not apply, because a person asked for that page |
| PerplexityBot | Perplexity | Surfaces and links websites in Perplexity results; “not used to crawl content for AI foundation models” | The site stops being surfaced in Perplexity results |
| Perplexity-User | Perplexity | Visits a web page while Perplexity answers a specific user question | It “generally ignores robots.txt rules”, because the fetch is user-triggered |
| Claude-SearchBot | Anthropic | Analyses web content to improve the quality of Claude’s search results | Anthropic states its bots honour robots.txt directives |
| ClaudeBot | Anthropic | Collects web content used for training and developing Claude | Anthropic states ClaudeBot honours robots.txt directives |
| Claude-User | Anthropic | Fetches a web page when a Claude user’s question needs it | Controlled by its own robots.txt user-agent entry, separately from ClaudeBot |
What should a small business check before writing anything new?
Five free checks come before new content, and they decide whether anything published later can be seen at all. Most small-business sites fail at least one of them. No amount of new writing compensates for a page an assistant can’t read, so the checks come first, in this order.
Two of those five checks deserve emphasis. The JavaScript check is where modern, expensive-looking websites fail hardest: a site built as a single-page app can rank perfectly well on Google (which does render scripts) and still be a blank page to an assistant’s crawler. The hosting check matters because it produces the most confusing symptom of all: everything looks fine in your browser, nothing is wrong in robots.txt, and no crawler ever arrives. Testing with a fake user-agent from your own computer proves nothing either way, because verified crawlers are recognised by signature, not by the name in the user-agent string. Search Console showing a page indexed is evidence. A request you sent yourself, with a name you typed, is not.
- 1 Open your page with JavaScript disabled and see what survives. In Chrome: open DevTools, press Cmd/Ctrl+Shift+P, run “Disable JavaScript” and reload. Whatever text disappears is text AI crawlers most likely never see — the major ones fetch JavaScript files but don’t execute them.
- 2 Read your robots.txt. Confirm the search-oriented crawlers are allowed: an explicit `User-agent: OAI-SearchBot` followed by `Allow: /` leaves no room for interpretation. Then check that you haven’t blocked them while intending to opt out of model training only.
- 3 Check your hosting or CDN settings, not just robots.txt. Some platforms block AI crawlers by default, so a page can be open in robots.txt while the network layer turns the request away. On Cloudflare this lives under Security → Settings (AI bot policies); on most hosts it’s a firewall or bot-protection setting.
- 4 Connect Google Search Console and Bing Webmaster Tools, and submit your sitemap in both. This is the only reliable way to know whether your pages are indexed, not merely published.
- 5 Make sure every page you care about is reachable by a link from another page. Pages that exist only in the sitemap get discovered late and treated as unimportant.
What kind of page actually gets quoted?
The page that gets quoted answers one question completely, in the words a customer would use, with specifics that can be checked. Assistants quote passages, not whole websites, so the unit of work is a self-contained section that still makes sense when it’s lifted out and pasted into an answer.
One thing deserves bluntness, because the industry is currently loud about it: writing a separate “machine-friendly” version of your content isn’t the job. Serving crawlers something different from what people see is cloaking, and Google’s spam policies are explicit that it can cost a site its rankings entirely. The page that gets quoted is the page a human would find useful. Structure of that kind simply makes the usefulness easier to extract; it is formatting applied to something true, not a substitute for having something to say.
- One clear question per page, phrased the way a person asks it — not a slogan and not a product name.
- An answer in the first two or three sentences of every section. If the answer arrives in paragraph six, the passage that gets extracted will be the wrong one.
- Real numbers, dates, names and units: a price with a currency, a timeframe in days, a date the information was checked. Vagueness is the main reason a page is passed over in favour of a competitor’s.
- Tables for anything comparative. We find a table row survives being quoted; the same information buried in flowing prose usually doesn’t.
- Sections that stand alone: no “as mentioned above”, no pronouns pointing back at the heading, and the name of the thing repeated instead of “it”.
- A visible “last updated” date that is true. Fresher pages tend to win over stale ones, but faking the date is manipulation, and it’s checkable.
- Honesty about when you’re the wrong choice. It costs one paragraph, and it’s the biggest single difference between a page that reads like a source and a page that reads like an advert.
Why isn’t a good website enough on its own?
A good website isn’t enough on its own, because most recommendations don’t come from your own site. When someone asks an assistant to suggest a studio, a clinic or a supplier, the answer is usually assembled from places where businesses are listed and discussed: directories, professional profiles, forums and review platforms. A business that exists only on its own domain is hard to recommend.
Off-site presence is the part small businesses postpone, because it’s admin work rather than creative work, and it’s usually the part that matters most. It’s also the part nobody can do for you: the accounts are yours, the verification is yours, and the details have to be true. Budget a day for it, do it once properly, and then leave it alone except when something changes — a new address, a new phone number, a new service. The failure mode here isn’t doing it badly. It’s doing it once, half-way, and leaving three different versions of your phone number scattered across the internet.
- A company profile on the professional network you use, filled in properly instead of left as a placeholder.
- Two or three relevant industry directories, chosen because your clients use them — not because they accept anyone. Registering with twenty irrelevant ones is a link-spam pattern and is treated as one.
- A Google Business Profile, plus an Apple Business Connect listing if your customers are on iPhones, assuming you have an address or a service area.
- The same name, address and contact details everywhere. Inconsistent details make one business look like several half-matching ones.
- Genuine reviews from real clients on platforms you don’t control. This is the only legitimate route to reputation signals, and the reason buying reviews is both fraud and a poor investment.
What if your customers speak more than one language?
Publish a proper version per language, each at its own address, and write it. Do not run it through a translator. In our experience, assistants answering in Polish or German lean on sources written in those languages, so a business with only an English page simply does not appear in a large part of its own market.
Two practical rules follow. Link the language versions to each other with hreflang tags (self-referencing, plus an x-default), so each address is understood as the same page in another language instead of as duplicate content. And publish all your languages at once: a half-translated site advertises versions that don’t exist, which produces errors in Search Console instead of reach.
The harder rule is editorial. A translated page reads like a translated page: the idioms are slightly off, the examples are foreign, the prices are in the wrong currency. That is exactly the tone a reader in that market discounts. If you can’t write it natively or have it written natively, it’s better to have fewer languages done properly. Two convincing languages beat four that sound like a machine, both for the person reading and for the assistant deciding whose page to trust.
What should you avoid, even though it’s being sold to you?
Avoid anything that manufactures signals instead of earning them. These tactics are the same search spam the industry has always had, wearing new vocabulary, and they carry the same consequences. At worst a site stops appearing in results at all, which takes months to recover from.
One more item belongs here, gently: the `llms.txt` file. It costs nothing to publish and it’s fine to have — we publish one ourselves. But no major vendor has committed to reading it as a ranking or citation signal, and it is no substitute for pages a reader would quote. Treat it as tidiness, not as a lever. The same goes for most tools sold as “AI visibility platforms”: the useful ones measure what’s happening, and that measurement earns its price once you have something to measure. None of them can make an assistant cite a page that doesn’t answer anything.
- Fake or incentivised reviews, and star ratings marked up on your own site with no verifiable source behind them.
- Bulk registration with low-quality directories for the sake of links instead of customers.
- Networks of thin, generated sites pointing at your domain.
- Text hidden for crawlers, or a different version of a page served to bots than to people.
- Instructions embedded in a page that try to manipulate the assistant reading it. Google’s spam policies explicitly cover attempts to manipulate generative AI responses in Search, and treat them like any other spam.
- Paying anyone who promises “guaranteed placement in ChatGPT”. OpenAI does sell clearly labelled sponsored placements through its own ads manager, but that’s advertising, marked as such, and a separate thing from the organic citations this guide is about. Being quoted as a source works by a different mechanism entirely, and no amount of money moves it.
How long does it honestly take?
Getting crawled and indexed is fast — days, once the checks in this guide pass and the consoles are connected. Being quoted is slow. The figures that follow are our working estimates, not vendor numbers. Assistants that fetch pages live can pick up a new page within weeks, while a new business name appearing routinely in competitive recommendations is a matter of months.
The order of events tells you whether to worry. First the crawlers arrive. Then the pages get indexed. Then pages start being fetched while somebody is in the middle of asking a question. That is the first firm sign anything is working, and it appears in your server logs before it appears anywhere else. Recommendations by name come last, and they depend far more on your off-site presence than on how well the page is written.
If nothing at all has changed after two months of a page being live, with no crawler visits and no index entries, the problem usually isn’t the writing. It’s either a technical block or a domain with nothing pointing at it yet, so the fix is the JavaScript and hosting checks, or the off-site profiles — not more pages. Publishing more pages is the standard response, and it’s the wrong one: ten invisible pages are exactly as invisible as one.
How can you tell whether it’s working?
Four measurements tell you whether AI visibility work is having any effect, and none of them cost money. Take a baseline before publishing anything. Without one there is nothing to compare against later, and the temptation to see progress that isn’t there is strong.
Analytics will under-report all of this, and it’s better to know that in advance: we consistently see a share of assistant traffic arrive without a referrer and land in “direct”. So the cleanest measurement remains the least technical one: asking new enquiries how they found you, and writing the answer down somewhere you will look at it later. Six months of those notes beat any dashboard, and they’re the only record that connects visibility to money instead of to activity.
- 1 Write down a fixed list of the questions your customers ask, in each language you sell in. Ask them in the assistants today and record what comes back. That’s your baseline, and it’s usually zero mentions.
- 2 Watch indexing in Search Console and Bing Webmaster Tools: how many of your pages are actually indexed, not how many you published.
- 3 Watch your server or CDN logs for AI crawlers, and separate the indexing bots from the live user-triggered fetches. On Cloudflare that’s the Security → Events log; on most hosts it’s the raw access log. Search the log for the eight crawler user-agent strings listed in this guide. Those ending in -User are live fetches, and a visit from ChatGPT-User or Perplexity-User means an assistant went to your page while answering somebody. That is the earliest evidence there is.
- 4 Re-run the same question list every two weeks, unchanged. Changing the wording feels productive and destroys the comparison.
Everything you might be wondering.
- I keep being told I need a special file for AI. Do I?
- No. Google’s documentation is explicit that no extra files or markup are needed for its AI features; the link is in the sources below. Structured data is still useful for describing your business accurately, but it isn’t an entry ticket, and no file makes an assistant quote a page that doesn’t answer anything.
- If I block AI crawlers from training on my content, do I disappear from ChatGPT?
- Not necessarily, because the crawlers are separate. OpenAI documents GPTBot for model training and OAI-SearchBot for surfacing sites in ChatGPT’s search features, each controlled independently in robots.txt. Blocking GPTBot doesn’t by itself remove you from ChatGPT’s search answers. Blocking OAI-SearchBot does.
- Do I have to pay OpenAI or Perplexity to be included?
- Not for organic citations — there’s no submission form and nothing to buy. OpenAI does sell labelled sponsored placements in ChatGPT through its ads manager, but that is advertising, not a citation. Anyone charging you for “guaranteed inclusion” in AI answers is selling something that doesn’t exist.
- ChatGPT recommends my competitor and not me. Why?
- Usually because your competitor exists in more places than you do. Assistants assemble recommendations from directories, profiles, reviews and articles as much as from company websites, so the business that’s listed, reviewed and written about wins — even when its own site is worse than yours. That’s the off-site work, and it is the least glamorous part of the job.
- An assistant says something wrong about my business. How do I fix it?
- Correct it at the source instead of arguing with the assistant. Find where the wrong information lives: an outdated directory entry, an old profile, a stale page of your own. Fix it there, then state the correct version plainly on your own site. Assistants reflect what the web says; changing the answer means changing the web.
- Will AI answers kill my website traffic?
- They change it more than they kill it. Some questions that used to bring a visit now get answered without one, and the visits that remain tend to arrive later in the decision, from people who already know what they want. It makes the quality of what you publish matter more than the quantity of visits it attracts.
- Is a blog necessary?
- Regular publishing isn’t the point — coverage is. A handful of pages that genuinely answer the questions your customers ask will do more than a blog updated weekly with company news. Update those pages when something changes, and say plainly when you last did.
- We’re a small local business. Is any of this realistic without an agency?
- Yes, for most of it. The checks and the profiles are admin, not expertise — anyone methodical can do them. The writing is the hard part, and a studio can do the technical layer, the structure and the drafting. But the raw specifics that get quoted — your prices, your timeframes, what you refuse to do — exist only in your head, so that part is yours either way.
Read next
- What a small business website really costs — what drives the price, and what the site costs to keep once it is live
- Website builder or a site built for you — where a platform limits what you can publish, and what you can take with you if you leave
Sources
- OpenAI — Bots and crawlers documentation ↗ — OAI-SearchBot, GPTBot and ChatGPT-User: purposes, effect of disallowing, ~24-hour propagation.
- Google Search Central — AI features and your website ↗ — No special markup or AI files required; supporting links must be indexed and snippet-eligible.
- Anthropic — Does Anthropic crawl data from the web? ↗ — ClaudeBot, Claude-User and Claude-SearchBot, and their robots.txt compliance.
- Perplexity — Perplexity Crawlers ↗ — PerplexityBot is not used for foundation-model training; Perplexity-User generally ignores robots.txt.
- Vercel — The rise of the AI crawler ↗ — Large-scale log analysis (published December 2024) finding that the major AI crawlers fetch JavaScript files but do not execute them.
- Google Search Central — Spam policies for Google web search ↗ — Cloaking, hidden text, link spam and manipulation of generative AI responses.
- Google Search Central — Structured data policies ↗ — Rules on fake, incentivised and self-serving review markup.