For the complete documentation index, see llms.txt. Every page on this site is also served as Markdown: append `.md` to any URL, or send `Accept: text/markdown`.

ChatGPT-User vs Perplexity-User: two live fetchers that robots.txt may not stop

Both agents fetch a page because a person asked a question, and both operators state that robots.txt rules may not apply to them. That agreement is the story.

Last updated: 2026-09-03By Karl-Gustav Kallasmaa
openai.com logo

ChatGPT-User

by OpenAI

OpenAI's user-initiated agent. It visits a page when somebody asks ChatGPT or a custom GPT a question, and it also backs GPT Actions. OpenAI states it is not used for automatic crawling.

Checked 2026-09-03T00:00:00Z
perplexity.ai logo

Perplexity-User

by Perplexity

Perplexity's user-initiated agent. It visits a web page to help answer a question a person has asked, and Perplexity states the response may include a link back to that page.

Checked 2026-09-03T00:00:00Z

Which one should you choose?

These two agents are the closest thing to twins in this family. Both fetch on a person's behalf, both are documented as outside ordinary robots.txt control, and both publish an address list you can enforce against if you must. The real decision is not which to allow but whether you want to be in the answer at all.

Choose ChatGPT-User when

Reason about ChatGPT-User when your question is about GPT Actions and custom GPTs as well as chat. It is the only agent here that also backs an application-integration path, so it can reach endpoints a search crawler never would.

Choose Perplexity-User when

Reason about Perplexity-User when attribution is what you care about. Perplexity documents that the response may include a link to the page it visited, which is a return on the fetch that the ChatGPT-User documentation does not describe.

When neither is the right answer

For nearly every commercial site the right action is none. Both agents arrive because a person asked about you; refusing them does not stop the answer being given, it only removes your own page from the sources it is built on.

What is specific to this comparison

  • This is the only pairing in the family where the two operators agree with each other and disagree with a third: OpenAI and Perplexity both place user-initiated fetches outside robots.txt, while Anthropic explicitly places Claude-User inside it.
  • ChatGPT-User is the only agent on this site documented as backing an application-integration path — GPT Actions — which means it can reach endpoints that exist for software rather than for readers.
  • Perplexity documents an attribution link back to the fetched page and OpenAI's ChatGPT-User description does not, so the two agents differ in what the publisher gets in return for the fetch.
  • Both operators publish a dedicated address list for the agent that will not obey robots.txt, which is the practical admission that enforcement, not request, is the mechanism they expect a determined publisher to use.

ChatGPT-User vs Perplexity-User, criterion by criterion

Control
Operator's stated robots.txt position
NoRules may not applySource, checked 2026-09-03T00:00:00Z
Purpose
What triggers the fetch
A question in ChatGPT or a custom GPTSource, checked 2026-09-03T00:00:00Z
Scope
Also used for application integrations
YesYes — GPT ActionsSource, checked 2026-09-03T00:00:00Z
Value to you
Documented as linking back to your page
Not documentedNot stated in the crawler documentationSource, checked 2026-09-03T00:00:00Z
Identification
Published user-agent string
Yes...compatible; ChatGPT-User/1.0; +https://openai.com/botSource, checked 2026-09-03T00:00:00Z
Verification
Published address range list
Yesopenai.com/chatgpt-user.jsonSource, checked 2026-09-03T00:00:00Z
Cost of blocking
Affects whether you appear in the engine's search index
NoNo — that is OAI-SearchBot's jobSource, checked 2026-09-03T00:00:00Z
Nature
Described as an automatic crawler
NoExplicitly notSource, checked 2026-09-03T00:00:00Z

The short answer

ChatGPT-User and Perplexity-User do the same thing: they fetch a page because a person asked a question that the page might answer. Both operators state, in their own documentation, that robots.txt may not govern them. OpenAI writes that because these actions are initiated by a user, robots.txt rules may not apply. Perplexity writes that since a user requested the fetch, its fetcher generally ignores robots.txt rules.

If you have been told that a Disallow line will stop AI assistants reading your site, this is the pairing that shows why that advice is incomplete.

Two operators agreeing, and one that does not

The agreement between OpenAI and Perplexity is more interesting than either statement alone. Independently, both companies have decided that the robots exclusion standard governs automated crawling and not a fetch a human being requested. That is a coherent reading. Robots.txt was designed when the thing on the other end was a crawler working through a queue, not an assistant retrieving one URL because somebody typed a question.

Anthropic reached the opposite conclusion for the same situation.[^anthropic-contrast] Its documentation states that Claude-User honours robots.txt and describes the token as the mechanism by which site owners control which sites can be accessed through user-initiated requests. Same problem, same year, opposite answer.

That disagreement is the single most important fact a publisher can hold about this category, because it means there is no industry norm to rely on. Each operator's own words are the specification, and they differ. Advice that says "AI bots respect robots.txt" or "AI bots ignore robots.txt" is wrong about at least one of the three largest.

What each agent actually does

ChatGPT-User covers more ground than its name suggests. OpenAI documents it as the agent used when somebody asks ChatGPT or a custom GPT a question and it visits a page — and also for interactions with external applications via GPT Actions. That second role is unique in this family. Every other agent here is trying to read a document. ChatGPT-User can also be calling an interface on behalf of a user, which means it may touch endpoints that were never designed with a reader in mind.

OpenAI is also careful to fence off what ChatGPT-User does not do. It states the agent is not used for crawling the web in an automatic fashion, and that it is not used to determine whether content may appear in Search, directing site owners to OAI-SearchBot for search opt-outs and automatic crawl. So a rule about ChatGPT-User is not a rule about your ChatGPT search visibility — a distinction that gets lost constantly.

Perplexity-User is narrower and, in one respect, more generous. Perplexity documents it as visiting a web page to help provide an accurate answer, and states the response may include a link to the page. That link is the return on the transaction. OpenAI's ChatGPT-User description does not make an equivalent statement, which is not evidence that no link appears — it is simply not documented, and this page marks it unknown rather than filling the gap with an assumption.

The economics of a live fetch

It is worth being direct about why most sites should welcome both agents.

A training crawler visits speculatively. A live fetcher visits because a specific person, at a specific moment, is trying to find something out — quite possibly about you. That is the highest-intent moment in the entire funnel, and the only question is whether the answer they get is assembled from your current page or from a model's recollection of an older one.

Refusing the fetch does not remove the question. It removes your page from the set of sources the answer is built on, and it hands the description of your product to whoever else wrote about it. For a company whose commercial problem is being described inaccurately in AI answers, blocking a live fetcher is close to the exact opposite of the correct move.

The honest counter-case belongs to publishers whose revenue is the visit. A summarised answer with an attribution link does not pay the same as a reader on the page, and for a news site or a paid research archive that difference is the whole business. Those publishers have a real grievance, and the agents' robots.txt position makes it harder for them to act on it. That is a fair thing to be angry about; it is just not most readers' situation.

GPT Actions makes ChatGPT-User a different kind of problem

The line about GPT Actions in OpenAI's documentation deserves more attention than it usually gets, because it takes ChatGPT-User out of the category of "thing that reads pages" entirely.

An assistant reading a marketing page is a low-stakes event. An assistant calling an interface on a user's behalf is not. The same user agent covers both, which means a rule about ChatGPT-User is simultaneously a rule about whether people can use your product through an assistant. For a company that publishes an API, that is a product decision wearing the costume of a crawler decision, and it will usually be made by whoever is editing robots.txt rather than by whoever owns the API.

The practical implication is that the token deserves a different review than the others on this site. The questions are the ones you would ask about any integration path. Which endpoints can be reached this way? Do they require authentication, and does that authentication behave sensibly when the caller is an assistant rather than a browser? Are the rate limits appropriate for a caller that may retry? Is there anything reachable that assumes a human is on the other end?

Those are good questions to have answered regardless of your position on AI crawling, and this is a reasonable occasion to answer them. What would be a mistake is to resolve them by blanket-disallowing the token and assuming the matter is closed — both because the operator has said robots.txt may not apply, and because the thing you would be turning away is a person trying to use your product.

Neither agent is a substitute for being indexed

A recurring confusion worth heading off: live fetching and indexing are different systems, and being good at one does not compensate for being absent from the other.

An assistant fetches a URL when it already has a reason to believe that URL is relevant — a link in the conversation, a result from a search step, something in the model's own knowledge. Live fetching is retrieval of a known candidate, not discovery. If nothing in the engine's index points at your page, the live fetcher will never be pointed at it either, no matter how permissive your rules are.

This is why the tokens that actually govern discovery — OAI-SearchBot at OpenAI, PerplexityBot at Perplexity — are the ones with the higher stakes, and why OpenAI's documentation redirects site owners to OAI-SearchBot for search opt-outs rather than letting a ChatGPT-User rule stand in for one. A site that disallows the index crawler and allows the live fetcher has kept the door open and removed the sign from the street.

The correct configuration for almost everybody is both: allow the crawler that gets you into the candidate set, and accept the fetcher that arrives when a person's question makes you a candidate.

If you decide to block anyway

Robots.txt will not do it, by both operators' own accounts. What will is enforcement at the server or edge, and both companies publish what you need to write an accurate rule: openai.com/chatgpt-user.json and perplexity.com/perplexity-user.json.

Blocking by user-agent string alone is weak in both directions. It over-blocks, because any client can send that string, and it under-blocks for the same reason. Matching the published address ranges is the version that works, and the fact that both operators publish those ranges for the agent that will not obey robots.txt is a quiet admission that they expect determined publishers to use them.

Write the robots directives anyway if you want the record. They document intent, they cost nothing, and if an operator's policy changes they are already in place:

plain text
User-agent: ChatGPT-User
Disallow: /

User-agent: Perplexity-User
Disallow: /

Just do not report that as a block until the logs agree.

Measuring what you cannot control

When the file is not the authority, the log is. Both agents identify themselves clearly — match on openai.com/bot and perplexity.ai/perplexity-user, the self-identifying URLs, rather than on version numbers that move. Then check the addresses against the published lists before counting anything.

What you learn is genuinely useful beyond compliance. Live-fetch traffic to a specific URL is evidence that people are asking questions your page is a candidate answer for. It is one of the few available signals about demand at the moment of the question rather than at the moment of the search.

Attensira's crawler logs record these fetches by agent and URL, which is how a live fetch becomes a number rather than a rumour. The bot access score shows what your current rules permit, and the robots.txt generator will write the file — with the caveat, stated here rather than buried, that for these two agents the file is a preference and not a control.

Perplexity's own split between its obedient crawler and its disobedient fetcher is covered in PerplexityBot vs Perplexity-User. Anthropic's opposite policy is in ClaudeBot vs Claude-User. For the tokens that genuinely do control search visibility, read GPTBot vs OAI-SearchBot and OAI-SearchBot vs PerplexityBot.

Where Attensira fits, and where it does not

Because neither agent is reliably governed by robots.txt, the only honest measurement is the log. Attensira's crawler logs record which agent fetched which URL and when, which is what tells you these fetches are happening at all.

See how Attensira compares to both

Questions people ask

Both operators say no, in slightly different words. OpenAI writes that because these actions are initiated by a user, robots.txt rules may not apply. Perplexity writes that since a user requested the fetch, its fetcher generally ignores robots.txt rules.

No. OpenAI states that ChatGPT-User is not used to determine whether content may appear in Search, and directs site owners to use OAI-SearchBot in robots.txt for managing Search opt-outs and automatic crawl.

Both operators say it is not. OpenAI states that ChatGPT-User is not used for crawling the web in an automatic fashion, and Perplexity describes its equivalent as visiting a web page to help answer a specific question. The distinction is what both companies use to justify the robots.txt position.

Enforcement at the server or edge, matched against the published address lists — openai.com/chatgpt-user.json and perplexity.com/perplexity-user.json. Both operators publish those lists, which is what makes an accurate rule possible.

Usually not. These agents fire because somebody is already asking about you, and Perplexity states its fetcher may include a link to the page in its response. Turning them away removes your page from an answer that will be given anyway.

Sources

Every claim on this page, with the page it came from and the date that page was read. Prices and feature lists change; these are what the source said on the date shown, not timeless facts.

  1. OpenAI documents ChatGPT-User as the agent used when a person asks ChatGPT or a custom GPT a question and it visits a web page, and for GPT Actions.When users ask ChatGPT or a CustomGPT a question, it may visit a web page with a ChatGPT-User agent.https://developers.openai.com/api/docs/bots — read 2026-09-03T00:00:00Z
  2. OpenAI states that because ChatGPT-User actions are initiated by a user, robots.txt rules may not apply.Because these actions are initiated by a user, robots.txt rules may not apply.https://developers.openai.com/api/docs/bots — read 2026-09-03T00:00:00Z
  3. OpenAI states that ChatGPT-User is not used for crawling the web in an automatic fashion.ChatGPT-User is not used for crawling the web in an automatic fashion.https://developers.openai.com/api/docs/bots — read 2026-09-03T00:00:00Z
  4. OpenAI publishes ChatGPT-User's user-agent string as Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bothttps://developers.openai.com/api/docs/bots — read 2026-09-03T00:00:00Z
  5. OpenAI publishes ChatGPT-User's address ranges at openai.com/chatgpt-user.jsonhttps://developers.openai.com/api/docs/bots — read 2026-09-03T00:00:00Z
  6. OpenAI documents ChatGPT-User as also covering interactions with external applications via GPT Actions.ChatGPT users may also interact with external applications via GPT Actions.https://developers.openai.com/api/docs/bots — read 2026-09-03T00:00:00Z
  7. Perplexity documents Perplexity-User as visiting a web page to help provide an accurate answer when a user asks a question.When users ask Perplexity a question, it might visit a web page to help provide an accurate answerhttps://docs.perplexity.ai/guides/bots — read 2026-09-03T00:00:00Z
  8. Perplexity states that since a user requested the fetch, Perplexity-User generally ignores robots.txt rules.Since a user requested the fetch, this fetcher generally ignores robots.txt ruleshttps://docs.perplexity.ai/guides/bots — read 2026-09-03T00:00:00Z
  9. Perplexity publishes Perplexity-User's user-agent string ending in compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-userhttps://docs.perplexity.ai/guides/bots — read 2026-09-03T00:00:00Z
  10. Perplexity publishes Perplexity-User's address ranges at perplexity.com/perplexity-user.jsonhttps://docs.perplexity.ai/guides/bots — read 2026-09-03T00:00:00Z
  11. Perplexity's documentation describes Perplexity-User only in terms of answering user questions, with no application-integration role.https://docs.perplexity.ai/guides/bots — read 2026-09-03T00:00:00Z
  12. Anthropic takes the opposite position for its own user-initiated agent, stating that Claude-User honours robots.txt and that the token lets site owners control user-initiated access.https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler — read 2026-09-03T00:00:00Z