For the complete documentation index, see llms.txt. Every page on this site is also served as Markdown: append `.md` to any URL, or send `Accept: text/markdown`.

DuckAssistBot vs PerplexityBot: two crawlers built for cited answers

DuckDuckGo and Perplexity both crawl the web for AI answers that cite their sources. They differ on training, on timing, and on what blocking one actually costs you.

Last updated: 2026-09-04By Karl-Gustav Kallasmaa
duckduckgo.com logo

DuckAssistBot

by DuckDuckGo

DuckDuckGo's crawler for AI-assisted answers. DuckDuckGo documents it as crawling pages in real time for AI-assisted answers that prominently cite their sources, and states the data is not used to train AI models.

Checked 2026-09-04T00:00:00Z
perplexity.ai logo

PerplexityBot

by Perplexity

Perplexity's search crawler. Perplexity documents it as designed to surface and link websites in search results on Perplexity, and recommends that publishers allow it in robots.txt.

Checked 2026-09-04T00:00:00Z

Which one should you choose?

Both of these crawlers exist to produce answers that link back to the publisher, which makes them the two AI agents with the clearest publisher-side upside on the open web. The differences are in the guarantees: DuckDuckGo states in writing that the data does not train models and that opting out does not touch organic rankings, while Perplexity's documentation actively recommends allowing its crawler and separates it cleanly from a sibling that ignores robots.txt.

Choose DuckAssistBot when

Prioritise DuckAssistBot access when you need an explicit written boundary between being crawled and being trained on, since DuckDuckGo states plainly that the crawler's data is not used to train AI models — a commitment most operators leave unwritten.

Choose PerplexityBot when

Prioritise PerplexityBot access when citation-driven referral traffic is the goal, since Perplexity's product surfaces and links sources by design and its documentation asks publishers to allow the crawler rather than merely permitting them to.

When neither is the right answer

Blocking either one is hard to justify for a site that wants to be found. Both agents cite and link the sources they use, so refusing them removes attribution you would otherwise receive for free. If crawl load is the real concern, rate-limit at your edge rather than disallowing a crawler that sends traffic back.

What is specific to this comparison

  • DuckAssistBot is the only crawler in this family whose documentation promises that opting out will not affect organic search rankings or result inclusion, because DuckDuckGo runs both a conventional search index and an AI answer layer and had to separate them in writing.
  • DuckDuckGo publishes a specific 72-hour window for a robots.txt opt-out to take effect, which is the only numeric propagation commitment any operator on this site makes; every other crawler leaves the delay unstated.
  • Perplexity's crawler documentation actively recommends that publishers allow PerplexityBot, which is a stronger position than the neutral 'here is how to block it' framing every other operator uses.
  • Perplexity is the operator in this pairing with a documented robots-ignoring sibling agent, so a publisher who blocks PerplexityBot has not blocked all Perplexity-originated fetches, while a DuckAssistBot block has no documented equivalent escape hatch.

DuckAssistBot vs PerplexityBot, criterion by criterion

Purpose
Documented purpose
Real-time crawling for AI-assisted answers that cite sourcesSource, checked 2026-09-04T00:00:00Z
Purpose
Data used for model training?
NoNo — stated explicitly in the documentationSource, checked 2026-09-04T00:00:00Z
Control
Respects robots.txt
YesYes — opt-out documented via robots.txtSource, checked 2026-09-04T00:00:00Z
Control
Documented time for an opt-out to take effect
72 hoursSource, checked 2026-09-04T00:00:00Z
Cost of blocking
Effect of blocking on conventional search results
YesNone — organic rankings and inclusion unaffectedSource, checked 2026-09-04T00:00:00Z
Verification
Published address list
Yesduckduckgo.com/duckassistbot.jsonSource, checked 2026-09-04T00:00:00Z
Identification
Published user-agent string
DuckAssistBot/1.2; (+http://duckduckgo.com/duckassistbot.html)Source, checked 2026-09-04T00:00:00Z
Context
Sibling agent documented as ignoring robots.txt
NoNone documented on the crawler pageSource, checked 2026-09-04T00:00:00Z

The short answer

These are the two crawlers on the open web with the least ambiguous case for allowing them. DuckAssistBot crawls in real time for DuckDuckGo's AI-assisted answers, which DuckDuckGo says prominently cite their sources. PerplexityBot exists to surface and link websites in Perplexity's results. Both operators have built products whose output points at the publisher, and both publish address lists so you can prove the traffic is real.

If you are triaging AI crawler decisions by risk, these two belong at the bottom of the list. The interesting question is not whether to allow them. It is what each operator has been willing to put in writing, because the two documents differ in ways that tell you something about each company.

What DuckDuckGo committed to on paper

DuckDuckGo's crawler page does three things almost nobody else does.

It states that the data is not used in any way to train AI models. That is an explicit negative, not an omission a publisher has to interpret. Most search crawlers operated by companies that also build models leave the boundary implicit, and a cautious publisher fills the silence with suspicion.

It states that opting out does not affect organic search rankings or result inclusion. DuckDuckGo runs a conventional search index alongside an AI answer layer, so it faced a question its pure-answer-engine peers do not: can a publisher refuse the AI feature without being punished in ordinary search? The answer is written down. Publishers who want to be in search results but not in AI summaries have a documented path.

And it states a number: robots.txt changes take effect after 72 hours. That single figure resolves the most common false alarm in crawler management, which is a site owner shipping a disallow, seeing the crawler again the next morning, and concluding the operator ignores the standard. Three days is the published window. Wait it out before escalating, and use the address list to confirm that what you are seeing is genuinely DuckAssistBot rather than something wearing its name.

What Perplexity committed to on paper

Perplexity's documentation is shorter and takes a different posture. It describes PerplexityBot as designed to surface and link websites in search results on Perplexity, and it recommends allowing PerplexityBot in your robots.txt.

That recommendation is worth pausing on. Every other operator in this space writes in the register of permission — here is the token, here is how to block it if you wish. Perplexity writes in the register of advocacy: allow this, it is in your interest. Whether you find that refreshing or presumptuous, it is a coherent position for a product whose entire output is a set of citations. Perplexity has less to gain from crawling a site that does not want to be quoted than a training-oriented operator does.

The other half of Perplexity's page is the part publishers should read carefully. Perplexity documents a second agent, Perplexity-User, and says it generally ignores robots.txt rules because the fetch originates from a user request. That disclosure is honest and it is also consequential: a robots.txt block on PerplexityBot stops the systematic crawl, and does not stop a Perplexity user's fetch of a specific page. DuckDuckGo's crawler page documents no equivalent agent, so a DuckAssistBot block has no documented bypass.

Where the two documents leave gaps

Neither page is complete, and it is worth naming what is missing rather than pretending the sourcing is exhaustive.

Perplexity's crawler documentation is scoped to search surfacing and does not make a statement about training the way DuckDuckGo's does. That is an absence, not a denial, and it should be read as exactly that: the page does not say the data trains models, and it does not say it does not. A publisher who needs that assurance in writing has it from one of these two operators and not the other.

DuckDuckGo's page does not state a propagation window for anything other than robots changes, and neither operator documents crawl-delay support. So for both agents, the vocabulary is allow or disallow, and any finer control is something you build at your own edge with rate limits.

Both publish addresses, which is the single most useful thing a crawler operator can do, and it means every claim about compliance on this page is checkable by a publisher with access logs and half an hour.

The robots.txt lines

Allow both, which is what the evidence supports for almost any site that wants to be found:

plain text
User-agent: DuckAssistBot
Allow: /

User-agent: PerplexityBot
Allow: /

Refuse DuckDuckGo's AI answers while staying in its conventional search index, which is the configuration DuckDuckGo's own documentation makes safe:

plain text
User-agent: DuckAssistBot
Disallow: /

Exclude a specific path from both — gated resources, customer lists, pricing experiments — while leaving your documentation and marketing pages available:

plain text
User-agent: DuckAssistBot
Disallow: /internal/

User-agent: PerplexityBot
Disallow: /internal/

Two mechanical notes. Robots.txt matches on the token, so the User-agent: line takes PerplexityBot and never the full Mozilla string. And rules are per host and per scheme: your documentation subdomain serves its own robots file and carries its own answer, and for most technical companies that host holds the pages an answer engine would most want to quote.

Version numbers, again

DuckAssistBot's published user agent carries /1.2; PerplexityBot's carries /1.0. Both numbers will change eventually, and any filter, firewall rule or dashboard written against the full string will stop matching on the day they do — silently, and in a way that looks exactly like a crawler losing interest in your site.

Match on the bare token and on the self-identifying URL, both of which are stable across versions. This is the single most common reason a company concludes that an AI crawler abandoned them, and it is a monitoring bug rather than a crawl event every time.

Two products, two very different reasons to be crawled

The crawlers look similar in a robots file. The products behind them are not, and the difference changes what allowing each one is worth to you.

DuckDuckGo is a search engine that added an answer layer. A user who asks it a question may get an AI-assisted answer with citations, or may get a conventional list of results, and DuckDuckGo has committed in writing that your presence in the second is unaffected by your decision about the first. The upside of allowing DuckAssistBot is therefore incremental: you were already in the index, and this adds a chance of appearing in the summarised answer above it.

Perplexity has no conventional results list to fall back on. The product is the answer, and the citations are the interface. A publisher who is not in Perplexity's crawl is not ranked lower — they are simply not among the sources the answer is built from. That makes the upside of allowing PerplexityBot categorical rather than incremental, and it explains why Perplexity's documentation asks rather than merely permits.

The practical consequence for prioritisation: if you have to argue one of these internally, argue Perplexity first, because the downside of being absent is total rather than partial. But the argument should rarely be necessary. Neither crawler asks for anything a publisher who wants readers should be reluctant to give.

The 72-hour rule is worth internalising

Most crawler escalations are premature. A team ships a robots.txt change, watches the logs for a day, sees the crawler again and starts drafting a complaint about an operator that ignores the standard.

DuckDuckGo has removed the ambiguity for its own crawler by publishing the window: 72 hours. Nothing you observe inside that period is evidence of anything. It is also a reasonable mental default for operators who publish no number at all, because the underlying mechanics are the same everywhere — robots files are fetched on a schedule and cached, and a crawl already scheduled against an older copy will still run.

The corollary is about the other direction, and it is the one people forget. An allow takes time to have an effect too. Opening access to a crawler that has been blocked for a year does not produce citations that week: the crawler has to return, fetch, and be selected as a source for a question somebody asks. Teams that flip the switch and check their numbers on Friday conclude the change did nothing. Give it a quarter, and measure whether the pages were fetched at all before drawing any conclusion about whether they were quoted.

What allowing them does not buy you

Being crawled is necessary for being cited and nowhere near sufficient. Both of these agents will happily fetch a page that neither product ever quotes, because the decision about what to quote is made downstream against whatever else is available on the question.

The work that closes that gap is content work, not robots work: does the page answer a question somebody actually asks, in a form that can be lifted into an answer without a human editor rewriting it, better than the pages currently being quoted. No directive in a robots file shortens that path, and any tool that implies otherwise is selling you a robots.txt editor with a dashboard attached.

It is also worth being precise about limits. Neither directive controls copies of your content that somebody else hosts under their own robots file. Neither governs what a person pastes into a chat window — Perplexity documents a separate agent for exactly that traffic. And neither is a licence: a robots directive is a machine-readable preference made to a well-behaved client, which both of these operators demonstrably are.

Confirming access, per URL

A robots.txt edit is an intention; your access log is the outcome. Both operators publish address lists, so for these two agents you can do the full verification: match the user-agent token, check the requesting address against the published file, and confirm which URLs were actually fetched.

That last part is where most of the value is. Knowing an answer-engine crawler reached your homepage tells you very little. Knowing whether it has read the twelve documentation pages that answer your buyers' real questions tells you whether you are a candidate source at all. Attensira's crawler logs record which AI agent fetched which URL and when. To read your current rules before changing them, the robots.txt generator and the bot access score will tell you what your file permits today.

For PerplexityBot against the crawler behind ChatGPT's search features, read OAI-SearchBot vs PerplexityBot. For Perplexity's internal split between the crawler that honours robots.txt and the fetcher that does not, see PerplexityBot vs Perplexity-User. And for the two largest assistant vendors' search crawlers head to head, see Claude-SearchBot vs OAI-SearchBot.

Where Attensira fits, and where it does not

Attensira's crawler logs record which of these agents fetched which URL and when, which is how you tell an answer-engine crawler that is reading your documentation from one that only ever sees your homepage.

See how Attensira compares to both

Questions people ask

DuckDuckGo's help page for the crawler states that the data is not used in any way to train AI models. That is an explicit negative in the operator's own documentation, which is unusual — most search crawlers leave the boundary between indexing and training to be inferred.

DuckDuckGo's documentation says opting out does not affect organic search rankings or result inclusion. The crawler is scoped to AI-assisted answers, so blocking it removes you from those answers while leaving DuckDuckGo's conventional search results alone.

DuckDuckGo documents that robots.txt changes take effect after 72 hours, and points publishers at crawling@duckduckgo.com for opt-out issues. That published delay is worth knowing before you conclude a directive was ignored.

Perplexity's crawler documentation describes PerplexityBot as respecting robots.txt and recommends allowing it in your site's robots.txt file. Its sibling agent, Perplexity-User, is documented differently — Perplexity says that fetcher generally ignores robots.txt rules because the fetch originates from a user request.

Both operators publish machine-readable address lists — duckduckgo.com/duckassistbot.json for DuckAssistBot and perplexity.com/perplexitybot.json for PerplexityBot. That puts both ahead of operators who publish only a user-agent string, which anyone can forge.

Sources

Every claim on this page, with the page it came from and the date that page was read. Prices and feature lists change; these are what the source said on the date shown, not timeless facts.

  1. DuckDuckGo documents DuckAssistBot as a web crawler for DuckDuckGo Search that crawls pages in real time for AI-assisted answers, which prominently cite their sources.DuckAssistBot is a web crawler for DuckDuckGo Search that crawls pages in real-time for our AI-assisted answers, which prominently cite their sources.https://duckduckgo.com/duckduckgo-help-pages/results/duckassistbot — read 2026-09-04T00:00:00Z
  2. DuckDuckGo states that the crawler's data is not used in any way to train AI models.not used in any way to train AI modelshttps://duckduckgo.com/duckduckgo-help-pages/results/duckassistbot — read 2026-09-04T00:00:00Z
  3. DuckDuckGo publishes DuckAssistBot's user-agent string as DuckAssistBot/1.2; (+http://duckduckgo.com/duckassistbot.html)https://duckduckgo.com/duckduckgo-help-pages/results/duckassistbot — read 2026-09-04T00:00:00Z
  4. DuckDuckGo documents that a robots.txt change disallowing DuckAssistBot takes effect after 72 hours.https://duckduckgo.com/duckduckgo-help-pages/results/duckassistbot — read 2026-09-04T00:00:00Z
  5. DuckDuckGo states that opting out of DuckAssistBot does not affect organic search rankings or result inclusion.https://duckduckgo.com/duckduckgo-help-pages/results/duckassistbot — read 2026-09-04T00:00:00Z
  6. DuckDuckGo publishes DuckAssistBot's addresses as a machine-readable list at duckduckgo.com/duckassistbot.json.https://duckduckgo.com/duckduckgo-help-pages/results/duckassistbot — read 2026-09-04T00:00:00Z
  7. Perplexity documents PerplexityBot as designed to surface and link websites in search results on Perplexity.designed to surface and link websites in search results on Perplexityhttps://docs.perplexity.ai/docs/resources/perplexity-crawlers — read 2026-09-04T00:00:00Z
  8. Perplexity publishes PerplexityBot's user-agent string as Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)https://docs.perplexity.ai/docs/resources/perplexity-crawlers — read 2026-09-04T00:00:00Z
  9. Perplexity documents PerplexityBot as respecting robots.txt and recommends allowing PerplexityBot in a site's robots.txt file.https://docs.perplexity.ai/docs/resources/perplexity-crawlers — read 2026-09-04T00:00:00Z
  10. Perplexity publishes PerplexityBot's address ranges at perplexity.com/perplexitybot.json.https://docs.perplexity.ai/docs/resources/perplexity-crawlers — read 2026-09-04T00:00:00Z
  11. Perplexity documents a second agent, Perplexity-User, which it says generally ignores robots.txt rules because the fetch originates from a user request.https://docs.perplexity.ai/docs/resources/perplexity-crawlers — read 2026-09-04T00:00:00Z