AhrefsBot vs SemrushBot: two link-index crawlers you cannot verify the same way
Both crawl your site to build a commercial SEO index. One publishes its addresses and honours crawl-delay; the other says not to block it by IP at all.
AhrefsBot
by Ahrefs
The Ahrefs link crawler. Ahrefs documents it as powering the database for both Ahrefs, a marketing intelligence platform, and Yep, an independent privacy-focused search engine, and as strictly obeying disallow and crawl-delay directives.
Checked 2026-09-04T00:00:00ZSemrushBot
by Semrush
The Semrush crawler family. Semrush documents SemrushBot as collecting data for its public backlink index, alongside a set of separately-named tokens for Site Audit, Backlink Audit, on-page checks, content tools and other products.
Checked 2026-09-04T00:00:00ZWhich one should you choose?
Neither of these crawlers touches your visibility in AI answers or in Google, and both are blockable through robots.txt. They differ on everything a site reliability engineer would care about: Ahrefs publishes addresses and reverse-DNS hostnames and obeys crawl-delay as written, while Semrush publishes no addresses, caps crawl-delay at ten seconds, and spreads its crawling across a family of tokens that a single rule does not cover.
Choose AhrefsBot when
Allow AhrefsBot when you want a verifiable crawler on your site — the published IP endpoints and the ahrefs.com and ahrefs.net reverse-DNS suffixes mean you can prove any given request is genuine, which is rare and useful when you are triaging unexplained traffic.
Choose SemrushBot when
Allow SemrushBot when your own team uses Semrush tools against your site, since several of the documented tokens exist to serve audits and checks that you initiated. Blocking the family wholesale can break the tooling your marketing team is paying for.
When neither is the right answer
Block both if neither product is in your stack and crawl budget is genuinely tight. Unlike a search or answer-engine crawler, refusing these costs you no distribution — the only loss is that your backlink profile becomes less visible inside two commercial products, which chiefly affects competitors researching you.
What is specific to this comparison
- SemrushBot is the only crawler documented on this site whose operator explicitly advises publishers not to block it by IP, on the grounds that it does not use consecutive IP blocks — the opposite of the published-address-list posture every AI crawler operator has adopted.
- Semrush documents a cap on crawl-delay: values above ten seconds are cut down to ten. No other operator on this site documents a ceiling on a directive it claims to honour, which means a publisher asking for a sixty-second gap is not getting one.
- Ahrefs is the only crawler here whose verification story includes a reverse-DNS suffix as well as published addresses, giving two independent checks on a forged user agent.
- These are the only two crawlers in the compare family whose blocking decision has no effect on any answer surface at all — refusing them changes what two commercial SEO products know about your backlinks, and nothing a user or an assistant ever sees.
AhrefsBot vs SemrushBot, criterion by criterion
The short answer
These two crawlers are the largest non-search, non-AI consumers of your bandwidth, and they are the ones most often confused with agents that affect visibility. They do not. AhrefsBot powers the Ahrefs database and the Yep search engine. SemrushBot feeds Semrush's public backlink index and its associated tools. Blocking either changes what a commercial SEO product knows about your link profile. It does not change what Google ranks, and it does not change what any assistant answers.
That makes this a pure operations decision, which is refreshing, and it is decided on operational grounds: how much load each crawler imposes, how precisely you can control it, and whether you can verify that the traffic is genuine.
The token problem
The first practical difference is that "SemrushBot" is not one crawler.
Semrush documents a family of separately-named tokens, each attached to a different product: SiteAuditBot, SemrushBot-BA, SemrushBot-SI, SemrushBot-SWA, SplitSignalBot, SemrushBot-OCOB, SemrushBot-FT, RyteBot and SemrushBot-ESI, alongside the bare SemrushBot that feeds the backlink index. Robots.txt matching is on the token, so a Disallow written under SemrushBot governs that token and nothing else. A team that blocks it and then wonders why Semrush-branded traffic continues has not been ignored — they have written a rule for one of the crawlers.
This is not obfuscation. The granularity has a real benefit: your own team's Site Audit runs under a token you can allow while refusing the backlink crawl you did not ask for. But it does mean the block is a list, not a line, and the list changes as products are added.
AhrefsBot is documented as a single primary token, which makes the rule trivial to write and trivial to review.
Verification, and Semrush's unusual position
Ahrefs publishes crawler IP ranges and individual crawler IPs through public API endpoints, and documents that all of its crawler's reverse DNS hostnames end with ahrefs.com or ahrefs.net. That is two independent verification paths on top of the user-agent string, and it is stronger than what several AI crawler operators offer.
Semrush takes the opposite approach and says so directly: do not try to block SemrushBot via IP, because it does not use any consecutive IP blocks. The documented control is robots.txt, per token.
Take that at face value about the architecture — a crawler distributed across non-contiguous addresses is a real engineering choice, not evasion. But be honest about the consequence. Without a published address list or a documented reverse-DNS suffix, a request identifying as SemrushBot cannot be verified. Anyone can send that header, and scrapers routinely wear the user agents of reputable crawlers because a recognised name draws fewer blocks and fewer challenges.
So if you disallow SemrushBot and keep seeing it in logs, the supportable conclusion is not that Semrush ignores robots.txt. It is that somebody is using the name and you have no way to tell who. Handle unverifiable branded traffic as a capacity question — rate-limit at the edge — rather than building a compliance argument on a string.
Crawl-delay, and the cap nobody notices
Both operators honour crawl-delay, and one of them puts a ceiling on it.
Ahrefs documents AhrefsBot as strictly obeying both disallow rules and crawl-delay directives, and says robots.txt changes are picked up before the next scheduled crawl. It also documents automatic backoff: 4xx and 5xx responses are treated as signals to reduce crawl speed. That is the behaviour you want from a crawler on a site with a fragile origin — it responds to distress without being asked.
Semrush documents that its crawler can take intervals of up to 10 seconds between requests, and that higher values will be cut down to that 10-second limit. It also documents that if no delay is specified it adjusts request frequency according to current server load, and that a robots.txt change is discovered within up to one hour or 100 requests.
The cap is the detail worth internalising. A site owner who writes Crawl-delay: 60 because their origin is struggling has expressed one request per minute and will receive one per ten seconds — six times the traffic they asked for, with no error, no warning and nothing in the file to indicate the difference. It is honestly documented and almost never read.
What blocking actually costs you
Nothing that a user sees. That is the unusual and liberating property of this pairing.
Your backlink profile becomes less complete inside those two products. The people most affected are competitors researching you, agencies pitching you, and your own team if you are a customer. If your marketing team pays for Semrush and runs Site Audit against your site, blocking the audit token breaks a tool you are already buying — which is the most common self-inflicted wound in this area, and it usually happens when an infrastructure team blocks a crawler family during an incident without asking who uses it.
There is a second-order effect worth naming honestly rather than overstating: link data from these indexes circulates into third-party research, agency reports and public write-ups, some of which is read by people evaluating you and some of which ends up as web content that AI systems later read. That is a long, indirect chain and nobody should present it as a visibility mechanism. It is a reason not to treat blocking as entirely free, not a reason to keep a crawler you cannot afford.
The robots.txt lines
Allow Ahrefs, throttled, using a directive it obeys as written:
User-agent: AhrefsBot
Crawl-delay: 10Block the Semrush backlink crawl while keeping the audit tooling your team paid for:
User-agent: SemrushBot
Disallow: /
User-agent: SiteAuditBot
Allow: /Refuse both entirely, which is a defensible position for a site with tight origin capacity and neither product in its stack:
User-agent: AhrefsBot
Disallow: /
User-agent: SemrushBot
Disallow: /Two mechanics. Robots.txt matches on the token, so the User-agent: line takes AhrefsBot and not the full Mozilla string with the version in it. And rules are per host and per scheme — every host you serve carries its own file and its own answer, and a rule on the apex domain does nothing for a documentation subdomain.
Measuring the load before deciding anything
The argument for blocking an SEO crawler is almost always framed as a load argument, and it is almost never measured first.
The measurement is straightforward. Take a week of access logs, group by user-agent token, and produce three numbers per crawler: requests, bytes served, and the share of requests that hit an expensive path — anything database-backed, anything with query parameters, anything your cache misses. Most sites discover the distribution is nothing like the intuition. Two or three crawlers account for the bulk of requests, and a large share of those requests land on a handful of generated paths that no human ever visits.
That result usually points at a better fix than a block. If ninety percent of a crawler's requests are hitting faceted search URLs or an infinite calendar, disallowing those paths for every crawler solves the load problem, keeps every index intact, and improves how real search engines crawl you at the same time. Blocking a single named crawler treats the symptom and leaves the trap in place for the next one.
It also gives you a defensible answer when someone in marketing asks why their tool stopped seeing the site. "We measured, this crawler was eight percent of bot requests, and the real cost was a crawl trap we have now fixed" is a different conversation from "we blocked some bots during an incident."
Reviewing the token list on a cadence
Because Semrush's crawling is spread across a family of tokens tied to products, the list is not stable. Products get added, tools get renamed, and a robots file written to be exhaustive two years ago now names some of the crawlers and misses others.
Put both operator pages on the same quarterly review as your AI crawler tokens. Re-fetch, diff the documented token list against what your file names, and note the tokens you are deliberately silent about. Silence is a legitimate choice — it just should not be an accident, and with a family of nine or ten named agents, accident is the default state.
Do not confuse these with the agents that matter
The reason this comparison sits alongside pages about AI crawlers is that the two categories get conflated constantly, usually during an incident. Someone sees heavy bot traffic, reaches for a blocklist copied from a forum post, and ships a commit that blocks AhrefsBot, SemrushBot, GPTBot, OAI-SearchBot and ClaudeBot in one go. The first two were the actual load problem. The next three were the site's distribution, and now they are gone.
The categories genuinely differ. SEO crawlers build commercial indexes about your site; refusing them is a cost question. Answer-engine crawlers decide whether your pages can be quoted in the answers your buyers read; refusing them is a distribution question. They share a file format and nothing else, and the only way to keep them apart in a stressful moment is to have decided about each in a calm one.
If load is the pressure, the tools that fit are here: a crawl-delay for the operator that honours it as written, an edge rate limit for the one you cannot verify, and a path scope for the heavy sections — faceted search, calendars, generated archives — that account for most crawler expense on most sites.
Confirming what is actually fetching you
A robots.txt edit is an intention; your access log is the fact. For AhrefsBot you can verify end to end: match the token, check the address against the published endpoints, confirm the reverse-DNS suffix. For SemrushBot you have the token and your own rate data, which is a weaker but workable basis for a capacity decision.
Attensira's crawler logs record which agent fetched which URL and when, which is how SEO-tool traffic stops being mixed in with the AI agents that determine whether you get cited. To read your current rules before changing them — including which tokens your file does not mention at all — the robots.txt generator and the bot access score will tell you what your file permits today.
Related comparisons
For a crawler that builds a public archive rather than a commercial index, read CCBot vs GPTBot. For the operators who publish per-crawler address files and what that buys you, see Amzn-SearchBot vs OAI-SearchBot. And for a crawler family where one token bundles two jobs, see meta-webindexer vs meta-externalagent.
Where Attensira fits, and where it does not
Attensira's crawler logs record which agent fetched which URL and when, which is how you separate SEO-tool crawling from the AI agents that actually determine whether you get cited.
See how Attensira compares to bothQuestions people ask
Sources
Every claim on this page, with the page it came from and the date that page was read. Prices and feature lists change; these are what the source said on the date shown, not timeless facts.
- Ahrefs documents AhrefsBot as powering the database for both Ahrefs, a marketing intelligence platform, and Yep, an independent privacy-focused search engine.Powers the database for both Ahrefs, a marketing intelligence platform, and Yep, an independent, privacy-focused search engine.https://ahrefs.com/robot — read 2026-09-04T00:00:00Z
- Ahrefs publishes AhrefsBot's user-agent string as Mozilla/5.0 (compatible; AhrefsBot/7.0; +http://ahrefs.com/robot/)https://ahrefs.com/robot — read 2026-09-04T00:00:00Z
- Ahrefs documents AhrefsBot as strictly obeying both disallow rules and crawl-delay directives, with robots.txt changes picked up before the next scheduled crawl.https://ahrefs.com/robot — read 2026-09-04T00:00:00Z
- Ahrefs publishes crawler IP ranges and individual crawler IPs through public API endpoints at api.ahrefs.com.https://ahrefs.com/robot — read 2026-09-04T00:00:00Z
- Ahrefs documents that all of its crawler's reverse DNS hostnames end with ahrefs.com or ahrefs.net.https://ahrefs.com/robot — read 2026-09-04T00:00:00Z
- Ahrefs documents that AhrefsBot recognises 4xx and 5xx status codes as signals to reduce its crawling speed automatically.https://ahrefs.com/robot — read 2026-09-04T00:00:00Z
- Semrush documents data collected by SemrushBot as feeding a public backlink search engine index maintained as a dedicated tool called Backlink Analytics, alongside other tools including Site Audit, Backlink Audit, Link Building and content analysis.https://www.semrush.com/bot/ — read 2026-09-04T00:00:00Z
- Semrush documents a family of separately-named crawler tokens, including SiteAuditBot, SemrushBot-BA, SemrushBot-SI, SemrushBot-SWA, SplitSignalBot, SemrushBot-OCOB, SemrushBot-FT, RyteBot and SemrushBot-ESI, each tied to a different tool.https://www.semrush.com/bot/ — read 2026-09-04T00:00:00Z
- Semrush documents that its crawler can take intervals of up to 10 seconds between requests to a site, and that higher crawl-delay values will be cut down to this 10-second limit.can take intervals of up to 10 seconds between requests to a site. Higher values will be cut down to this 10-second limit.https://www.semrush.com/bot/ — read 2026-09-04T00:00:00Z
- Semrush states that publishers should not try to block SemrushBot via IP because it does not use any consecutive IP blocks, and publishes no IP list.Do not try to block SemrushBot via IP as we do not use any consecutive IP blocks.https://www.semrush.com/bot/ — read 2026-09-04T00:00:00Z
- Semrush documents that a robots.txt change is discovered within up to one hour or 100 requests.https://www.semrush.com/bot/ — read 2026-09-04T00:00:00Z