You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何验证Ahrefs、Semrush及Uptime.com爬虫(Bot)身份?

Verifying Crawlers Like Ahrefs, Semrush, and Uptime.com (When Reverse DNS Fails)

Great question—this is a super common headache when dealing with non-Google crawlers that don’t follow the clean reverse/forward DNS verification pattern we rely on for Googlebot. The good news is there are reliable workarounds to validate these tools and block unauthenticated crawlers. Here’s how to approach it:

1. Leverage Official Published IP Ranges

Most legitimate crawler providers (including Ahrefs, Semrush, and Uptime.com) publicly list their crawler IP address ranges. This is the most trustworthy method because IPs are harder to spoof than User-Agents.

  • How to use: Download the latest IP ranges from the crawler’s official documentation, then configure your server firewall, Nginx/Apache rules, or WAF to only allow requests from these IPs.
  • Example Nginx snippet:
    allow 54.36.149.0/24; # Replace with Semrush's actual IP range
    allow 178.128.0.0/16; # Replace with Ahrefs' actual IP range
    deny all;
    
  • Pro tip: Set up a monthly reminder to check for IP range updates—many providers expand their ranges periodically to avoid service disruptions.

2. Validate User-Agent Strings (With IP Context)

While User-Agents can be faked, pairing them with IP range checks makes this method much more reliable. Every legitimate crawler uses a unique, documented User-Agent string:

  • Ahrefs: AhrefsBot/7.0; +http://ahrefs.com/robot/
  • Semrush: SemrushBot/7~bl; +http://www.semrush.com/bot.html
  • Uptime.com: UptimeBot/1.0; +https://uptime.com/
  • Workflow: First confirm the request IP is in the provider’s published range, then verify the User-Agent matches the official format (use regex for flexibility if versions change).

3. Use Crawler-Specific Verification Tokens or Files

Some crawlers offer dedicated verification methods to prove their identity:

  • Verification files: Tools like Uptime.com might ask you to place a unique text file in your site’s root directory (e.g., uptime-verification-1234.txt). The crawler will check for this file before accessing your site, so you can block requests that don’t first validate against this file.
  • Custom request headers: A few providers let you configure a secret token that their crawler includes in a request header (like X-Crawler-Verify: your-secret-token). You can add server rules to block any requests missing this valid header.

4. Reverse DNS Pattern Matching (Even for "Unmeaningful" Hostnames)

While the reverse DNS might look generic (like ip170.ip-54-36-149.eu), some providers still embed subtle patterns in their reverse DNS entries:

  • For example, Ahrefs’ crawler IPs often resolve to hostnames containing ahrefs (e.g., crawl-001.ahrefs.com).
  • How to implement: Use a regex to match these patterns (e.g., .*ahrefs\.com$ or .*semrush\.net$), then follow up with a forward DNS lookup to confirm the resolved IP matches the original request IP.

Final Recommendation

No single method is 100% foolproof, but combining IP range validation + User-Agent checks is the most robust approach. This layered strategy minimizes false positives (so you don’t block legitimate crawlers) and ensures unauthorized bots are kept out.

内容的提问来源于stack exchange,提问作者Vahid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:35:58