You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

IIS Rewrite Module配置:拦截多数爬虫但允许GoogleBot访问的解决方案

IIS Rewrite Module配置:拦截多数爬虫但允许GoogleBot访问的解决方案

最近我在折腾IIS的Rewrite规则,想尽可能拦截各类爬虫,一开始用的规则是这样的:

<rule name="BotBlock" stopProcessing="true">
  <match url=".*" />
  <conditions>
    <add input="{HTTP_USER_AGENT}" pattern="^$|bot|crawl|spider" />
  </conditions>
  <action type="CustomResponse" statusCode="403" statusReason="Forbidden" statusDescription="Forbidden" />
</rule>

这个规则的作用很明确:拦截所有User-Agent为空,或者包含bot、crawl、spider关键词的请求,效果本来还不错,但问题来了——它把GoogleBot也给拦了,这可不是我想要的结果。

我自己试了好几种正则写法,比如:

  • ^$|!googlebot|bot|crawl|spider
  • ^$|(?!googlebot)|bot|crawl|spider
  • ^(?!googlebot)$|bot|crawl|spider
  • ^$|(!googlebot)|bot|crawl|spider

结果要么是把所有User-Agent都拦了,要么还是没法让GoogleBot正常访问,属实头疼。

后来多亏了The fourth bird的帮忙,终于找到了正确的解决方案,修改后的规则条件部分是这样的:

<add input="{HTTP_USER_AGENT}" pattern="^$|\b(?!.*googlebot.*\b)\w*(?:bot|crawl|spider)\w*" />

这个正则的逻辑是:要么匹配空的User-Agent,要么匹配包含bot/crawl/spider但不包含googlebot的关键词,这样就完美实现了“拦截多数爬虫,但允许GoogleBot访问”的需求。

备注:内容来源于stack exchange,提问作者VDWWD

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.20 02:37:57