IIS Rewrite Rule正则优化:允许含其他bot字段的Googlebot访问
IIS Rewrite规则优化:允许Googlebot并阻止其他爬虫
问题背景
我需要配置IIS Rewrite规则,实现阻止所有包含bot、crawl或spider关键词的User Agent,但允许Googlebot正常访问。当前使用的规则存在误判:当Googlebot的UA字符串中包含其他位置的bot关键词(比如URL里的bot.html)时,请求会被错误拦截。
原规则代码:
<rule name="BotBlock" stopProcessing="true"> <match url=".*" /> <conditions> <add input="{HTTP_USER_AGENT}" pattern="^$|\b(?!.*googlebot.*\b)\w*(?:bot|crawl|spider)\w*" /> </conditions> <action type="CustomResponse" statusCode="403" statusReason="Forbidden" statusDescription="Forbidden" /> </rule>
问题示例
- 正常情况:
Googlebot/2.1 (+http://www.google.com)→ 被允许(符合预期) - 异常情况:
Googlebot/2.1 (+http://www.google.com/bot.html)→ 被拦截(不符合预期,因为是Googlebot) - 正常情况:
KHTML, like Gecko; compatible; bingbot→ 被拦截(符合预期)
解决方案
修改正则表达式,核心逻辑调整为:仅当UA不包含googlebot(不区分大小写)时,才匹配bot/crawl/spider关键词,同时保留拦截空UA的规则。
修改后的完整规则:
<rule name="BotBlock" stopProcessing="true"> <match url=".*" /> <conditions> <add input="{HTTP_USER_AGENT}" pattern="^$|^(?!.*googlebot).*\b(?:bot|crawl|spider)\b" ignoreCase="true" /> </conditions> <action type="CustomResponse" statusCode="403" statusReason="Forbidden" statusDescription="Forbidden" /> </rule>
规则说明
ignoreCase="true":开启大小写不敏感匹配,确保Googlebot、googlebot等变体都能被识别^(?!.*googlebot):正向否定预查,确保整个UA字符串中不存在googlebot关键词.*\b(?:bot|crawl|spider)\b:匹配包含bot/crawl/spider完整单词的UA^$:拦截空UA
验证效果
Googlebot/2.1 (+http://www.google.com/bot.html)→ 被允许(符合预期)KHTML, like Gecko; compatible; bingbot→ 被拦截(符合预期)- 空UA → 被拦截(符合预期)
Mozilla/5.0 (compatible; Baiduspider/2.0; +http://www.baidu.com/search/spider.html)→ 被拦截(符合预期)
内容的提问来源于stack exchange,提问作者VDWWD
相关产品推荐
相关产品推荐

