如何阻止特定浏览器User Agent?解决网站数据采集Bot问题
应对固定时段Bot数据抓取的有效方案
最近被一个烦人的Bot折腾坏了——每天固定时段爬我的网站,不仅消耗大量带宽,还把Google Analytics的统计数据搞失真了。这Bot一开始用amazonaws的IP访问,后来换了其他主机,但User Agent居然没改!我最初尝试通过User Agent拦截但失败了,折腾一阵后终于找到可行方案,分享给大家:
最初的失败尝试
我一开始写了这样的Rewrite规则,期望返回503状态码阻止Bot,但没生效:
RewriteEngine On RewriteCond %{HTTP_USER_AGENT} Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Ubuntu HeadlessChrome HeadlessChrome Safari/537.36 RewriteRule .* - [R=503,L]
最终生效的完整.htaccess配置
后来在社区大佬MrWhite的帮助下,我调整了规则并整合到现有配置中,现在Bot被成功拦截了。更新后的完整配置如下:
<LocationMatch .*> <IfModule mod_security2.c> SecRuleRemoveById 211170 SecRuleRemoveById 211180 </IfModule> </LocationMatch> Options +FollowSymlinks Options -Indexes <FilesMatch "(?i)((\.tpl|\.ini|\.log|(?<!robots)\.txt))"> Require all denied </FilesMatch> # SEO URL Settings RewriteEngine On RewriteCond %{HTTP_USER_AGENT} "=Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Ubuntu HeadlessChrome HeadlessChrome Safari/537.36" RewriteRule .* - [F] RewriteBase / RewriteRule ^sitemap.xml$ index.php?route=extension/feed/google_sitemap [L] RewriteRule ^googlebase.xml$ index.php?route=extension/feed/google_base [L] RewriteRule ^system/download/(.*) index.php?route=error/not_found [L] RewriteCond %{REQUEST_FILENAME} !-f RewriteCond %{REQUEST_FILENAME} !-d RewriteCond %{REQUEST_URI} !.*\.(ico|gif|jpg|jpeg|png|js|css) RewriteRule ^([^?]*) index.php?_route_=$1 [L,QSA] <Files 403.shtml> order allow,deny allow from all </Files>
关键调整说明
- 将原来的
[R=503,L]替换为[F],直接返回403 Forbidden状态码,比503更直接地拒绝Bot访问 - 给User Agent匹配规则加上了双引号和等号
"=...",确保是精确匹配,避免部分匹配导致的误拦截或漏拦截 - 把Bot拦截规则放在了所有Rewrite规则的最前面,保证拦截逻辑优先执行
内容的提问来源于stack exchange,提问作者Jeof
相关产品推荐
相关产品推荐

