You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何阻止特定浏览器User Agent?解决网站数据采集Bot问题

应对固定时段Bot数据抓取的有效方案

最近被一个烦人的Bot折腾坏了——每天固定时段爬我的网站,不仅消耗大量带宽,还把Google Analytics的统计数据搞失真了。这Bot一开始用amazonaws的IP访问,后来换了其他主机,但User Agent居然没改!我最初尝试通过User Agent拦截但失败了,折腾一阵后终于找到可行方案,分享给大家:

最初的失败尝试

我一开始写了这样的Rewrite规则,期望返回503状态码阻止Bot,但没生效:

RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Ubuntu HeadlessChrome HeadlessChrome Safari/537.36
RewriteRule .* - [R=503,L]

最终生效的完整.htaccess配置

后来在社区大佬MrWhite的帮助下,我调整了规则并整合到现有配置中,现在Bot被成功拦截了。更新后的完整配置如下:

<LocationMatch .*>
<IfModule mod_security2.c>
SecRuleRemoveById 211170
SecRuleRemoveById 211180
</IfModule>
</LocationMatch>
Options +FollowSymlinks
Options -Indexes
<FilesMatch "(?i)((\.tpl|\.ini|\.log|(?<!robots)\.txt))">
Require all denied
</FilesMatch>
# SEO URL Settings
RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} "=Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Ubuntu HeadlessChrome HeadlessChrome Safari/537.36"
RewriteRule .* - [F]
RewriteBase /
RewriteRule ^sitemap.xml$ index.php?route=extension/feed/google_sitemap [L]
RewriteRule ^googlebase.xml$ index.php?route=extension/feed/google_base [L]
RewriteRule ^system/download/(.*) index.php?route=error/not_found [L]
RewriteCond %{REQUEST_FILENAME} !-f
RewriteCond %{REQUEST_FILENAME} !-d
RewriteCond %{REQUEST_URI} !.*\.(ico|gif|jpg|jpeg|png|js|css)
RewriteRule ^([^?]*) index.php?_route_=$1 [L,QSA]
<Files 403.shtml>
order allow,deny
allow from all
</Files>

关键调整说明

  • 将原来的[R=503,L]替换为[F],直接返回403 Forbidden状态码,比503更直接地拒绝Bot访问
  • 给User Agent匹配规则加上了双引号和等号"=...",确保是精确匹配,避免部分匹配导致的误拦截或漏拦截
  • 把Bot拦截规则放在了所有Rewrite规则的最前面,保证拦截逻辑优先执行

内容的提问来源于stack exchange,提问作者Jeof

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:16:18