如何配置Nginx拦截特定爬虫但允许其访问robots.txt?
Got it, let's break down why your modified config isn't working and get this sorted. The core issue here is Nginx's execution priority: the if block in the server context runs before any location matching happens. So even though you added the /robots.txt location first, the if ($block_ua) check triggers a 403 response before Nginx ever gets to process that specific location.
Here are two solid solutions to fix this:
Solution 1: Override the block variable in the /robots.txt location
This approach keeps your existing server-level if block but resets the block flag specifically for /robots.txt requests:
map $http_user_agent $block_ua { default 0; ~*yandexbot 1; } server { # Allow all bots to access robots.txt by resetting the block flag location = /robots.txt { set $block_ua 0; try_files $uri $uri/ /index.php?$args; } # Block bad bots for all other requests if ($block_ua) { return 403; } location / { try_files $uri $uri/ /index.php?$args; } # other location blocks below... }
When a request hits /robots.txt, the set $block_ua 0 line overrides the value from the map, so the server-level if block won't trigger a 403.
Solution 2: Move the bot block logic into the root location
This is a cleaner approach that leverages Nginx's location matching priority (exact matches like = /robots.txt run first):
map $http_user_agent $block_ua { default 0; ~*yandexbot 1; } server { # Exact match for robots.txt - no blocking here location = /robots.txt { try_files $uri $uri/ /index.php?$args; } # Block bad bots only for requests that hit the root location location / { if ($block_ua) { return 403; } try_files $uri $uri/ /index.php?$args; } # other location blocks below... }
Since /robots.txt matches the exact location first, it skips the blocking logic entirely. All other requests fall into the root location and get checked for the bot flag.
Testing & Applying the Fix
After updating your config, reload Nginx to apply changes:
sudo nginx -s reload
Then test with curl to verify:
# Should return 200 OK for robots.txt curl -I -A "yandexbot" https://your-domain.com/robots.txt # Should return 403 Forbidden for other paths curl -I -A "yandexbot" https://your-domain.com/
内容的提问来源于stack exchange,提问作者Kok Hui

