You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何配置Nginx拦截特定爬虫但允许其访问robots.txt?

Fixing Nginx Bot Block to Allow Access to /robots.txt

Got it, let's break down why your modified config isn't working and get this sorted. The core issue here is Nginx's execution priority: the if block in the server context runs before any location matching happens. So even though you added the /robots.txt location first, the if ($block_ua) check triggers a 403 response before Nginx ever gets to process that specific location.

Here are two solid solutions to fix this:

Solution 1: Override the block variable in the /robots.txt location

This approach keeps your existing server-level if block but resets the block flag specifically for /robots.txt requests:

map $http_user_agent $block_ua {
    default 0;
    ~*yandexbot 1;
}
server {
    # Allow all bots to access robots.txt by resetting the block flag
    location = /robots.txt {
        set $block_ua 0;
        try_files $uri $uri/ /index.php?$args;
    }

    # Block bad bots for all other requests
    if ($block_ua) {
        return 403;
    }

    location / {
        try_files $uri $uri/ /index.php?$args;
    }

    # other location blocks below...
}

When a request hits /robots.txt, the set $block_ua 0 line overrides the value from the map, so the server-level if block won't trigger a 403.

Solution 2: Move the bot block logic into the root location

This is a cleaner approach that leverages Nginx's location matching priority (exact matches like = /robots.txt run first):

map $http_user_agent $block_ua {
    default 0;
    ~*yandexbot 1;
}
server {
    # Exact match for robots.txt - no blocking here
    location = /robots.txt {
        try_files $uri $uri/ /index.php?$args;
    }

    # Block bad bots only for requests that hit the root location
    location / {
        if ($block_ua) {
            return 403;
        }
        try_files $uri $uri/ /index.php?$args;
    }

    # other location blocks below...
}

Since /robots.txt matches the exact location first, it skips the blocking logic entirely. All other requests fall into the root location and get checked for the bot flag.

Testing & Applying the Fix

After updating your config, reload Nginx to apply changes:

sudo nginx -s reload

Then test with curl to verify:

# Should return 200 OK for robots.txt
curl -I -A "yandexbot" https://your-domain.com/robots.txt

# Should return 403 Forbidden for other paths
curl -I -A "yandexbot" https://your-domain.com/

内容的提问来源于stack exchange,提问作者Kok Hui

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:17:57