You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python requests设置allow_redirects=False仍跳转,如何仅获取robots.txt数据

根因分析

allow_redirects=False仅会阻止requests自动跟随3xx重定向,不会拦截服务端返回的3xx状态码。而raise_for_status()默认仅对4xx、5xx状态码抛出异常,3xx重定向响应不会被捕获,会直接传入extract_robots方法执行。

解决方案

在调用extract_robots前新增状态码、响应特征双重校验,仅符合robots.txt特征的响应才进入解析逻辑,修改后代码如下:

import urllib.parse
import requests
from requests import exceptions

for url in urls:
    robots_url = urllib.parse.urljoin(url, "robots.txt")
    try: 
        r = requests.get(robots_url, headers=headers, allow_redirects=False) 
        r.raise_for_status()
        # 校验1:仅处理200状态码响应,直接跳过3xx重定向
        if r.status_code != 200:
            continue
        # 可选校验2:确认响应为文本类型,避免拿到重定向HTML页面
        content_type = r.headers.get("Content-Type", "")
        if not content_type.startswith("text/plain"):
            continue
        # 可选校验3:确认响应路径为robots.txt,规避特殊路由规则干扰
        if not r.url.endswith("robots.txt"):
            continue
        extract_robots(r)
    except (exceptions.RequestException, exceptions.HTTPError, exceptions.Timeout) as err:
        handle_exeption(err)

如果需要兼容部分跳转到合法robots.txt的场景,可以补充逻辑:3xx状态码下先提取响应头Location字段,判断跳转目标为robots.txt文件时主动发起新请求拉取内容,否则直接跳过即可。

内容的提问来源于stack exchange,提问作者Pierre56

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 01:24:04