You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用requests.history仍无法下载两次重定向的PDF文件求助

解决TI文档链接重定向下载失败的问题

兄弟,我太懂这种折腾一天的崩溃感了!TI的这些文档链接确实有点坑,看似简单的跳转背后藏着小门道,给你两个靠谱的解决方案:

方案一:正确处理重定向(带上必要请求头)

TI的服务器会校验请求的合法性,如果你直接用默认的requests请求,很可能因为缺少浏览器标识被拦截,导致重定向失败。试试给请求加上完整的请求头,模拟真实浏览器访问:

import requests
from bs4 import BeautifulSoup

# 替换成你筛选出的目标PDF链接列表
pdf_links = ['http://www.ti.com/lit/sboa263']

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8'
}

for link in pdf_links:
    try:
        # 显式允许重定向(requests默认开启,这里声明更清晰)
        response = requests.get(link, headers=headers, allow_redirects=True, timeout=10)
        # 校验响应是否为PDF资源
        if 'application/pdf' in response.headers.get('Content-Type', ''):
            # 提取文件名,兼容URL末尾是否带.pdf的情况
            filename = response.url.split('/')[-1] if response.url.endswith('.pdf') else f"{response.url.split('/')[-1]}.pdf"
            with open(filename, 'wb') as f:
                f.write(response.content)
            print(f"成功下载:{filename}")
        else:
            print(f"链接{link}未跳转到有效PDF资源")
    except Exception as e:
        print(f"下载{link}失败:{str(e)}")

方案二:直接构造真实PDF URL(跳过重定向)

TI的文档链接有固定规律:原链接是http://www.ti.com/lit/[文档号],对应的真实PDF地址是https://www.ti.com/lit/pdf/[文档号],直接构造这个URL请求,完全不用处理重定向,效率更高也更稳定:

import requests

# 替换成你筛选出的目标PDF链接列表
pdf_links = ['http://www.ti.com/lit/sboa263']
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}

for link in pdf_links:
    # 从原链接提取文档号
    doc_id = link.split('/')[-1]
    # 构造真实PDF下载地址
    real_pdf_url = f'https://www.ti.com/lit/pdf/{doc_id}'
    try:
        response = requests.get(real_pdf_url, headers=headers, timeout=10)
        if response.status_code == 200 and 'application/pdf' in response.headers.get('Content-Type', ''):
            filename = f'{doc_id}.pdf'
            with open(filename, 'wb') as f:
                f.write(response.content)
            print(f"成功下载:{filename}")
        else:
            print(f"构造的URL{real_pdf_url}无效或无法访问")
    except Exception as e:
        print(f"下载{real_pdf_url}失败:{str(e)}")

额外小提醒

  • 记得把代码里的pdf_links替换成你用正则筛选出来的实际链接列表
  • 如果需要批量下载,建议在循环里加个time.sleep(1)的延时,避免触发TI的反爬机制导致IP被封禁
  • 优先选方案二,因为直接构造URL不受重定向规则变化影响,稳定性拉满

内容的提问来源于stack exchange,提问作者jar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:34:19