使用requests.history仍无法下载两次重定向的PDF文件求助
解决TI文档链接重定向下载失败的问题
兄弟,我太懂这种折腾一天的崩溃感了!TI的这些文档链接确实有点坑,看似简单的跳转背后藏着小门道,给你两个靠谱的解决方案:
方案一:正确处理重定向(带上必要请求头)
TI的服务器会校验请求的合法性,如果你直接用默认的requests请求,很可能因为缺少浏览器标识被拦截,导致重定向失败。试试给请求加上完整的请求头,模拟真实浏览器访问:
import requests from bs4 import BeautifulSoup # 替换成你筛选出的目标PDF链接列表 pdf_links = ['http://www.ti.com/lit/sboa263'] headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8' } for link in pdf_links: try: # 显式允许重定向(requests默认开启,这里声明更清晰) response = requests.get(link, headers=headers, allow_redirects=True, timeout=10) # 校验响应是否为PDF资源 if 'application/pdf' in response.headers.get('Content-Type', ''): # 提取文件名,兼容URL末尾是否带.pdf的情况 filename = response.url.split('/')[-1] if response.url.endswith('.pdf') else f"{response.url.split('/')[-1]}.pdf" with open(filename, 'wb') as f: f.write(response.content) print(f"成功下载:{filename}") else: print(f"链接{link}未跳转到有效PDF资源") except Exception as e: print(f"下载{link}失败:{str(e)}")
方案二:直接构造真实PDF URL(跳过重定向)
TI的文档链接有固定规律:原链接是http://www.ti.com/lit/[文档号],对应的真实PDF地址是https://www.ti.com/lit/pdf/[文档号],直接构造这个URL请求,完全不用处理重定向,效率更高也更稳定:
import requests # 替换成你筛选出的目标PDF链接列表 pdf_links = ['http://www.ti.com/lit/sboa263'] headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } for link in pdf_links: # 从原链接提取文档号 doc_id = link.split('/')[-1] # 构造真实PDF下载地址 real_pdf_url = f'https://www.ti.com/lit/pdf/{doc_id}' try: response = requests.get(real_pdf_url, headers=headers, timeout=10) if response.status_code == 200 and 'application/pdf' in response.headers.get('Content-Type', ''): filename = f'{doc_id}.pdf' with open(filename, 'wb') as f: f.write(response.content) print(f"成功下载:{filename}") else: print(f"构造的URL{real_pdf_url}无效或无法访问") except Exception as e: print(f"下载{real_pdf_url}失败:{str(e)}")
额外小提醒
- 记得把代码里的
pdf_links替换成你用正则筛选出来的实际链接列表 - 如果需要批量下载,建议在循环里加个
time.sleep(1)的延时,避免触发TI的反爬机制导致IP被封禁 - 优先选方案二,因为直接构造URL不受重定向规则变化影响,稳定性拉满
内容的提问来源于stack exchange,提问作者jar
相关产品推荐
相关产品推荐

