使用Python Requests下载PDF后损坏无法打开的技术求助
问题描述
尝试用Python从以下两个链接下载PDF文件,但下载后的文件无法打开:
- https://fnet.bmfbovespa.com.br/fnet/publico/exibirDocumento?id=693676
- https://fnet.bmfbovespa.com.br/fnet/publico/downloadDocumento?id=693676
使用的requests代码如下:
import requests i = ["https://fnet.bmfbovespa.com.br/fnet/publico/exibirDocumento?id=693676", "https://fnet.bmfbovespa.com.br/fnet/publico/downloadDocumento?id=693676"] l =0 for k in i: l += 1 user_agent = "scrapping_script/1.0" headers = {'User-Agent': user_agent} download = requests.get(k, headers=headers) with open(f"/Users/renato/Documents/{l}.pdf", 'wb') as f: f.write(download.content)
已尝试过urllib以及修改请求头,但问题依旧,寻求解决建议。
解决建议
- 区分有效下载链接:第一个链接
exibirDocumento是网页展示页面,返回的是HTML内容而非PDF,保存成PDF自然无法打开,只需要使用downloadDocumento这个真正的下载链接即可。 - 维持会话状态:该网站可能需要Cookie验证才能正常下载,建议用
requests.Session()保持会话,先访问展示页面获取必要Cookie,再发起下载请求:import requests download_url = "https://fnet.bmfbovespa.com.br/fnet/publico/downloadDocumento?id=693676" view_url = "https://fnet.bmfbovespa.com.br/fnet/publico/exibirDocumento?id=693676" # 使用主流浏览器的User-Agent,避免被反爬拦截 user_agent = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" with requests.Session() as session: session.headers.update({'User-Agent': user_agent}) # 先访问展示页面获取会话Cookie session.get(view_url) # 发起下载请求 response = session.get(download_url) # 检查请求是否成功 if response.status_code == 200 and response.headers.get('Content-Type') == 'application/pdf': with open("/Users/renato/Documents/downloaded.pdf", 'wb') as f: f.write(response.content) else: print(f"请求失败,状态码:{response.status_code},内容类型:{response.headers.get('Content-Type')}") - 验证响应内容:下载前检查响应的
Content-Type是否为application/pdf,如果不是,说明返回的是错误页面或验证页面,需要排查请求头、会话是否符合网站要求。 - 替换User-Agent:避免使用自定义的
scrapping_script/1.0这类易被识别的UA,改用主流浏览器的UA字符串,降低被反爬机制拦截的概率。
内容的提问来源于stack exchange,提问作者Renato Krouss
相关产品推荐
相关产品推荐

