使用Python下载联合国决议PDF时的跳转问题求助
如何用Python正确下载联合国决议PDF
你遇到的问题是因为https://undocs.org/en/A/RES/76/307并非直接的PDF文件链接,而是一个HTML跳转页——它通过<meta>标签实现页面刷新跳转,而requests库默认只会自动处理HTTP 3xx状态码的重定向,不会解析HTML里的跳转指令,所以你拿到的是跳转页的HTML内容,不是PDF。
以下是两种可行的解决方案:
方法1:解析META跳转链接,自动获取PDF地址
用BeautifulSoup解析跳转页的HTML,提取<meta>标签中的跳转URL,拼接成绝对地址后再请求:
import requests from bs4 import BeautifulSoup # 初始跳转页URL initial_url = "https://undocs.org/en/A/RES/76/307" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } # 请求初始页面 response = requests.get(initial_url, headers=headers) soup = BeautifulSoup(response.text, "html.parser") # 提取meta刷新标签里的跳转URL refresh_tag = soup.find("meta", {"http-equiv": "refresh"}) if refresh_tag: # 解析CONTENT属性,格式为"1; URL=/tmp/xxx.html" content = refresh_tag.get("content") redirect_url = content.split("URL=")[-1] # 拼接成绝对URL absolute_redirect_url = f"https://undocs.org{redirect_url}" # 请求跳转后的页面,requests会自动处理3xx重定向到PDF地址 pdf_response = requests.get(absolute_redirect_url, headers=headers, allow_redirects=True) # 保存PDF with open("A_RES_76_307.pdf", "wb") as f: f.write(pdf_response.content) print("PDF下载完成") else: print("未找到跳转标签")
方法2:直接请求实际PDF地址
既然你已经获取了实际的PDF链接,直接请求该地址即可,但需添加请求头模拟浏览器,避免被服务器拦截:
import requests pdf_url = "https://documents-dds-ny.un.org/doc/UNDOC/GEN/N22/587/47/PDF/N2258747.pdf?OpenElement" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } response = requests.get(pdf_url, headers=headers) # 校验响应是否为有效PDF if response.status_code == 200 and response.headers.get("Content-Type") == "application/pdf": with open("A_RES_76_307.pdf", "wb") as f: f.write(response.content) print("PDF下载完成") else: print(f"下载失败,状态码:{response.status_code},内容类型:{response.headers.get('Content-Type')}")
注意事项
- 必须添加
User-Agent请求头:联合国文档服务器会拦截无标识的请求,模拟浏览器的请求头可避免被拒绝。 - 方法1适合批量下载多个决议的场景,无需手动查找每个文件的实际链接;方法2更直接,适合单个文件下载。
内容的提问来源于stack exchange,提问作者Jan
相关产品推荐
相关产品推荐

