如何使用python3下载指定网页链接中的PDF文件
Python3下载目标站点PDF的可行实现方案
前置注意事项
- 目标URL末尾的
#74l属于页内锚点标识,发起请求时可直接剔除,不影响资源正常拉取 - 该站点配置了基础反爬规则,无标识的裸请求大概率返回403拦截,必须携带模拟浏览器的请求头
- 拿到响应后优先校验
Content-Type字段是否为application/pdf,避免下载到错误提示页、跳转页等无效内容
方法1:基于Python标准库urllib实现(无第三方依赖)
无需额外安装任何依赖包,适合权限受限、无法随意安装第三方库的运行环境,参考代码:
import urllib.request target_url = "https://qingarchives.npm.edu.tw/index.php?act=Display/image/207469Zh18QEz" request_headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36", "Referer": "https://qingarchives.npm.edu.tw/" } req = urllib.request.Request(target_url, headers=request_headers) with urllib.request.urlopen(req, timeout=30) as response: if response.getheader("Content-Type") == "application/pdf": with open("archive_file.pdf", "wb") as f: f.write(response.read()) print("PDF下载完成") else: print("返回内容非PDF格式,下载失败")
方法2:基于requests第三方库实现(代码简洁易维护)
requests是Python生态最常用的HTTP请求工具,异常处理、编码适配逻辑更完善,使用前先执行pip install requests安装依赖,参考代码:
import requests target_url = "https://qingarchives.npm.edu.tw/index.php?act=Display/image/207469Zh18QEz" request_headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36", "Referer": "https://qingarchives.npm.edu.tw/" } response = requests.get(target_url, headers=request_headers, timeout=30) if response.status_code == 200 and response.headers.get("Content-Type") == "application/pdf": with open("archive_file.pdf", "wb") as f: f.write(response.content) print("PDF下载完成") else: print(f"下载失败,状态码:{response.status_code},响应类型:{response.headers.get('Content-Type')}")
方法3:基于浏览器自动化工具实现(适配动态校验场景)
如果前两种方法拿到的始终是HTML页面而非PDF资源,说明站点存在前端JS校验、动态生成资源链接的逻辑,此时可使用playwright模拟真实浏览器环境触发下载,使用前先执行pip install playwright && playwright install chromium安装依赖,参考代码:
from playwright.sync_api import sync_playwright target_url = "https://qingarchives.npm.edu.tw/index.php?act=Display/image/207469Zh18QEz#74l" save_path = "./archive_file.pdf" with sync_playwright() as p: browser = p.chromium.launch(headless=True) page = browser.new_page() with page.expect_download() as download_task: page.goto(target_url, timeout=60000) download = download_task.value download.save_as(save_path) browser.close() print(f"PDF已保存至路径:{save_path}")
内容的提问来源于stack exchange,提问作者chino
相关产品推荐
相关产品推荐

