You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用python3下载指定网页链接中的PDF文件

Python3下载目标站点PDF的可行实现方案

前置注意事项

  • 目标URL末尾的#74l属于页内锚点标识,发起请求时可直接剔除,不影响资源正常拉取
  • 该站点配置了基础反爬规则,无标识的裸请求大概率返回403拦截,必须携带模拟浏览器的请求头
  • 拿到响应后优先校验Content-Type字段是否为application/pdf,避免下载到错误提示页、跳转页等无效内容

方法1:基于Python标准库urllib实现(无第三方依赖)

无需额外安装任何依赖包,适合权限受限、无法随意安装第三方库的运行环境,参考代码:

import urllib.request

target_url = "https://qingarchives.npm.edu.tw/index.php?act=Display/image/207469Zh18QEz"
request_headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36",
    "Referer": "https://qingarchives.npm.edu.tw/"
}

req = urllib.request.Request(target_url, headers=request_headers)
with urllib.request.urlopen(req, timeout=30) as response:
    if response.getheader("Content-Type") == "application/pdf":
        with open("archive_file.pdf", "wb") as f:
            f.write(response.read())
        print("PDF下载完成")
    else:
        print("返回内容非PDF格式,下载失败")

方法2:基于requests第三方库实现(代码简洁易维护)

requests是Python生态最常用的HTTP请求工具,异常处理、编码适配逻辑更完善,使用前先执行pip install requests安装依赖,参考代码:

import requests

target_url = "https://qingarchives.npm.edu.tw/index.php?act=Display/image/207469Zh18QEz"
request_headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36",
    "Referer": "https://qingarchives.npm.edu.tw/"
}

response = requests.get(target_url, headers=request_headers, timeout=30)
if response.status_code == 200 and response.headers.get("Content-Type") == "application/pdf":
    with open("archive_file.pdf", "wb") as f:
        f.write(response.content)
    print("PDF下载完成")
else:
    print(f"下载失败,状态码:{response.status_code},响应类型:{response.headers.get('Content-Type')}")

方法3:基于浏览器自动化工具实现(适配动态校验场景)

如果前两种方法拿到的始终是HTML页面而非PDF资源,说明站点存在前端JS校验、动态生成资源链接的逻辑,此时可使用playwright模拟真实浏览器环境触发下载,使用前先执行pip install playwright && playwright install chromium安装依赖,参考代码:

from playwright.sync_api import sync_playwright

target_url = "https://qingarchives.npm.edu.tw/index.php?act=Display/image/207469Zh18QEz#74l"
save_path = "./archive_file.pdf"

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    with page.expect_download() as download_task:
        page.goto(target_url, timeout=60000)
    download = download_task.value
    download.save_as(save_path)
    browser.close()
    print(f"PDF已保存至路径:{save_path}")

内容的提问来源于stack exchange,提问作者chino

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 14:48:21