You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python requests下载指定URL PDF失败,求有效替代方案

解决PDF下载损坏问题的可行方案

方案1:补充完整请求头字段

很多网站会校验除User-Agent外的其他请求头字段,比如Accept、Referer,补充后可避免返回非PDF内容:

import requests

url = 'https://www.adgm.com/documents/operating-in-adgm/ongoing-obligation/enforcement/alpha-development-middle-east-ltd-penalty-notice-redacted.pdf?la=en&hash=5EA2DA7D1492D105375580EEF2FB088F'
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.114 Safari/537.36',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8',
    'Accept-Language': 'en-US,en;q=0.5',
    'Referer': 'https://www.adgm.com/',
    'Connection': 'keep-alive'
}

response = requests.get(url, headers=headers, timeout=10)
# 先验证请求状态与返回内容类型
if response.status_code == 200:
    if 'application/pdf' in response.headers.get('Content-Type', ''):
        with open('sample.pdf', 'wb') as f:
            f.write(response.content)
    else:
        print("返回内容非PDF,可能被反爬拦截")
else:
    print(f"请求失败,状态码:{response.status_code}")

方案2:使用Session保持会话

部分网站会通过会话Cookie验证请求合法性,用requests.Session可自动管理Cookie:

import requests

url = 'https://www.adgm.com/documents/operating-in-adgm/ongoing-obligation/enforcement/alpha-development-middle-east-ltd-penalty-notice-redacted.pdf?la=en&hash=5EA2DA7D1492D105375580EEF2FB088F'
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.114 Safari/537.36',
    'Referer': 'https://www.adgm.com/'
}

session = requests.Session()
# 先访问网站主页获取必要Cookie
session.get('https://www.adgm.com/', headers=headers)
# 再发起PDF下载请求
response = session.get(url, headers=headers, timeout=10)

if response.status_code == 200 and 'application/pdf' in response.headers.get('Content-Type', ''):
    with open('sample.pdf', 'wb') as f:
        f.write(response.content)
else:
    print("下载失败,返回内容异常")

方案3:用Playwright模拟真实浏览器下载

如果网站存在JS反爬或渲染验证,Playwright可模拟真实浏览器行为,确保获取完整PDF:

先安装依赖:

pip install playwright
playwright install chrome

编写下载代码:

from playwright.sync_api import sync_playwright
import time

url = 'https://www.adgm.com/documents/operating-in-adgm/ongoing-obligation/enforcement/alpha-development-middle-east-ltd-penalty-notice-redacted.pdf?la=en&hash=5EA2DA7D1492D105375580EEF2FB088F'

with sync_playwright() as p:
    browser = p.chromium.launch(headless=False)  # 可设置headless=True后台运行
    page = browser.new_page()
    # 指定下载路径
    page.set_download_path("./")
    # 访问目标URL触发下载
    page.goto(url)
    # 根据文件大小调整等待时间
    time.sleep(5)
    browser.close()

内容的提问来源于stack exchange,提问作者Lewis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 19:32:42