You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python的wget和requests包下载PDF文件时出现不完整问题

解决Python下载PDF时部分文件损坏/大小异常的问题

核心原因分析

你遇到的小文件无法打开的情况,大概率是服务器拦截了请求(返回错误页面/空白内容)、网络中断导致下载不完整,或者请求头不符合网站要求,服务器没有返回真实的PDF数据。


具体解决方案

1. 模拟浏览器请求头,避免被拦截

很多网站会检测请求的User-Agent、Accept等字段,默认的requests/wget请求头过于简单,容易被识别为爬虫,返回403页面(大小通常几KB)。给请求添加完整的浏览器头:

import requests
import pandas as pd

# 模拟Chrome浏览器的请求头,可根据实际浏览器调整
HEADERS = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    'Accept': 'application/pdf,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8',
    'Accept-Language': 'zh-CN,zh;q=0.8,en-US;q=0.5,en;q=0.3',
    'Referer': 'https://www.example.com/'  # 替换为PDF所在的网站域名
}

def download_pdf(url, save_path):
    try:
        # 开启流式下载,避免大文件占用过多内存
        response = requests.get(url, headers=HEADERS, stream=True, timeout=30, allow_redirects=True)
        # 主动抛出HTTP错误(比如403、404)
        response.raise_for_status()
        
        # 检查返回内容是否为PDF
        content_type = response.headers.get('Content-Type', '')
        if 'application/pdf' not in content_type:
            print(f"跳过 {url}:返回内容类型为 {content_type},非PDF")
            return False
        
        # 分块写入文件
        with open(save_path, 'wb') as f:
            for chunk in response.iter_content(chunk_size=8192):
                f.write(chunk)
        return True
    except Exception as e:
        print(f"下载失败 {url}:{str(e)}")
        return False

# 遍历DataFrame中的URL
df = pd.DataFrame({'pdf_url': ['https://www.example.pdf']})
for idx, row in df.iterrows():
    save_path = f"downloaded_pdf_{idx}.pdf"
    download_pdf(row['pdf_url'], save_path)

2. 增加重试机制,应对网络波动

网络不稳定或服务器临时限流会导致下载中断,只获取了部分内容。可以用重试逻辑确保下载完整:

from tenacity import retry, stop_after_attempt, wait_exponential

# 最多重试3次,每次等待时间指数增长(2s→4s→8s)
@retry(stop=stop_after_attempt(3), wait=wait_exponential(multiplier=1, min=2, max=10))
def download_with_retry(url, save_path):
    return download_pdf(url, save_path)

3. 验证下载文件的有效性

下载后主动检查文件是否可用,避免无效文件留存:

from PyPDF2 import PdfReader
import os

def is_valid_pdf(file_path):
    try:
        reader = PdfReader(file_path)
        # 能读取到页面数即为有效PDF
        return len(reader.pages) > 0
    except Exception:
        return False

# 下载后验证
for idx, row in df.iterrows():
    save_path = f"downloaded_pdf_{idx}.pdf"
    if download_with_retry(row['pdf_url'], save_path):
        if not is_valid_pdf(save_path):
            print(f"删除无效文件:{save_path}")
            os.remove(save_path)

4. 用异步工具提升批量下载效率

如果是大量PDF下载,推荐用httpx的异步客户端,同时处理多个请求,且兼容性更好:

import httpx
import asyncio

async def async_download_pdf(url, save_path):
    async with httpx.AsyncClient(headers=HEADERS) as client:
        response = await client.get(url, follow_redirects=True, timeout=30)
        response.raise_for_status()
        if 'application/pdf' not in response.headers.get('Content-Type', ''):
            return False
        with open(save_path, 'wb') as f:
            f.write(response.content)
    return True

# 批量异步下载
async def batch_download(df):
    tasks = []
    for idx, row in df.iterrows():
        save_path = f"async_pdf_{idx}.pdf"
        tasks.append(async_download_pdf(row['pdf_url'], save_path))
    await asyncio.gather(*tasks)

asyncio.run(batch_download(df))

内容的提问来源于stack exchange,提问作者S Shekhar Mishra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 14:32:02