使用Python的wget和requests包下载PDF文件时出现不完整问题
解决Python下载PDF时部分文件损坏/大小异常的问题
核心原因分析
你遇到的小文件无法打开的情况,大概率是服务器拦截了请求(返回错误页面/空白内容)、网络中断导致下载不完整,或者请求头不符合网站要求,服务器没有返回真实的PDF数据。
具体解决方案
1. 模拟浏览器请求头,避免被拦截
很多网站会检测请求的User-Agent、Accept等字段,默认的requests/wget请求头过于简单,容易被识别为爬虫,返回403页面(大小通常几KB)。给请求添加完整的浏览器头:
import requests import pandas as pd # 模拟Chrome浏览器的请求头,可根据实际浏览器调整 HEADERS = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept': 'application/pdf,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8', 'Accept-Language': 'zh-CN,zh;q=0.8,en-US;q=0.5,en;q=0.3', 'Referer': 'https://www.example.com/' # 替换为PDF所在的网站域名 } def download_pdf(url, save_path): try: # 开启流式下载,避免大文件占用过多内存 response = requests.get(url, headers=HEADERS, stream=True, timeout=30, allow_redirects=True) # 主动抛出HTTP错误(比如403、404) response.raise_for_status() # 检查返回内容是否为PDF content_type = response.headers.get('Content-Type', '') if 'application/pdf' not in content_type: print(f"跳过 {url}:返回内容类型为 {content_type},非PDF") return False # 分块写入文件 with open(save_path, 'wb') as f: for chunk in response.iter_content(chunk_size=8192): f.write(chunk) return True except Exception as e: print(f"下载失败 {url}:{str(e)}") return False # 遍历DataFrame中的URL df = pd.DataFrame({'pdf_url': ['https://www.example.pdf']}) for idx, row in df.iterrows(): save_path = f"downloaded_pdf_{idx}.pdf" download_pdf(row['pdf_url'], save_path)
2. 增加重试机制,应对网络波动
网络不稳定或服务器临时限流会导致下载中断,只获取了部分内容。可以用重试逻辑确保下载完整:
from tenacity import retry, stop_after_attempt, wait_exponential # 最多重试3次,每次等待时间指数增长(2s→4s→8s) @retry(stop=stop_after_attempt(3), wait=wait_exponential(multiplier=1, min=2, max=10)) def download_with_retry(url, save_path): return download_pdf(url, save_path)
3. 验证下载文件的有效性
下载后主动检查文件是否可用,避免无效文件留存:
from PyPDF2 import PdfReader import os def is_valid_pdf(file_path): try: reader = PdfReader(file_path) # 能读取到页面数即为有效PDF return len(reader.pages) > 0 except Exception: return False # 下载后验证 for idx, row in df.iterrows(): save_path = f"downloaded_pdf_{idx}.pdf" if download_with_retry(row['pdf_url'], save_path): if not is_valid_pdf(save_path): print(f"删除无效文件:{save_path}") os.remove(save_path)
4. 用异步工具提升批量下载效率
如果是大量PDF下载,推荐用httpx的异步客户端,同时处理多个请求,且兼容性更好:
import httpx import asyncio async def async_download_pdf(url, save_path): async with httpx.AsyncClient(headers=HEADERS) as client: response = await client.get(url, follow_redirects=True, timeout=30) response.raise_for_status() if 'application/pdf' not in response.headers.get('Content-Type', ''): return False with open(save_path, 'wb') as f: f.write(response.content) return True # 批量异步下载 async def batch_download(df): tasks = [] for idx, row in df.iterrows(): save_path = f"async_pdf_{idx}.pdf" tasks.append(async_download_pdf(row['pdf_url'], save_path)) await asyncio.gather(*tasks) asyncio.run(batch_download(df))
内容的提问来源于stack exchange,提问作者S Shekhar Mishra
相关产品推荐
相关产品推荐

