Python脚本从网站获取PDF速度过慢的优化方案求助
本地Python脚本下载PDF速度极慢的问题与优化需求
我编写了一个轻量Python脚本,用于打开托管PDF的网站、下载PDF并提取文本,选用requests和PyPDF2库。PyPDF2提取文本的效率尚可,但获取PDF数据的过程速度极慢。
以下是我测试的两段代码:
1. 基于urllib.request的原始代码
import urllib.request from urllib.parse import urlparse import time url = "https://www.ohchr.org/sites/default/files/UDHR/Documents/UDHR_Translations/eng.pdf" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36", "Accept": "application/pdf", # Indicating we want a PDF file } # Extract the base domain from the URL to set as the Referer header parsed_url = urlparse(url) referer = f"{parsed_url.scheme}://{parsed_url.netloc}" # Extract base domain (e.g., "https://example.com") # Update the headers with dynamic Referer headers["Referer"] = referer start_time=time.time() # Step 1: Fetch PDF content directly from the URL with headers req = urllib.request.Request(url, headers=headers) with urllib.request.urlopen(req) as response: pdf_data = response.read() print(time.time() - start_time)
该代码下载耗时约65.54秒,而浏览器打开该URL速度极快。
2. 带urllib3重试逻辑的requests版本
import requests import time from requests.adapters import HTTPAdapter from urllib3.util.retry import Retry url = "https://www.ohchr.org/sites/default/files/UDHR/Documents/UDHR_Translations/eng.pdf" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36", "Accept": "application/pdf", "Cache-Control": "no-cache", "Pragma": "no-cache", } start_time = time.time() # Configure retries for requests session = requests.Session() retries = Retry(total=3, backoff_factor=0.3, status_forcelist=[500, 502, 503, 504]) adapter = HTTPAdapter(max_retries=retries) session.mount("https://", adapter) response = session.get(url, headers=headers, timeout=5) if response.status_code == 200: pdf_data = response.content print(f"Time taken: {time.time() - start_time:.2f} seconds") else: print(f"Failed to fetch the PDF. Status code: {response.status_code}")
该版本耗时约105.47秒。
两种方法均可下载PDF,但速度无法满足生产需求——例如联合国相关URL,浏览器1-2秒即可加载,脚本却耗时极久。我的网络稳定且速度快,调整User-Agent等请求头也无效。
更新情况
在Colab中运行代码速度很快,但本地Windows10环境下,优化后的代码仍耗时68.41秒。希望了解:
- 本地配置可能存在的问题
- 可行的优化方法、替代库或策略
内容的提问来源于stack exchange,提问作者R_Student
相关产品推荐
相关产品推荐

