You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python脚本从网站获取PDF速度过慢的优化方案求助

本地Python脚本下载PDF速度极慢的问题与优化需求

我编写了一个轻量Python脚本,用于打开托管PDF的网站、下载PDF并提取文本,选用requests和PyPDF2库。PyPDF2提取文本的效率尚可,但获取PDF数据的过程速度极慢。

以下是我测试的两段代码:

1. 基于urllib.request的原始代码

import urllib.request
from urllib.parse import urlparse
import time


url = "https://www.ohchr.org/sites/default/files/UDHR/Documents/UDHR_Translations/eng.pdf"

headers = {
      "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36",
      "Accept": "application/pdf",  # Indicating we want a PDF file
  }

# Extract the base domain from the URL to set as the Referer header
parsed_url = urlparse(url)
referer = f"{parsed_url.scheme}://{parsed_url.netloc}"  # Extract base domain (e.g., "https://example.com")

# Update the headers with dynamic Referer
headers["Referer"] = referer

start_time=time.time()
# Step 1: Fetch PDF content directly from the URL with headers
req = urllib.request.Request(url, headers=headers)
with urllib.request.urlopen(req) as response:
    pdf_data = response.read()
   
print(time.time() - start_time)

该代码下载耗时约65.54秒,而浏览器打开该URL速度极快。

2. 带urllib3重试逻辑的requests版本

import requests
import time
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

url = "https://www.ohchr.org/sites/default/files/UDHR/Documents/UDHR_Translations/eng.pdf"

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36",
    "Accept": "application/pdf",
    "Cache-Control": "no-cache",
    "Pragma": "no-cache",
}

start_time = time.time()

# Configure retries for requests
session = requests.Session()
retries = Retry(total=3, backoff_factor=0.3, status_forcelist=[500, 502, 503, 504])
adapter = HTTPAdapter(max_retries=retries)
session.mount("https://", adapter)

response = session.get(url, headers=headers, timeout=5)

if response.status_code == 200:
    pdf_data = response.content
    print(f"Time taken: {time.time() - start_time:.2f} seconds")
else:
    print(f"Failed to fetch the PDF. Status code: {response.status_code}")

该版本耗时约105.47秒。

两种方法均可下载PDF,但速度无法满足生产需求——例如联合国相关URL,浏览器1-2秒即可加载,脚本却耗时极久。我的网络稳定且速度快,调整User-Agent等请求头也无效。

更新情况

在Colab中运行代码速度很快,但本地Windows10环境下,优化后的代码仍耗时68.41秒。希望了解:

  • 本地配置可能存在的问题
  • 可行的优化方法、替代库或策略

内容的提问来源于stack exchange,提问作者R_Student

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 04:42:33