You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页PDF爬取求助:从搜索结果页自动获取PDF文件

问题分析与解决方案

你的代码存在几个关键问题,导致无法批量获取PDF链接:

  1. 绝对XPATH脆弱性:你用了绝对路径定位元素,页面结构稍有变化就会失效,应该改用相对定位方式。
  2. 未遍历所有结果:只定位了单个下载按钮,没有处理搜索结果列表里的所有条目。
  3. 未处理分页:搜索结果可能分布在多页,原代码没有翻页逻辑。
  4. Selenium版本兼容问题:旧版的executable_path参数已被弃用,建议使用Service类管理驱动。

改进后的代码实现

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import TimeoutException, NoSuchElementException
import requests
import os

# 初始化浏览器(Selenium 4+ 推荐用Service,自动管理chromedriver)
driver = webdriver.Chrome(service=Service())
driver.get('https://www.coroners.nsw.gov.au/coronial-findings-search.html?searchtext=death%20in%20custody&searchYear=All')

# 创建保存PDF的文件夹
save_dir = 'nsw_coroner_pdfs'
os.makedirs(save_dir, exist_ok=True)

def download_pdf(url, save_path):
    # 用你已实现的下载逻辑,这里用requests示例
    response = requests.get(url, stream=True)
    with open(save_path, 'wb') as f:
        for chunk in response.iter_content(chunk_size=8192):
            f.write(chunk)

try:
    while True:
        # 等待当前页所有结果行加载完成
        WebDriverWait(driver, 20).until(
            EC.presence_of_all_elements_located((By.XPATH, "//div[contains(@class,'findings-row')]//a"))
        )
        
        # 获取当前页所有PDF链接
        pdf_links = driver.find_elements(By.XPATH, "//div[contains(@class,'findings-row')]//a")
        for link in pdf_links:
            pdf_url = link.get_attribute('href')
            # 确认是PDF链接
            if pdf_url.endswith('.pdf'):
                # 提取文件名
                filename = pdf_url.split('/')[-1]
                save_path = os.path.join(save_dir, filename)
                print(f"下载中: {filename}")
                download_pdf(pdf_url, save_path)
        
        # 尝试点击下一页
        try:
            next_btn = WebDriverWait(driver, 10).until(
                EC.element_to_be_clickable((By.XPATH, "//a[contains(text(), 'Next')]"))
            )
            # 检查下一页是否可点击(如果是最后一页,按钮可能禁用)
            if 'disabled' in next_btn.get_attribute('class'):
                break
            next_btn.click()
            # 等待页面加载
            WebDriverWait(driver, 10).until(
                EC.staleness_of(pdf_links[0])  # 等待旧页面元素失效,确认新页加载
            )
        except (TimeoutException, NoSuchElementException):
            # 没有下一页,退出循环
            break

except TimeoutException:
    print("页面加载超时")
finally:
    driver.quit()

关键改进点说明

  • 相对定位:用//div[contains(@class,'findings-row')]//a定位所有结果里的链接,不受页面整体结构变化影响。
  • 分页处理:循环检测并点击"Next"按钮,直到没有下一页为止。
  • Selenium兼容:使用Service类管理Chrome驱动,无需手动指定驱动路径(Selenium 4.6+版本支持自动下载匹配的驱动)。
  • 批量下载:遍历所有PDF链接,调用你已实现的下载逻辑(示例中用requests实现基础下载)。

注意事项

  • 确保安装了最新版Selenium:pip install --upgrade selenium
  • 网站可能有反爬机制,建议添加适当的等待时间,避免请求过于频繁。
  • 如果部分链接不是直接指向PDF,而是跳转至详情页,需要额外添加逻辑:点击进入详情页后,定位页面内的PDF下载链接再提取。

内容的提问来源于stack exchange,提问作者Rokit87

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 08:40:00