Python网页PDF爬取求助:从搜索结果页自动获取PDF文件
问题分析与解决方案
你的代码存在几个关键问题,导致无法批量获取PDF链接:
- 绝对XPATH脆弱性:你用了绝对路径定位元素,页面结构稍有变化就会失效,应该改用相对定位方式。
- 未遍历所有结果:只定位了单个下载按钮,没有处理搜索结果列表里的所有条目。
- 未处理分页:搜索结果可能分布在多页,原代码没有翻页逻辑。
- Selenium版本兼容问题:旧版的
executable_path参数已被弃用,建议使用Service类管理驱动。
改进后的代码实现
from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import TimeoutException, NoSuchElementException import requests import os # 初始化浏览器(Selenium 4+ 推荐用Service,自动管理chromedriver) driver = webdriver.Chrome(service=Service()) driver.get('https://www.coroners.nsw.gov.au/coronial-findings-search.html?searchtext=death%20in%20custody&searchYear=All') # 创建保存PDF的文件夹 save_dir = 'nsw_coroner_pdfs' os.makedirs(save_dir, exist_ok=True) def download_pdf(url, save_path): # 用你已实现的下载逻辑,这里用requests示例 response = requests.get(url, stream=True) with open(save_path, 'wb') as f: for chunk in response.iter_content(chunk_size=8192): f.write(chunk) try: while True: # 等待当前页所有结果行加载完成 WebDriverWait(driver, 20).until( EC.presence_of_all_elements_located((By.XPATH, "//div[contains(@class,'findings-row')]//a")) ) # 获取当前页所有PDF链接 pdf_links = driver.find_elements(By.XPATH, "//div[contains(@class,'findings-row')]//a") for link in pdf_links: pdf_url = link.get_attribute('href') # 确认是PDF链接 if pdf_url.endswith('.pdf'): # 提取文件名 filename = pdf_url.split('/')[-1] save_path = os.path.join(save_dir, filename) print(f"下载中: {filename}") download_pdf(pdf_url, save_path) # 尝试点击下一页 try: next_btn = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, "//a[contains(text(), 'Next')]")) ) # 检查下一页是否可点击(如果是最后一页,按钮可能禁用) if 'disabled' in next_btn.get_attribute('class'): break next_btn.click() # 等待页面加载 WebDriverWait(driver, 10).until( EC.staleness_of(pdf_links[0]) # 等待旧页面元素失效,确认新页加载 ) except (TimeoutException, NoSuchElementException): # 没有下一页,退出循环 break except TimeoutException: print("页面加载超时") finally: driver.quit()
关键改进点说明
- 相对定位:用
//div[contains(@class,'findings-row')]//a定位所有结果里的链接,不受页面整体结构变化影响。 - 分页处理:循环检测并点击"Next"按钮,直到没有下一页为止。
- Selenium兼容:使用
Service类管理Chrome驱动,无需手动指定驱动路径(Selenium 4.6+版本支持自动下载匹配的驱动)。 - 批量下载:遍历所有PDF链接,调用你已实现的下载逻辑(示例中用
requests实现基础下载)。
注意事项
- 确保安装了最新版Selenium:
pip install --upgrade selenium - 网站可能有反爬机制,建议添加适当的等待时间,避免请求过于频繁。
- 如果部分链接不是直接指向PDF,而是跳转至详情页,需要额外添加逻辑:点击进入详情页后,定位页面内的PDF下载链接再提取。
内容的提问来源于stack exchange,提问作者Rokit87
相关产品推荐
相关产品推荐

