You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python基于Selenium下载网站下拉选项对应REPORTS板块PDF问题

WSOP赛事PDF批量采集Selenium修复方案

问题根因

原有代码失效核心是4个共性问题:

  • 页面元素为异步联动加载,下拉选择后DOM会动态刷新,未做等待直接操作元素会触发找不到元素、点击无响应问题
  • GO按钮为submit类型的input标签,不是a链接,用LINK_TEXT定位完全无法匹配;REPORTS标签用固定序号XPath,页面tab顺序变动就会失效
  • 未做REPORTS板块内容加载等待,也没有遍历板块内所有PDF链接的逻辑,只能拿到默认加载的首个文件
  • 旧版Selenium的find_element_by_*系列方法在高版本已经完全弃用,直接调用会抛异常

修复后完整代码

from selenium import webdriver
from webdriver_manager.chrome import ChromeDriverManager
from selenium.webdriver.support.ui import Select
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.chrome.service import Service
import os
import time
import requests
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC

# 基础配置
pdf_save_path = os.path.join(os.getcwd(), "PDFs")
os.makedirs(pdf_save_path, exist_ok=True) # 不存在则自动创建存储文件夹

options = Options()
options.headless = True
# 配置Chrome自动下载PDF,不弹出预览窗口
options.add_experimental_option("prefs", {
    "download.default_directory": pdf_save_path,
    "download.prompt_for_download": False,
    "download.directory_upgrade": True,
    "plugins.always_open_pdf_externally": True,
    "safebrowsing.enabled": True
})

# 初始化浏览器
service = Service(ChromeDriverManager().install())
driver = webdriver.Chrome(service=service, options=options)
wait = WebDriverWait(driver, 15) # 统一显式等待时长15秒

driver.get('https://www.wsop.com/tournaments/results/')

# 三级联动下拉遍历:年份->赛事网格->具体赛事
# 先获取所有年份选项
year_select = Select(wait.until(EC.presence_of_element_located((By.ID, "CPHbody_aid"))))
year_options = [opt.get_attribute("value") for opt in year_select.options if opt.get_attribute("value")]

for year_val in year_options:
    # 重新定位下拉元素,避免DOM刷新后引用失效
    year_select = Select(wait.until(EC.presence_of_element_located((By.ID, "CPHbody_aid"))))
    year_select.select_by_value(year_val)
    time.sleep(1) # 等待联动选项刷新

    # 获取当前年份下所有赛事网格选项
    grid_select = Select(wait.until(EC.presence_of_element_located((By.ID, "CPHbody_grid"))))
    grid_options = [opt.get_attribute("value") for opt in grid_select.options if opt.get_attribute("value")]

    for grid_val in grid_options:
        grid_select = Select(wait.until(EC.presence_of_element_located((By.ID, "CPHbody_grid"))))
        grid_select.select_by_value(grid_val)
        time.sleep(1) # 等待赛事列表刷新

        # 获取当前网格下所有具体赛事选项
        tour_select = Select(wait.until(EC.presence_of_element_located((By.ID, "CPHbody_tid"))))
        tour_options = [opt.get_attribute("value") for opt in tour_select.options if opt.get_attribute("value")]

        for tour_val in tour_options:
            print(f"正在采集年份{year_val}、网格{grid_val}、赛事{tour_val}的报告")
            tour_select = Select(wait.until(EC.presence_of_element_located((By.ID, "CPHbody_tid"))))
            tour_select.select_by_value(tour_val)

            # 修复GO按钮点击问题:用name属性定位submit按钮,等可点击后操作
            go_btn = wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, 'input[name="GO"]')))
            go_btn.click()

            # 检查REPORTS标签是否存在,不存在直接跳过
            try:
                # 修复REPORTS定位问题:用文本模糊匹配,不依赖tab顺序
                report_tab = wait.until(EC.element_to_be_clickable((By.XPATH, '//a[contains(normalize-space(text()), "REPORTS")]')))
                report_tab.click()
            except:
                print(f"赛事{tour_val}无REPORTS板块,跳过")
                driver.get('https://www.wsop.com/tournaments/results/')
                time.sleep(1)
                continue

            # 修复单页PDF批量下载问题
            try:
                report_section = wait.until(EC.presence_of_element_located((By.ID, "reports")))
                # 提取板块内所有PDF链接
                pdf_links = report_section.find_elements(By.CSS_SELECTOR, 'a[href$=".pdf"]')
                if not pdf_links:
                    print(f"赛事{tour_val}REPORTS板块无PDF文件,跳过")
                    driver.get('https://www.wsop.com/tournaments/results/')
                    time.sleep(1)
                    continue

                # 带浏览器Cookie下载,避免鉴权拦截,效率比Selenium点击高
                cookies = {c["name"]: c["value"] for c in driver.get_cookies()}
                for link in pdf_links:
                    pdf_url = link.get_attribute("href")
                    pdf_name = pdf_url.split("/")[-1]
                    save_path = os.path.join(pdf_save_path, pdf_name)
                    if os.path.exists(save_path): # 已下载文件自动跳过
                        print(f"文件{pdf_name}已存在,跳过")
                        continue
                    resp = requests.get(pdf_url, cookies=cookies, timeout=20)
                    with open(save_path, "wb") as f:
                        f.write(resp.content)
                    print(f"已下载:{pdf_name}")
            except Exception as e:
                print(f"采集赛事{tour_val}PDF出错:{str(e)}")
            
            # 返回选择页准备下一个赛事采集
            driver.get('https://www.wsop.com/tournaments/results/')
            time.sleep(1)

driver.quit()
print("全量PDF采集完成")

关键逻辑说明

  • 联动下拉处理:每次选择上一级选项后,等待下一级选项加载完成再操作,且每次操作前重新定位元素,避免DOM刷新导致元素引用失效
  • 元素定位:所有可交互元素都加显式等待,元素加载完成、可点击后再操作,避免页面未渲染完导致的点击失效
  • 下载优化:用带Cookie的requests直接请求PDF链接,不会触发页面跳转、PDF预览等问题,下载速度远快于Selenium模拟点击;自动跳过已下载文件,避免重复采集
  • 容错处理:某个赛事无REPORTS板块、无PDF文件、请求出错时直接跳过,不会中断全量采集流程

后续解析提示

PDF采集完成后,可以用pdfplumber库读取PDF内的文本、表格内容,按年份、赛事ID、选手排名、筹码量、获奖金额等字段做规整后,直接转成pandas DataFrame即可开展后续数据分析。注意不同年份的赛事PDF格式可能存在差异,需要加字段匹配的容错逻辑。如果网络环境较差导致requests下载被拦截,可以把下载逻辑替换为Selenium点击PDF链接,提前配置的浏览器参数会自动触发下载,不会弹出预览窗口。

内容的提问来源于stack exchange,提问作者Python404

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 19:27:16