You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Fangraphs动态数据爬取问题:Selenium代码无法更新下拉框值

Fangraphs WPA Inquirer 爬取问题解决

问题描述

编程基础薄弱,仅了解scrape概念,借助ChatGPT生成Selenium代码,意图爬取Fangraphs网站WPA Inquirer工具中,由Run Environment、Base Situation、Inning、Outs、Run Differential这5个下拉框控制的灰色背景区域的Leverage Index等动态数据。代码执行无报错,但下拉框值未实际更新,导致Home/Away胜率及LI值始终不变,无法获取目标数据。

原因分析

原代码通过JS直接修改下拉框输入框的value属性,但该网站的下拉框是自定义组件(非原生<select>标签),仅修改value不会触发页面的变更事件(如onChange),因此页面不会重新计算并更新Leverage Index等数据。

修正后的代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import pandas as pd

# 初始化浏览器
driver = webdriver.Chrome()
url = 'https://www.fangraphs.com/tools/wpa-inquirer'
driver.get(url)
wait = WebDriverWait(driver, 10)

def select_dropdown(dropdown_id, target_value):
    # 点击触发下拉框展开
    dropdown_trigger = wait.until(EC.element_to_be_clickable((By.ID, dropdown_id)))
    dropdown_trigger.click()
    
    # 找到对应选项并点击
    option_xpath = f"//div[@id='{dropdown_id}_DropDown']//li[text()='{target_value}']"
    target_option = wait.until(EC.element_to_be_clickable((By.XPATH, option_xpath)))
    target_option.click()

def get_leverage_index(run_env, base_situation, inning, outs, run_differential):
    select_dropdown('rcbRun', run_env)
    select_dropdown('rcbBase', base_situation)
    select_dropdown('rcbInning', inning)
    select_dropdown('rcbOuts', outs)
    select_dropdown('rcbScore', run_differential)
    
    # 等待数据更新后再获取
    leverage_index = wait.until(EC.visibility_of_element_located(
        (By.XPATH, '//td[text()="Leverage Index"]/following-sibling::td')
    )).text
    return leverage_index

data = []

# 定义所有要遍历的参数值
run_env_values = ['3.0', '3.5', '4.0', '4.5', '5.0', '5.5', '6.0', '6.5']
base_situation_values = ['_ _ _', '1 _ _', '_ 2 _', '1 2 _', '_ _ 3', '1 _ 3', '_ 2 3', '1 2 3']
inning_values = ['1 (Top)', '1 (Bottom)', '2 (Top)', '2 (Bottom)', '3 (Top)', '3 (Bottom)',
                 '4 (Top)', '4 (Bottom)', '5 (Top)', '5 (Bottom)', '6 (Top)', '6 (Bottom)',
                 '7 (Top)', '7 (Bottom)', '8 (Top)', '8 (Bottom)', '>= 9 (Top)', '>= 9 (Bottom)']
outs_values = ['0', '1', '2']
run_differential_values = ['-10', '-9', '-8', '-7', '-6', '-5', '-4', '-3', '-2', '-1', '0', '1', '2', '3', '4', '5', '6', '7', '8', '9', '10']

total_count = len(run_env_values) * len(base_situation_values) * len(inning_values) * len(outs_values) * len(run_differential_values)
progress_count = 0

# 遍历所有参数组合
for run_env in run_env_values:
    for base_situation in base_situation_values:
        for inning in inning_values:
            for outs in outs_values:
                for run_differential in run_differential_values:
                    progress_count += 1
                    print(f'({progress_count}/{total_count})')
                    try:
                        leverage_index = get_leverage_index(run_env, base_situation, inning, outs, run_differential)
                        data.append([run_env, base_situation, inning, outs, run_differential, leverage_index])
                    except Exception as e:
                        print(f'获取数据失败: {e}')
                        data.append([run_env, base_situation, inning, outs, run_differential, '获取失败'])

# 关闭浏览器并保存数据
driver.quit()

df = pd.DataFrame(data, columns=['Run Environment', 'Base Situation', 'Inning', 'Outs', 'Run Differential', 'Leverage Index'])
df.to_excel('leverage_index_data.xlsx', index=False)

关键修改说明

  • 模拟真实交互:不再直接修改输入框value,而是点击下拉框触发展开,再点击目标选项,完全模拟用户操作,确保触发页面的变更事件。
  • 添加显式等待:使用WebDriverWait等待元素可点击/可见,避免因页面加载延迟导致的元素定位失败。
  • 异常处理:添加try-except块,避免单个参数组合出错导致整个程序中断。
  • 等待数据更新:获取Leverage Index前等待元素可见,确保页面已完成数据计算。

内容的提问来源于stack exchange,提问作者BlueBulbLight

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 21:45:25