You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

遍历网页不同选择器并将数据存入统一大DataFrame的问题

解决方案

要实现抓取所有Game Recency选项的数据并新增对应列,需在原有代码基础上嵌套遍历Game Recency下拉选项,具体修改如下:

关键修改点

  • 定位Game Recency下拉菜单,获取所有可选项
  • 嵌套循环:先遍历每个Recency选项,再遍历各位置表格
  • 为每个数据帧新增Game_Recency列,记录当前选中的时间范围
  • 优化等待逻辑,确保元素加载完成后再操作

修改后的完整代码

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import Select
import time as t
import pandas as pd 

pd.set_option('display.max_columns', None)
pd.set_option('display.max_colwidth', None)
big_df = pd.DataFrame()
chrome_options = Options()
chrome_options.add_argument("--no-sandbox")
chrome_options.add_argument('disable-notifications')
chrome_options.add_argument("window-size=1280,720")

webdriver_service = Service(r'chromedriver\chromedriver') ## 替换为你的chromedriver路径
driver = webdriver.Chrome(service=webdriver_service, options=chrome_options)
wait = WebDriverWait(driver, 20)
url = "https://www.fantasypros.com/daily-fantasy/nba/fanduel-defense-vs-position.php"
driver.get(url)
t.sleep(5)  # 可根据网络情况调整初始等待时间

# 定位Game Recency下拉菜单并获取所有选项
recency_select = wait.until(EC.presence_of_element_located((By.ID, 'game-recency')))
select = Select(recency_select)
recency_options = [option.text.strip() for option in select.options]

# 嵌套遍历:先遍历每个Game Recency选项,再遍历各位置表格
for recency in recency_options:
    select.select_by_visible_text(recency)
    print(f'已选中Game Recency: {recency}')
    t.sleep(2)
    
    # 获取所有位置选项
    tables_list = wait.until(EC.presence_of_all_elements_located((By.XPATH, '//ul[@class="pills pos-filter pull-left"]/li')))
    for x in tables_list:
        x.click()
        print(f'  已选中位置: {x.text}')
        t.sleep(2)
        table = wait.until(EC.element_to_be_clickable((By.XPATH, '//table[@id="data-table"]')))
        df = pd.read_html(table.get_attribute('outerHTML'))[0]
        df['Category'] = x.text.strip()
        df['Game_Recency'] = recency  # 新增Game Recency标记列
        big_df = pd.concat([big_df, df], axis=0, ignore_index=True)
        print(f'  已完成该位置数据抓取')

print(big_df)
big_df.to_csv('fanduel_full_data.csv', index=False)
driver.quit()

代码说明

  1. 使用Select类处理下拉菜单,简化选项选择操作
  2. 嵌套循环确保每个Game Recency选项下的所有位置表格都被抓取
  3. 新增Game_Recency列,明确标记每条数据对应的时间范围
  4. 优化等待逻辑,减少不必要的长等待,提升抓取效率
  5. 最后调用driver.quit()关闭浏览器,释放资源

内容的提问来源于stack exchange,提问作者mmoore0323

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 14:45:41