遍历网页不同选择器并将数据存入统一大DataFrame的问题
解决方案
要实现抓取所有Game Recency选项的数据并新增对应列,需在原有代码基础上嵌套遍历Game Recency下拉选项,具体修改如下:
关键修改点
- 定位Game Recency下拉菜单,获取所有可选项
- 嵌套循环:先遍历每个Recency选项,再遍历各位置表格
- 为每个数据帧新增
Game_Recency列,记录当前选中的时间范围 - 优化等待逻辑,确保元素加载完成后再操作
修改后的完整代码
from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.support.ui import Select import time as t import pandas as pd pd.set_option('display.max_columns', None) pd.set_option('display.max_colwidth', None) big_df = pd.DataFrame() chrome_options = Options() chrome_options.add_argument("--no-sandbox") chrome_options.add_argument('disable-notifications') chrome_options.add_argument("window-size=1280,720") webdriver_service = Service(r'chromedriver\chromedriver') ## 替换为你的chromedriver路径 driver = webdriver.Chrome(service=webdriver_service, options=chrome_options) wait = WebDriverWait(driver, 20) url = "https://www.fantasypros.com/daily-fantasy/nba/fanduel-defense-vs-position.php" driver.get(url) t.sleep(5) # 可根据网络情况调整初始等待时间 # 定位Game Recency下拉菜单并获取所有选项 recency_select = wait.until(EC.presence_of_element_located((By.ID, 'game-recency'))) select = Select(recency_select) recency_options = [option.text.strip() for option in select.options] # 嵌套遍历:先遍历每个Game Recency选项,再遍历各位置表格 for recency in recency_options: select.select_by_visible_text(recency) print(f'已选中Game Recency: {recency}') t.sleep(2) # 获取所有位置选项 tables_list = wait.until(EC.presence_of_all_elements_located((By.XPATH, '//ul[@class="pills pos-filter pull-left"]/li'))) for x in tables_list: x.click() print(f' 已选中位置: {x.text}') t.sleep(2) table = wait.until(EC.element_to_be_clickable((By.XPATH, '//table[@id="data-table"]'))) df = pd.read_html(table.get_attribute('outerHTML'))[0] df['Category'] = x.text.strip() df['Game_Recency'] = recency # 新增Game Recency标记列 big_df = pd.concat([big_df, df], axis=0, ignore_index=True) print(f' 已完成该位置数据抓取') print(big_df) big_df.to_csv('fanduel_full_data.csv', index=False) driver.quit()
代码说明
- 使用
Select类处理下拉菜单,简化选项选择操作 - 嵌套循环确保每个Game Recency选项下的所有位置表格都被抓取
- 新增
Game_Recency列,明确标记每条数据对应的时间范围 - 优化等待逻辑,减少不必要的长等待,提升抓取效率
- 最后调用
driver.quit()关闭浏览器,释放资源
内容的提问来源于stack exchange,提问作者mmoore0323
相关产品推荐
相关产品推荐

