如何抓取RotoWire NBA球员投注数据并转为DataFrame?
解决方案:抓取Rotowire NBA球员投注数据并转为DataFrame
一、利用订阅者专属的「导出CSV」功能(优先推荐)
作为订阅者,官方导出的CSV数据最完整、最稳定,完全无需复杂爬取:
- 操作步骤:
- 用Selenium登录账号(若页面需验证),打开目标页面
- 定位每个表格右上角的「Export CSV」按钮
- 循环点击按钮下载对应市场的CSV,读取后添加
market列标注投注类型,最后合并成统一DataFrame
代码示例:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import pandas as pd import os import time # 初始化Chrome浏览器(需对应驱动) driver = webdriver.Chrome() driver.get("https://www.rotowire.com/betting/nba/player-props.php") # 登录逻辑(如果需要) # driver.find_element(By.ID, "username").send_keys("你的账号") # driver.find_element(By.ID, "password").send_keys("你的密码") # driver.find_element(By.XPATH, "//button[@type='submit']").click() # 等待所有表格加载完成 WebDriverWait(driver, 10).until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, ".table-rotowire"))) all_dfs = [] download_dir = os.path.expanduser("~/Downloads") # 遍历所有表格 for table in driver.find_elements(By.CSS_SELECTOR, ".table-rotowire"): # 获取当前表格对应的投注市场名称(比如"得分"、"助攻") market_name = table.find_element(By.XPATH, "./preceding-sibling::h3").text.strip() # 点击导出按钮 table.find_element(By.CSS_SELECTOR, ".table-export").click() # 等待下载完成 time.sleep(2) # 读取最新下载的CSV latest_csv = max([os.path.join(download_dir, f) for f in os.listdir(download_dir) if f.endswith(".csv")], key=os.path.getctime) df = pd.read_csv(latest_csv) df["market"] = market_name all_dfs.append(df) # 可选:删除已处理的CSV避免重复 os.remove(latest_csv) # 合并所有数据 final_df = pd.concat(all_dfs, ignore_index=True) driver.quit() # 查看结果 print(final_df.head())
二、修复Selenium爬取页面数据(无导出按钮时备用)
如果不想用导出功能,需要处理页面动态加载的隐藏行和多表格遍历问题:
- 问题根源:表格行需滚动到视图内才会加载,多表格需逐个定位处理
- 解决代码:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.common.action_chains import ActionChains import pandas as pd driver = webdriver.Chrome() driver.get("https://www.rotowire.com/betting/nba/player-props.php") driver.maximize_window() driver.implicitly_wait(10) all_dfs = [] # 遍历每个投注市场区块 for section in driver.find_elements(By.CSS_SELECTOR, ".prop-market"): market_name = section.find_element(By.TAG_NAME, "h3").text.strip() # 滚动到区块位置,触发隐藏行加载 ActionChains(driver).move_to_element(section).perform() # 提取表格数据 table = section.find_element(By.CSS_SELECTOR, ".table-rotowire") headers = [th.text.strip() for th in table.find_elements(By.TAG_NAME, "th")] rows = table.find_elements(By.TAG_NAME, "tr") table_data = [] for row in rows[1:]: # 跳过表头行 cols = [td.text.strip() for td in row.find_elements(By.TAG_NAME, "td")] if cols: table_data.append(cols) df = pd.DataFrame(table_data, columns=headers) df["market"] = market_name all_dfs.append(df) final_df = pd.concat(all_dfs, ignore_index=True) driver.quit() print(final_df.head())
三、为什么requests/BeautifulSoup会失败?
Rotowire页面数据是JavaScript动态渲染的,初始HTML仅含页面框架,实际表格数据通过AJAX请求加载。若坚持用requests,需要抓包找到真实数据接口,并携带登录后的认证Cookie、正确请求头,但这种方法极易因接口变更失效,不推荐。
内容的提问来源于stack exchange,提问作者AMJ
相关产品推荐
相关产品推荐

