请求修复Python Selenium爬取Prizepicks全NBA props的代码
修复Prizepicks NBA Props爬取代码并添加定时刷新功能
问题排查
- TypeError原因:你之前误将单个
stat-containerWebElement当作可迭代对象遍历,正确做法是获取该容器下所有stat子元素。 - Mac环境适配:原代码使用Windows的ChromeDriver路径,Mac下需改用驱动自动管理或正确路径。
- 反爬绕过:必须集成
selenium-stealth规避Prizepicks的反爬检测。 - 数据字段错误:原代码混淆了Prop类型(统计项名称)和数值,需修正字段映射。
修复后的完整代码
from selenium import webdriver from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium_stealth import stealth import time import pandas as pd def scrape_prizepicks_nba(): # 初始化Chrome驱动(适配Mac,自动管理版本) options = webdriver.ChromeOptions() options.add_argument("start-maximized") options.add_argument("--disable-blink-features=AutomationControlled") driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options) # Selenium-Stealth配置,绕过反爬 stealth(driver, languages=["en-US", "en"], vendor="Google Inc.", platform="MacIntel", webgl_vendor="Intel Inc.", renderer="Intel Iris OpenGL Engine", fix_hairline=True) driver.get("https://app.prizepicks.com/") try: # 关闭弹窗(如果存在) WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.CLASS_NAME, "close"))).click() # 进入NBA板块 WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.XPATH, "//div[@class='name'][normalize-space()='NBA']"))).click() # 等待统计项容器加载 stat_container = WebDriverWait(driver, 10).until(EC.visibility_of_element_located((By.CLASS_NAME, "stat-container"))) # 获取所有统计项元素 stat_elements = stat_container.find_elements(By.CSS_SELECTOR, "div.stat") nba_data = [] for stat in stat_elements: # 点击统计项,等待内容加载 stat_text = stat.text.strip() stat.click() time.sleep(2) # 等待投影数据加载 # 获取所有投影卡片 projections = WebDriverWait(driver, 15).until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, ".projection"))) for proj in projections: try: player_name = proj.find_element(By.XPATH, ".//div[@class='name']").text prop_line = proj.find_element(By.XPATH, ".//div[@class='presale-score']").text prop_over_under = proj.find_element(By.XPATH, ".//div[@class='text']").text nba_data.append({ "Player Name": player_name, "Prop Type": stat_text, "Line": prop_line, "Over/Under": prop_over_under }) except Exception as e: print(f"抓取单条数据失败: {e}") continue # 转为DataFrame并输出 df = pd.DataFrame(nba_data) print(df) # 可选:保存到CSV df.to_csv("prizepicks_nba_props.csv", mode='a', header=False, index=False) except Exception as e: print(f"爬取过程出错: {e}") finally: driver.quit() # 定时刷新功能(每30分钟刷新一次,可修改间隔) if __name__ == "__main__": refresh_interval = 1800 # 单位:秒,30分钟=1800秒 while True: print(f"开始新一轮爬取,当前时间: {time.strftime('%Y-%m-%d %H:%M:%S')}") scrape_prizepicks_nba() print(f"等待{refresh_interval//60}分钟后再次爬取...") time.sleep(refresh_interval)
代码说明
- 驱动管理:使用
webdriver-manager自动下载匹配Chrome版本的驱动,无需手动指定路径,完美适配Mac。 - 反爬规避:
selenium-stealth模拟真实浏览器环境,避免被Prizepicks检测为自动化工具。 - 统计项遍历:从
stat-container容器内获取所有统计项(Points、Rebounds、Assists等),逐个点击抓取对应数据。 - 数据准确性:修正字段映射,明确区分球员姓名、Prop类型、数值线和Over/Under方向。
- 定时刷新:通过
while循环+time.sleep实现定时爬取,默认30分钟刷新一次,可按需修改refresh_interval参数。 - 错误处理:添加异常捕获,避免单条数据抓取失败导致整个程序崩溃。
内容的提问来源于stack exchange,提问作者joey
相关产品推荐
相关产品推荐

