使用BeautifulSoup爬取ESPN赛事赔率与Money Line失败求助
问题排查与修复方案
你的代码无法提取球队名称和金钱线,核心问题是页面选择器失效和ESPN部分内容动态加载,以下是具体修复步骤:
一、选择器失效问题
ESPN的网页类名、ID会频繁更新,你使用的rteQ类和topOdd ID已不是当前页面的有效标识。通过浏览器开发者工具查看目标页面,当前正确的选择器如下:
1. 提取球队名称
球队名称在带有Team__TeamName类的元素中,主客队可通过父元素的away/home类区分:
# 提取客队和主队名称 away_team = bs.find('div', class_='Team away').find('span', class_='Team__TeamName').text.strip() home_team = bs.find('div', class_='Team home').find('span', class_='Team__TeamName').text.strip()
2. 提取金钱线(Money Line)
金钱线数据位于带有Bet__Odd类的元素中,需先定位到"Money Line"对应的投注板块:
# 找到所有投注选项板块 bet_sections = bs.find_all('div', class_='Bet__Market') money_line_section = None for section in bet_sections: if 'Money Line' in section.text: money_line_section = section break if money_line_section: # 提取主客队金钱线 odds = money_line_section.find_all('div', class_='Bet__Odd') away_money_line = odds[0].text.strip() home_money_line = odds[1].text.strip() else: away_money_line = home_money_line = None
二、动态内容加载问题
ESPN部分投注数据通过JavaScript动态渲染,requests.get仅能获取静态HTML,可能拿不到完整数据。可通过两种方式解决:
方案1:添加请求头模拟浏览器
给requests.get加上浏览器User-Agent,部分场景下可触发静态数据返回:
headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } response = requests.get(url, headers=headers)
方案2:使用Selenium处理动态渲染
若请求头无效,用Selenium模拟浏览器加载页面,确保获取完整渲染后的HTML:
from selenium import webdriver from selenium.webdriver.chrome.options import Options def extract_money_lines(url): # 配置无头浏览器 chrome_options = Options() chrome_options.add_argument('--headless=new') driver = webdriver.Chrome(options=chrome_options) driver.get(url) bs = BeautifulSoup(driver.page_source, 'html.parser') driver.quit() # 后续提取逻辑同前 # ...
三、完整修复后的代码
import bs4 from bs4 import BeautifulSoup import requests import pandas as pd def extract_money_lines(url): headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } response = requests.get(url, headers=headers) bs = BeautifulSoup(response.text, 'html.parser') # 提取球队名称 home_team = away_team = None try: away_team = bs.find('div', class_='Team away').find('span', class_='Team__TeamName').text.strip() home_team = bs.find('div', class_='Team home').find('span', class_='Team__TeamName').text.strip() except AttributeError: pass # 提取金钱线 home_money_line = away_money_line = None bet_sections = bs.find_all('div', class_='Bet__Market') for section in bet_sections: if 'Money Line' in section.text: odds = section.find_all('div', class_='Bet__Odd') if len(odds) >= 2: away_money_line = odds[0].text.strip() home_money_line = odds[1].text.strip() break return home_team, away_team, home_money_line, away_money_line # Example NBA game ID game_id = "401585432" url = f"https://www.espn.com/nba/game/_/gameId/{game_id}" # Extract money lines and team names home_team, away_team, home_money_line, away_money_line = extract_money_lines(url) # Create DataFrame data = { 'Home Team': [home_team], 'Home Money Line': [home_money_line], 'Away Team': [away_team], 'Away Money Line': [away_money_line] } df = pd.DataFrame(data) print(df) # Save DataFrame to CSV #df.to_csv('nba_money_lines.csv', index=False) #print("CSV file saved successfully.")
额外提示
- ESPN页面结构会定期更新,建议每次爬取前用浏览器开发者工具确认最新选择器。
- 频繁爬取可能触发反爬机制,建议添加请求间隔(
time.sleep(2))。 - 优先考虑使用ESPN官方API接口,数据更稳定且不需要解析HTML(可通过游戏ID构造请求路径)。
内容的提问来源于stack exchange,提问作者Chip Koch
相关产品推荐
相关产品推荐

