如何用Python完善ESPN NFL赛程爬取并生成指定DataFrame
问题描述
作为编程新手,我尝试用Python爬取ESPN网站2024年NFL第1周赛程(网址:https://www.espn.com/nfl/schedule/_/week/1/year/2024/seasontype/2)并存储为DataFrame。目前已成功提取主客场球队名称,但无法获取比赛时间、比赛地点、赔率信息,希望完善代码,生成包含「客场球队、主场球队、比赛时间、比赛地点、赔率」列的DataFrame。
现有代码:
url = 'https://www.espn.com/nfl/scoreboard/_/week/1/year/2024/seasontype/2' # Headers to make the request look like it's coming from a browser headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/58.0.3029.110 Safari/537.3" } # Send a GET request to the webpage with headers response = requests.get(url, headers=headers) src = (response.content) soup = BeautifulSoup(response.content, 'html.parser') # Find all the game containers game_containers = soup.find_all('a',class_="AnchorLink" ) team_names = soup.find_all('div', class_='ScoreCell__TeamName ScoreCell__TeamName--shortDisplayName truncate db') # List to hold the team names team_list = [team.text for team in team_names] # Pair the team names into away and home teams away_teams = team_list[::2] # Every other team starting from the first home_teams = team_list[1::2] # Every other team starting from the second # Create a DataFrame from the data df = pd.DataFrame({ 'Away Team': away_teams, 'Home Team': home_teams }) # Print the DataFrame print(df)
HTML结构分析:赛程内容位于<div class="mt3">下,单场比赛信息在<tr class="Table__TR Table__TR--sm Table__even" data-idx="0">标签内,其中比赛时间在<td class="date__col Table__TD">下的AnchorLink标签中,比赛地点在<td class="location__col Table__TD">下的div中,赔率信息在<td class="odds__col Table__TD">下的Odds__Message相关标签内。
解决方案
核心思路
不再单独提取所有球队名称,而是按单场比赛为单位遍历每个赛事容器,分别提取该场的客场/主场球队、时间、地点、赔率,这样能保证各字段对应关系准确,避免因页面结构变化导致的索引错位问题。
完整代码
import requests from bs4 import BeautifulSoup import pandas as pd url = 'https://www.espn.com/nfl/schedule/_/week/1/year/2024/seasontype/2' headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/58.0.3029.110 Safari/537.3" } # 发送请求并解析页面 response = requests.get(url, headers=headers) soup = BeautifulSoup(response.content, 'html.parser') # 找到所有单场比赛的容器(对应分析的tr标签) game_rows = soup.find_all('tr', class_=['Table__TR', 'Table__TR--sm', 'Table__even', 'Table__odd']) # 初始化存储数据的列表 games_data = [] for row in game_rows: # 跳过表头或空行 if not row.find('td', class_='date__col'): continue # 提取比赛时间 date_cell = row.find('td', class_='date__col') game_time = date_cell.find('a', class_='AnchorLink').text.strip() if date_cell else 'N/A' # 提取主客场球队 team_cells = row.find_all('div', class_='ScoreCell__TeamName ScoreCell__TeamName--shortDisplayName truncate db') if len(team_cells) >= 2: away_team = team_cells[0].text.strip() home_team = team_cells[1].text.strip() else: away_team = 'N/A' home_team = 'N/A' # 提取比赛地点 location_cell = row.find('td', class_='location__col') game_location = location_cell.find('div').text.strip() if location_cell else 'N/A' # 提取赔率信息 odds_cell = row.find('td', class_='odds__col') odds_message = odds_cell.find('div', class_='Odds__Message') game_odds = odds_message.text.strip() if odds_message else 'N/A' # 将单场数据加入列表 games_data.append({ '客场球队': away_team, '主场球队': home_team, '比赛时间': game_time, '比赛地点': game_location, '赔率': game_odds }) # 转换为DataFrame df = pd.DataFrame(games_data) print(df)
关键说明
- 按单场遍历:通过
game_rows获取所有赛事行,逐个处理每场比赛的所有字段,确保数据对应准确。 - 字段提取逻辑:
- 时间:从
date__col下的AnchorLink标签提取文本 - 地点:从
location__col下的div标签提取文本 - 赔率:从
odds__col下的Odds__Message标签提取文本,若不存在则显示N/A
- 时间:从
- 异常处理:添加了空值判断,避免因页面部分赛事信息缺失导致代码报错。
内容的提问来源于stack exchange,提问作者Yangxi Liu
相关产品推荐
相关产品推荐

