如何修改Python代码从Rotowire提取MLB预测球员姓名?
问题描述
我之前用Python代码从BaseballPress.com获取MLB官方先发阵容,但这些阵容通常要到赛前1小时左右才会发布。
原代码如下:
import requests import pandas as pd import openpyxl from bs4 import BeautifulSoup url = "https://www.baseballpress.com/lineups/2022-08-09" soup = BeautifulSoup(requests.get(url).content, "html.parser") def get_name(tag): if tag.select_one(".desktop-name"): return tag.select_one(".desktop-name").get_text() elif tag.select_one(".mobile-name"): return tag.select_one(".mobile-name").get_text() else: return tag.get_text() data = [] for card in soup.select(".lineup-card"): header = [ c.get_text(strip=True, separator=" ") for c in card.select(".lineup-card-header .c") ] h_p1, h_p2 = [ get_name(p) for p in card.select(".lineup-card-header .player") ] data.append([*header, h_p1, h_p2]) for p1, p2 in zip( card.select(".col--min:nth-of-type(1) .player"), card.select(".col--min:nth-of-type(2) .player"), ): p1 = get_name(p1).split(maxsplit=1)[-1] p2 = get_name(p2).split(maxsplit=1)[-1] data.append([*header, p1, p2]) df = pd.DataFrame( data, columns=["Team1", "Date", "Team2", "Player1", "Player2"] ) df.to_excel("MLB Games.xlsx", sheet_name='sheet1', index=False) print(df.head(10).to_markdown(index=False))
为了提前获取预测阵容,我转而使用Rotowire(它会提前24小时发布预测阵容),并修改了脚本,但不确定如何调整get_name()函数。修改后的代码如下:
import requests import pandas as pd import openpyxl from bs4 import BeautifulSoup url = "https://www.rotowire.com/baseball/daily-lineups.php" soup = BeautifulSoup(requests.get(url).content, "html.parser") def get_name(tag): if tag.select_one(".desktop-name"): return tag.select_one(".desktop-name").get_text() elif tag.select_one(".mobile-name"): return tag.select_one(".mobile-name").get_text() else: return tag.get_text() data = [] for card in soup.select(".lineup__main"): header = [ c.get_text(strip=True, separator=" ") for c in card.select(".lineup__teams .c") ] h_p1, h_p2 = [ get_name(p) for p in card.select(".lineup__teams .lineup__player") ] data.append([*header, h_p1, h_p2]) for p1, p2 in zip( card.select(".lineup__list is-visit:nth-of-type(1) .lineup__player"), card.select(".lineup__list is-home:nth-of-type(2) .lineup__player"), ): p1 = get_name(p1).split(maxsplit=1)[-1] p2 = get_name(p2).split(maxsplit=1)[-1] data.append([*header, p1, p2]) df = pd.DataFrame( data, columns=["Team1", "Date", "Team2", "Player1", "Player2"] ) df.to_excel("MLB Predicted Lineups.xlsx", sheet_name='sheet1', index=False) print(df.head(10).to_markdown(index=False))
解决方案
Rotowire的HTML结构和BaseballPress完全不同,原get_name()函数依赖的类名在Rotowire页面中不存在,需要针对性修改,同时还要修正代码里的其他选择器错误:
1. 调整get_name()函数
Rotowire的球员名称统一放在.lineup__player-name类的元素内,直接提取该元素的文本即可:
def get_name(tag): # 提取Rotowire页面中的球员名称 name_tag = tag.select_one(".lineup__player-name") if name_tag: return name_tag.get_text(strip=True) # 兼容特殊情况,直接返回标签文本 return tag.get_text(strip=True)
2. 修正选择器错误
原代码中的客队/主队列表选择器格式错误,且类名匹配不正确,需要调整为:
- 客队列表:
.lineup__list--visit - 主队列表:
.lineup__list--home
同时,球队和日期的提取逻辑也需要优化,确保数据准确。
3. 完整修正后的代码
import requests import pandas as pd import openpyxl from bs4 import BeautifulSoup url = "https://www.rotowire.com/baseball/daily-lineups.php" soup = BeautifulSoup(requests.get(url).content, "html.parser") def get_name(tag): name_tag = tag.select_one(".lineup__player-name") if name_tag: return name_tag.get_text(strip=True) return tag.get_text(strip=True) data = [] for card in soup.select(".lineup__main"): # 提取客队、主队和比赛日期 team_visit = card.select_one(".lineup__team--visit .lineup__team-name").get_text(strip=True) team_home = card.select_one(".lineup__team--home .lineup__team-name").get_text(strip=True) game_date = card.select_one(".lineup__date").get_text(strip=True) header = [team_visit, game_date, team_home] # 提取先发投手 pitcher_visit = get_name(card.select_one(".lineup__pitcher--visit .lineup__player")) pitcher_home = get_name(card.select_one(".lineup__pitcher--home .lineup__player")) data.append([*header, pitcher_visit, pitcher_home]) # 提取双方打线球员 for p1, p2 in zip( card.select(".lineup__list--visit .lineup__player"), card.select(".lineup__list--home .lineup__player"), ): p1_name = get_name(p1) p2_name = get_name(p2) data.append([*header, p1_name, p2_name]) df = pd.DataFrame( data, columns=["Team1", "Date", "Team2", "Player1", "Player2"] ) df.to_excel("MLB Predicted Lineups.xlsx", sheet_name='sheet1', index=False) print(df.head(10).to_markdown(index=False))
说明
- 完全适配Rotowire的页面结构,精准提取球员名称
- 优化了球队和日期的提取逻辑,避免原代码中可能出现的空值或错误数据
- 修正了列表选择器的格式和类名,确保能正确匹配客队和主队的打线
内容的提问来源于stack exchange,提问作者Ben Ahnen
相关产品推荐
相关产品推荐

