You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修改Python代码从Rotowire提取MLB预测球员姓名?

问题描述

我之前用Python代码从BaseballPress.com获取MLB官方先发阵容,但这些阵容通常要到赛前1小时左右才会发布。

原代码如下:

import requests
import pandas as pd
import openpyxl
from bs4 import BeautifulSoup

url = "https://www.baseballpress.com/lineups/2022-08-09"
soup = BeautifulSoup(requests.get(url).content, "html.parser")

def get_name(tag):
    if tag.select_one(".desktop-name"):
        return tag.select_one(".desktop-name").get_text()
    elif tag.select_one(".mobile-name"):
        return tag.select_one(".mobile-name").get_text()
    else:
       return tag.get_text()

data = []
for card in soup.select(".lineup-card"):
    header = [
        c.get_text(strip=True, separator=" ")
        for c in card.select(".lineup-card-header .c")
    ]
    h_p1, h_p2 = [
        get_name(p) for p in card.select(".lineup-card-header .player")
    ]
    data.append([*header, h_p1, h_p2])

    for p1, p2 in zip(
        card.select(".col--min:nth-of-type(1) .player"),
        card.select(".col--min:nth-of-type(2) .player"),
    ):
        p1 = get_name(p1).split(maxsplit=1)[-1]
        p2 = get_name(p2).split(maxsplit=1)[-1]

        data.append([*header, p1, p2])

df = pd.DataFrame(
    data, columns=["Team1", "Date", "Team2", "Player1", "Player2"]
)
df.to_excel("MLB Games.xlsx", sheet_name='sheet1', index=False)
print(df.head(10).to_markdown(index=False))

为了提前获取预测阵容,我转而使用Rotowire(它会提前24小时发布预测阵容),并修改了脚本,但不确定如何调整get_name()函数。修改后的代码如下:

import requests
import pandas as pd
import openpyxl
from bs4 import BeautifulSoup

url = "https://www.rotowire.com/baseball/daily-lineups.php"
soup = BeautifulSoup(requests.get(url).content, "html.parser")

def get_name(tag):
    if tag.select_one(".desktop-name"):
        return tag.select_one(".desktop-name").get_text()
    elif tag.select_one(".mobile-name"):
        return tag.select_one(".mobile-name").get_text()
    else:
       return tag.get_text()

data = []
for card in soup.select(".lineup__main"):
    header = [
        c.get_text(strip=True, separator=" ")
        for c in card.select(".lineup__teams .c")
    ]
    h_p1, h_p2 = [
        get_name(p) for p in card.select(".lineup__teams .lineup__player")
    ]
    data.append([*header, h_p1, h_p2])

    for p1, p2 in zip(
        card.select(".lineup__list is-visit:nth-of-type(1) .lineup__player"),
        card.select(".lineup__list is-home:nth-of-type(2) .lineup__player"),
    ):
        p1 = get_name(p1).split(maxsplit=1)[-1]
        p2 = get_name(p2).split(maxsplit=1)[-1]

        data.append([*header, p1, p2])

df = pd.DataFrame(
    data, columns=["Team1", "Date", "Team2", "Player1", "Player2"]
)
df.to_excel("MLB Predicted Lineups.xlsx", sheet_name='sheet1', index=False)
print(df.head(10).to_markdown(index=False))
解决方案

Rotowire的HTML结构和BaseballPress完全不同,原get_name()函数依赖的类名在Rotowire页面中不存在,需要针对性修改,同时还要修正代码里的其他选择器错误:

1. 调整get_name()函数

Rotowire的球员名称统一放在.lineup__player-name类的元素内,直接提取该元素的文本即可:

def get_name(tag):
    # 提取Rotowire页面中的球员名称
    name_tag = tag.select_one(".lineup__player-name")
    if name_tag:
        return name_tag.get_text(strip=True)
    # 兼容特殊情况,直接返回标签文本
    return tag.get_text(strip=True)

2. 修正选择器错误

原代码中的客队/主队列表选择器格式错误,且类名匹配不正确,需要调整为:

  • 客队列表:.lineup__list--visit
  • 主队列表:.lineup__list--home

同时,球队和日期的提取逻辑也需要优化,确保数据准确。

3. 完整修正后的代码

import requests
import pandas as pd
import openpyxl
from bs4 import BeautifulSoup

url = "https://www.rotowire.com/baseball/daily-lineups.php"
soup = BeautifulSoup(requests.get(url).content, "html.parser")

def get_name(tag):
    name_tag = tag.select_one(".lineup__player-name")
    if name_tag:
        return name_tag.get_text(strip=True)
    return tag.get_text(strip=True)

data = []
for card in soup.select(".lineup__main"):
    # 提取客队、主队和比赛日期
    team_visit = card.select_one(".lineup__team--visit .lineup__team-name").get_text(strip=True)
    team_home = card.select_one(".lineup__team--home .lineup__team-name").get_text(strip=True)
    game_date = card.select_one(".lineup__date").get_text(strip=True)
    header = [team_visit, game_date, team_home]
    
    # 提取先发投手
    pitcher_visit = get_name(card.select_one(".lineup__pitcher--visit .lineup__player"))
    pitcher_home = get_name(card.select_one(".lineup__pitcher--home .lineup__player"))
    data.append([*header, pitcher_visit, pitcher_home])

    # 提取双方打线球员
    for p1, p2 in zip(
        card.select(".lineup__list--visit .lineup__player"),
        card.select(".lineup__list--home .lineup__player"),
    ):
        p1_name = get_name(p1)
        p2_name = get_name(p2)
        data.append([*header, p1_name, p2_name])

df = pd.DataFrame(
    data, columns=["Team1", "Date", "Team2", "Player1", "Player2"]
)
df.to_excel("MLB Predicted Lineups.xlsx", sheet_name='sheet1', index=False)
print(df.head(10).to_markdown(index=False))

说明

  • 完全适配Rotowire的页面结构,精准提取球员名称
  • 优化了球队和日期的提取逻辑,避免原代码中可能出现的空值或错误数据
  • 修正了列表选择器的格式和类名,确保能正确匹配客队和主队的打线

内容的提问来源于stack exchange,提问作者Ben Ahnen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 21:39:21