You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

修改Python爬虫代码提取Baseball-Reference页面的boxscore href链接

提取网页中Boxscore链接并导出Excel的修改方案

问题描述

我想用Python爬虫抓取requests.get指定页面中所有的"boxscore"超链接,并导出到Excel表格。但当前程序会输出网页中class为"game"的<p>元素下的所有文本,请问需要修改哪些部分,才能仅提取class为"game"的<p>元素下<em>标签内的boxscore href链接?

原代码:

import requests
from bs4 import BeautifulSoup
import pandas as pd
from openpyxl import load_workbook

wb = load_workbook("tennis_input3.xlsx")
ws = wb.active

response = requests.get('https://www.baseball-reference.com/leagues/majors/2010-schedule.shtml')
webpage = response.content
soup = BeautifulSoup(response.text, "html.parser")
  
col1 = soup.find_all("p", class_="game")

print(pd.DataFrame({"MatchLink":col1}))
df = pd.DataFrame({"MatchLink":col1})

df.to_excel("tennis_3.xlsx", sheet_name="welcome")

修改要点

  1. 精准定位目标链接:原代码仅获取了class="game"的<p>元素,需要进一步遍历这些元素,找到内部<em>标签下、href包含"boxscore"的<a>标签,提取其链接属性。
  2. 删除冗余代码:原代码加载了tennis_input3.xlsx但未实际使用,直接移除这部分代码即可。
  3. 增加异常防护:遍历过程中加入判断逻辑,避免因部分元素无对应标签导致程序报错。

修改后的完整代码

import requests
from bs4 import BeautifulSoup
import pandas as pd

# 发起请求并解析页面
response = requests.get('https://www.baseball-reference.com/leagues/majors/2010-schedule.shtml')
soup = BeautifulSoup(response.text, "html.parser")

# 存储所有boxscore链接
boxscore_links = []

# 遍历每个class为game的p元素
for game_p in soup.find_all("p", class_="game"):
    # 找到p元素下的em标签
    em_tag = game_p.find("em")
    if em_tag:
        # 在em标签下筛选出href包含boxscore的a标签
        boxscore_a = em_tag.find("a", href=lambda href: href and "boxscore" in href)
        if boxscore_a:
            # 拼接完整可访问的URL
            full_link = f"https://www.baseball-reference.com{boxscore_a['href']}"
            boxscore_links.append(full_link)

# 转换为DataFrame并导出到Excel
df = pd.DataFrame({"MatchLink": boxscore_links})
df.to_excel("tennis_3.xlsx", sheet_name="welcome", index=False)

print(f"共抓取到{len(boxscore_links)}条boxscore链接,已导出到Excel")

代码说明

  • 用lambda href: href and "boxscore" in href精准筛选目标链接,避免抓取无关内容
  • 拼接域名生成完整可直接访问的URL,解决原相对路径无法打开的问题
  • 添加index=False,避免导出Excel时自动生成无用的索引列
  • 多层判断逻辑确保程序在遇到不规范页面元素时不会崩溃

内容的提问来源于stack exchange,提问作者NewGuy1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 03:45:52