修改Python爬虫代码提取Baseball-Reference页面的boxscore href链接
提取网页中Boxscore链接并导出Excel的修改方案
问题描述
我想用Python爬虫抓取requests.get指定页面中所有的"boxscore"超链接,并导出到Excel表格。但当前程序会输出网页中class为"game"的<p>元素下的所有文本,请问需要修改哪些部分,才能仅提取class为"game"的<p>元素下<em>标签内的boxscore href链接?
原代码:
import requests from bs4 import BeautifulSoup import pandas as pd from openpyxl import load_workbook wb = load_workbook("tennis_input3.xlsx") ws = wb.active response = requests.get('https://www.baseball-reference.com/leagues/majors/2010-schedule.shtml') webpage = response.content soup = BeautifulSoup(response.text, "html.parser") col1 = soup.find_all("p", class_="game") print(pd.DataFrame({"MatchLink":col1})) df = pd.DataFrame({"MatchLink":col1}) df.to_excel("tennis_3.xlsx", sheet_name="welcome")
修改要点
- 精准定位目标链接:原代码仅获取了
class="game"的<p>元素,需要进一步遍历这些元素,找到内部<em>标签下、href包含"boxscore"的<a>标签,提取其链接属性。 - 删除冗余代码:原代码加载了
tennis_input3.xlsx但未实际使用,直接移除这部分代码即可。 - 增加异常防护:遍历过程中加入判断逻辑,避免因部分元素无对应标签导致程序报错。
修改后的完整代码
import requests from bs4 import BeautifulSoup import pandas as pd # 发起请求并解析页面 response = requests.get('https://www.baseball-reference.com/leagues/majors/2010-schedule.shtml') soup = BeautifulSoup(response.text, "html.parser") # 存储所有boxscore链接 boxscore_links = [] # 遍历每个class为game的p元素 for game_p in soup.find_all("p", class_="game"): # 找到p元素下的em标签 em_tag = game_p.find("em") if em_tag: # 在em标签下筛选出href包含boxscore的a标签 boxscore_a = em_tag.find("a", href=lambda href: href and "boxscore" in href) if boxscore_a: # 拼接完整可访问的URL full_link = f"https://www.baseball-reference.com{boxscore_a['href']}" boxscore_links.append(full_link) # 转换为DataFrame并导出到Excel df = pd.DataFrame({"MatchLink": boxscore_links}) df.to_excel("tennis_3.xlsx", sheet_name="welcome", index=False) print(f"共抓取到{len(boxscore_links)}条boxscore链接,已导出到Excel")
代码说明
- 用
lambda href: href and "boxscore" in href精准筛选目标链接,避免抓取无关内容 - 拼接域名生成完整可直接访问的URL,解决原相对路径无法打开的问题
- 添加
index=False,避免导出Excel时自动生成无用的索引列 - 多层判断逻辑确保程序在遇到不规范页面元素时不会崩溃
内容的提问来源于stack exchange,提问作者NewGuy1
相关产品推荐
相关产品推荐

