You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python爬取baseball-reference.com比赛技术统计页面遇问题

解决方法

问题出在页面的技术统计表格内容被包裹在HTML注释中,BeautifulSoup默认不会解析注释里的元素,所以你之前的代码只能拿到页面可见的底部菜单链接。需要做以下修改:

  • 提取页面中所有包含球员数据的注释内容
  • 将注释内容转换成可解析的HTML对象
  • 在转换后的HTML中筛选目标球员链接

修改后的代码如下:

import requests
from bs4 import BeautifulSoup, Comment

# URL of the webpage
url = "https://www.baseball-reference.com/boxes/ANA/ANA202305210.shtml"

# Send a GET request to the webpage
response = requests.get(url)

# Check if the request was successful
if response.status_code == 200:
    # Parse the HTML content of the webpage using BeautifulSoup
    soup = BeautifulSoup(response.content, 'html.parser')
    
    # 提取页面中所有的HTML注释
    comments = soup.find_all(string=lambda text: isinstance(text, Comment))
    
    # 遍历注释,找到包含球员数据的表格部分
    player_links = []
    for comment in comments:
        # 将注释内容转为BeautifulSoup对象
        comment_soup = BeautifulSoup(comment, 'html.parser')
        # 查找该注释里所有包含/players/的链接
        links = comment_soup.find_all('a', href=lambda href: href and '/players/' in href)
        player_links.extend(links)
    
    # 去重并打印球员链接
    seen_hrefs = set()
    for link in player_links:
        href = link.get('href')
        if href not in seen_hrefs:
            seen_hrefs.add(href)
            print(href)
else:
    print("Failed to fetch the webpage.")

说明:

  • 引入Comment类来识别HTML注释内容
  • 遍历所有注释,将每个注释转为可解析的HTML结构,从中提取球员链接
  • 增加去重逻辑,避免重复输出同一球员的链接

内容的提问来源于stack exchange,提问作者NewGuy1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 02:43:10