You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BeautifulSoup无法获取全部stats_table类表格问题求助

解决方法

问题原因

Baseball Reference的部分统计表格被包裹在HTML注释(<!-- ... -->)中,默认情况下BeautifulSoup的html.parser不会解析注释内部的HTML结构,导致你只能提取到未被注释的2个表格,而手动查看时能看到注释里的其他7个表格。

修改后的代码

from bs4 import BeautifulSoup, Comment
import requests

# 获取统计表格
def get_hitting_stats(team, soup):
    # 提取未被注释的表格
    tables = soup.find_all("table", class_="stats_table")
    # 遍历所有注释节点,提取其中的表格
    for comment in soup.find_all(string=lambda text: isinstance(text, Comment)):
        comment_soup = BeautifulSoup(comment, 'html.parser')
        tables.extend(comment_soup.find_all("table", class_="stats_table"))
    print(f"找到{len(tables)}个stats_table表格")
    return tables
    
# 处理单场比赛数据
def process_game(gamelink, headers):
    req = requests.get(gamelink, headers)
    soup = BeautifulSoup(req.content, 'html.parser')
  
    home_hitting = get_hitting_stats("home", soup)
    away_hitting = get_hitting_stats("away", soup)
   
headers = {
    'User-Agent': 'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:52.0) Gecko/20100101 Firefox/52.0'
}
process_game("https://www.baseball-reference.com/boxes/CLE/CLE202208151.shtml", headers)

关键改动说明

  • 导入BeautifulSoup的Comment类,用于识别HTML注释节点
  • 在提取表格时,额外遍历所有注释内容,将注释内的HTML重新解析为soup对象,从中提取目标表格并合并到结果中
  • 移除了请求头中不必要的CORS相关字段,这些字段对客户端请求无实际作用

内容的提问来源于stack exchange,提问作者Justin Wade

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 21:39:21