You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何BeautifulSoup无法抓取baseball-reference.com的全部表格?

问题原因

Baseball Reference页面里,除前两个MVP投票表格外,后续的Cy Young、年度新秀等表格都被HTML注释(格式为<!-- ... -->)包裹,BeautifulSoup默认不会解析注释节点内的HTML内容,所以直接用find_all('table')只能拿到未被注释的前两个表格。

解决方法

先提取页面中的所有注释节点,再把注释里的HTML内容单独解析成BeautifulSoup对象,从中提取表格后和初始获取的表格合并。

修改后的代码如下:

import requests
from bs4 import BeautifulSoup, Comment

url = 'https://www.baseball-reference.com/awards/awards_2017.shtml'

page = requests.get(url)
soup = BeautifulSoup(page.text, 'html.parser')

# 获取页面中未被注释的表格
tables = soup.find_all('table')

# 提取所有HTML注释节点
comments = soup.find_all(string=lambda text: isinstance(text, Comment))

# 遍历注释,解析其中的表格并添加到列表中
for comment in comments:
    comment_soup = BeautifulSoup(comment, 'html.parser')
    comment_tables = comment_soup.find_all('table')
    tables.extend(comment_tables)

# 现在tables列表包含所有表格,可测试第三个表格(NL Cy Young Voting)
print(tables[2])
补充说明
  • 用Comment类识别页面中的注释节点,确保不会漏掉隐藏的表格内容
  • 每个注释内容单独解析,能完整提取其中的表格结构
  • 合并后的tables列表包含所有奖项投票表格,也可通过表格的id属性精准定位(比如table.find(id='nl_cy_young_voting'))

内容的提问来源于stack exchange,提问作者an-izq

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 05:52:04