为何BeautifulSoup无法抓取baseball-reference.com的全部表格?
问题原因
Baseball Reference页面里,除前两个MVP投票表格外,后续的Cy Young、年度新秀等表格都被HTML注释(格式为<!-- ... -->)包裹,BeautifulSoup默认不会解析注释节点内的HTML内容,所以直接用find_all('table')只能拿到未被注释的前两个表格。
解决方法
先提取页面中的所有注释节点,再把注释里的HTML内容单独解析成BeautifulSoup对象,从中提取表格后和初始获取的表格合并。
修改后的代码如下:
import requests from bs4 import BeautifulSoup, Comment url = 'https://www.baseball-reference.com/awards/awards_2017.shtml' page = requests.get(url) soup = BeautifulSoup(page.text, 'html.parser') # 获取页面中未被注释的表格 tables = soup.find_all('table') # 提取所有HTML注释节点 comments = soup.find_all(string=lambda text: isinstance(text, Comment)) # 遍历注释,解析其中的表格并添加到列表中 for comment in comments: comment_soup = BeautifulSoup(comment, 'html.parser') comment_tables = comment_soup.find_all('table') tables.extend(comment_tables) # 现在tables列表包含所有表格,可测试第三个表格(NL Cy Young Voting) print(tables[2])
补充说明
- 用
Comment类识别页面中的注释节点,确保不会漏掉隐藏的表格内容 - 每个注释内容单独解析,能完整提取其中的表格结构
- 合并后的
tables列表包含所有奖项投票表格,也可通过表格的id属性精准定位(比如table.find(id='nl_cy_young_voting'))
内容的提问来源于stack exchange,提问作者an-izq
相关产品推荐
相关产品推荐

