如何使用Beautiful Soup的find函数提取指定id或class的HTML表格元素
抓取失败原因
目标team_pitching表格被站点放在了HTML注释块中做懒加载优化,直接用find方法查询初始HTML时,注释内的DOM元素不会被识别,因此返回空值。而上方的team_batting表格属于首屏直接渲染的内容,所以可以正常提取。
解决方法
先提取页面内的所有注释内容,筛选出包含目标表格的注释块后单独解析,即可正常提取对应元素。同时支持通过id、class两种属性匹配表格。
完整可运行代码
import requests from bs4 import BeautifulSoup, Comment url = 'https://www.baseball-reference.com/register/team.cgi?id=17cdc2d2' res = requests.get(url) soup1 = BeautifulSoup(res.content, 'html.parser') # 提取页面所有注释,筛选包含目标表格的注释 comments = soup1.find_all(string=lambda text: isinstance(text, Comment)) pitching_comment = None for c in comments: if 'id="team_pitching"' in c: pitching_comment = c break # 解析注释内的HTML内容 pitching_soup = BeautifulSoup(pitching_comment, 'html.parser') # 方法1:通过id提取表格 table_by_id = pitching_soup.find('table', id='team_pitching') # 方法2:通过class属性提取表格,多class需用列表传入匹配 table_by_class = pitching_soup.find('table', class_=['sortable', 'stats_table', 'now_sortable'])
注意事项
- 多class属性匹配时不要直接传入完整字符串
sortable stats_table now_sortable,HTML中class属性的空格是分隔符,直接传入完整字符串会被识别为单个class名,无法匹配成功,传入列表形式的多class值是最稳妥的写法。 - 该站点大部分非首屏表格都采用了注释包裹的懒加载逻辑,后续抓取同站点其他表格可复用上述逻辑。
内容的提问来源于stack exchange,提问作者David
相关产品推荐
相关产品推荐

