BeautifulSoup无法获取全部stats_table类表格问题求助
解决方法
问题原因
Baseball Reference的部分统计表格被包裹在HTML注释(<!-- ... -->)中,默认情况下BeautifulSoup的html.parser不会解析注释内部的HTML结构,导致你只能提取到未被注释的2个表格,而手动查看时能看到注释里的其他7个表格。
修改后的代码
from bs4 import BeautifulSoup, Comment import requests # 获取统计表格 def get_hitting_stats(team, soup): # 提取未被注释的表格 tables = soup.find_all("table", class_="stats_table") # 遍历所有注释节点,提取其中的表格 for comment in soup.find_all(string=lambda text: isinstance(text, Comment)): comment_soup = BeautifulSoup(comment, 'html.parser') tables.extend(comment_soup.find_all("table", class_="stats_table")) print(f"找到{len(tables)}个stats_table表格") return tables # 处理单场比赛数据 def process_game(gamelink, headers): req = requests.get(gamelink, headers) soup = BeautifulSoup(req.content, 'html.parser') home_hitting = get_hitting_stats("home", soup) away_hitting = get_hitting_stats("away", soup) headers = { 'User-Agent': 'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:52.0) Gecko/20100101 Firefox/52.0' } process_game("https://www.baseball-reference.com/boxes/CLE/CLE202208151.shtml", headers)
关键改动说明
- 导入BeautifulSoup的
Comment类,用于识别HTML注释节点 - 在提取表格时,额外遍历所有注释内容,将注释内的HTML重新解析为soup对象,从中提取目标表格并合并到结果中
- 移除了请求头中不必要的CORS相关字段,这些字段对客户端请求无实际作用
内容的提问来源于stack exchange,提问作者Justin Wade
相关产品推荐
相关产品推荐

