You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取HTML表格无结果的问题及解决

问题描述

爬取目标网页的HTML表格元素时无法获取任何数据,但爬取表格外的内容完全正常。曾怀疑数据由JavaScript加载,尝试用Selenium仍无法获取数据,想知道遗漏了什么,以及其他爬取HTML表格数据的方法。

问题原因

目标网站的表格内容被包裹在HTML注释标签<!-- -->中,直接使用BeautifulSoup解析时,注释内的HTML结构会被当作注释内容忽略,导致无法通过选择器定位到表格内的元素。

原始尝试代码
# 抓取2020赛季数据
jz_2020_raw_stats = requests.get('https://www.basketball-reference.com/teams/UTA/2020.html').text
jz_2020_soup_stats = BeautifulSoup(jz_2020_raw_stats,'html.parser')
jz_2020_soup_stats.prettify()

# 提取所需列的值
pts = jz_2020_soup_stats.select("#team_and_opponent_sh > div > ul > li:nth-child(1) > span")[0].text
print(pts)
fg_pct = jz_2020_soup_stats.select("#team_and_opponent > tbody:nth-child(4) > tr:nth-child(1) > td:nth-child(6)")[0].text
print(fg_pct)
three_pct = jz_2020_soup_stats.select("#team_and_opponent > tbody:nth-child(4) > tr:nth-child(1) > td:nth-child(9)")[0].text
tov = jz_2020_soup_stats.select("#team_and_opponent > tbody:nth-child(4) > tr:nth-child(1) > td:nth-child(22)")[0].text
ft_pct = jz_2020_soup_stats.select("#team_and_opponent > tbody:nth-child(4) > tr:nth-child(1) > td:nth-child(15)")[0].text


# 打印数值
print('PTS:', pts)
print('FG%:', fg_pct)
print('3P%:', three_pct)
print('TOV:', tov)
print('FT%:', ft_pct)
可行解决方案代码
page = requests.get(jz_2020_url)
# 移除HTML注释标签,让表格内容可被解析
page = page.text.replace("<!--","").replace("-->","")
soup = BeautifulSoup(page, 'html.parser') 

# 定位目标表格并转为字符串
table_html = str(soup.find("table", {"id": "team_and_opponent"}))
# 使用pandas直接读取表格为DataFrame
df = pd.read_html(table_html)[0]

print(df)
补充说明
  • 移除注释标签后,原本被隐藏的表格结构会被正常识别,此时可以用BeautifulSoup定位表格,再通过pandas.read_html()快速将表格转为DataFrame,无需手动定位每个单元格。
  • 若遇到类似静态HTML中内容被注释包裹的情况,都可以先移除注释标签再进行解析。

内容的提问来源于stack exchange,提问作者Alex

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 03:58:08