使用BeautifulSoup爬取HTML表格无结果的问题及解决
问题描述
爬取目标网页的HTML表格元素时无法获取任何数据,但爬取表格外的内容完全正常。曾怀疑数据由JavaScript加载,尝试用Selenium仍无法获取数据,想知道遗漏了什么,以及其他爬取HTML表格数据的方法。
问题原因
目标网站的表格内容被包裹在HTML注释标签<!-- -->中,直接使用BeautifulSoup解析时,注释内的HTML结构会被当作注释内容忽略,导致无法通过选择器定位到表格内的元素。
原始尝试代码
# 抓取2020赛季数据 jz_2020_raw_stats = requests.get('https://www.basketball-reference.com/teams/UTA/2020.html').text jz_2020_soup_stats = BeautifulSoup(jz_2020_raw_stats,'html.parser') jz_2020_soup_stats.prettify() # 提取所需列的值 pts = jz_2020_soup_stats.select("#team_and_opponent_sh > div > ul > li:nth-child(1) > span")[0].text print(pts) fg_pct = jz_2020_soup_stats.select("#team_and_opponent > tbody:nth-child(4) > tr:nth-child(1) > td:nth-child(6)")[0].text print(fg_pct) three_pct = jz_2020_soup_stats.select("#team_and_opponent > tbody:nth-child(4) > tr:nth-child(1) > td:nth-child(9)")[0].text tov = jz_2020_soup_stats.select("#team_and_opponent > tbody:nth-child(4) > tr:nth-child(1) > td:nth-child(22)")[0].text ft_pct = jz_2020_soup_stats.select("#team_and_opponent > tbody:nth-child(4) > tr:nth-child(1) > td:nth-child(15)")[0].text # 打印数值 print('PTS:', pts) print('FG%:', fg_pct) print('3P%:', three_pct) print('TOV:', tov) print('FT%:', ft_pct)
可行解决方案代码
page = requests.get(jz_2020_url) # 移除HTML注释标签,让表格内容可被解析 page = page.text.replace("<!--","").replace("-->","") soup = BeautifulSoup(page, 'html.parser') # 定位目标表格并转为字符串 table_html = str(soup.find("table", {"id": "team_and_opponent"})) # 使用pandas直接读取表格为DataFrame df = pd.read_html(table_html)[0] print(df)
补充说明
- 移除注释标签后,原本被隐藏的表格结构会被正常识别,此时可以用BeautifulSoup定位表格,再通过
pandas.read_html()快速将表格转为DataFrame,无需手动定位每个单元格。 - 若遇到类似静态HTML中内容被注释包裹的情况,都可以先移除注释标签再进行解析。
内容的提问来源于stack exchange,提问作者Alex
相关产品推荐
相关产品推荐

