使用BeautifulSoup的.decompose()方法无法移除<tr class='thead'>元素的问题求助
使用BeautifulSoup的.decompose()方法无法移除元素的问题求助
嘿,我来帮你排查这个问题!你说用.decompose()删不掉<tr class="thead">的行,大概率是这两个小细节没处理到位:
问题1:只移除了第一个匹配的元素
soup.find()只会定位到页面中第一个符合tr.thead条件的元素,但NBA的分区统计表格里,每个分区(比如西北区、太平洋区)都会有一个这样的表头行,剩下的那些行依然留在表格里,所以用pandas读取时还是会看到它们。
问题2:没有限定查找范围到目标表格
你先全局查找tr.thead再移除,可能会误操作页面里其他无关的元素,或者反过来,目标表格里的tr.thead没被精准定位到。
修正后的代码
把你的代码改成这样试试:
dfs_player = [] for year in years: with open("player/{}.html".format(year)) as f: page = f.read() soup = BeautifulSoup(page, "html.parser") # 先精准定位到你要处理的表格容器 player_table_container = soup.find(id="div_per_game_stats") # 找到容器内所有的tr.thead元素,逐个移除 for thead_row in player_table_container.find_all("tr", class_="thead"): thead_row.decompose() # 再用pandas读取处理后的表格 player = pd.read_html(str(player_table_container))[0] player['Year'] = year dfs_player.append(player)
额外小提示
如果担心某些页面里没有这些tr.thead元素导致报错,可以加个简单的判断:
for thead_row in player_table_container.find_all("tr", class_="thead"): if thead_row: thead_row.decompose()
这样即使找不到目标元素,代码也能正常运行~
备注:内容来源于stack exchange,提问作者Rangga Buwana
相关产品推荐
相关产品推荐

