BeautifulSoup4重复打印数据问题(函数仅调用一次)
解决BeautifulSoup解析高分页面时记录重复打印的问题
问题现象
- 单次调用
is_in_highscores()函数,每条高分记录却被重复打印两次 - 测试用的递增计数显示正常,但输出结果每条重复一次
- 排查发现:每条记录对应的
<tr>节点内存在两个class="DoNotBreak"的<td>标签
原代码
def is_in_highscores(): world = "Monza" high_score_site = f"subtopic=highscores&world={world}&beprotection=-1&category={6}&profession=0¤tpage={6}" soup = high_score_suffix(high_score_site) high_score = soup.find_all(class_="DoNotBreak") for i in high_score: a = (i.findParent()) b = [element.text for element in i.findParent().contents] print(b)
问题根源
原代码通过find_all(class_="DoNotBreak")获取所有带该类的<td>标签,而每条记录的<tr>里有两个这样的<td>(角色名、职业),循环时同一个<tr>会被两次反查并输出,导致记录重复。
解决方案
直接定位每条记录对应的<tr>节点,避免从<td>反查父节点造成重复处理:
def is_in_highscores(): world = "Monza" high_score_site = f"subtopic=highscores&world={world}&beprotection=-1&category={6}&profession=0¤tpage={6}" soup = high_score_suffix(high_score_site) # 定位所有包含高分记录的<tr>节点 high_score_rows = soup.select("tr:has(.DoNotBreak)") for row in high_score_rows: # 提取该行所有<td>的文本内容 record = [td.text.strip() for td in row.find_all("td")] print(record)
修改说明
- 用
select("tr:has(.DoNotBreak)")精准筛选出包含目标<td>的记录行,确保只处理有效记录 - 直接遍历
<tr>节点,每个记录行仅被处理一次,彻底解决重复输出问题 - 使用
row.find_all("td")提取该行所有单元格文本,同时用strip()清理多余空格
补充信息
原网页中记录行的结构:
<tr style="background-color:#D4C0A1;"><td>300</td><td class="DoNotBreak"><a href="https://www.tibia.com/community/?subtopic=characters&name=Faje">Faje</a></td><td class="DoNotBreak">Elder Druid</td><td>Monza</td><td style="text-align: right;">637</td><td style="text-align: right;">4,283,568,941</td></tr>
原代码输出的重复结果:
['298', 'Loveable Wicked', 'Master Sorcerer', 'Monza', '638', '4,288,783,674'] ['298', 'Loveable Wicked', 'Master Sorcerer', 'Monza', '638', '4,288,783,674'] ['299', 'Sleepy Bzyku', 'Elite Knight', 'Monza', '637', '4,285,301,108'] ['299', 'Sleepy Bzyku', 'Elite Knight', 'Monza', '637', '4,285,301,108'] ['300', 'Faje', 'Elder Druid', 'Monza', '637', '4,283,568,941'] ['300', 'Faje', 'Elder Druid', 'Monza', '637', '4,283,568,941']
内容的提问来源于stack exchange,提问作者Kossano
相关产品推荐
相关产品推荐

