You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python2中BeautifulSoup处理子标签及去重技术问询

解决BeautifulSoup清理HTML的两个问题

我来帮你搞定这两个问题,咱们一步步拆解:

问题1:嵌套标签重复输出的处理

你观察得没错,最后一条输出里的重复是因为find_all(["a", "b", "p"])会把所有匹配的标签都捞出来——包括嵌套在目标标签里的子标签(比如<a>里面的<b>)。要解决这个,我们只需要保留顶级的目标标签,也就是那些父标签不属于我们要处理的标签列表的元素。

把你代码里的scrape_selected_tags那行替换成这个列表推导式:

scrape_selected_tags = [tag for tag in soup.find_all(["a", "b", "p"]) if tag.parent.name not in ["a", "b", "p"]]

这样一来,嵌套在<a>里的<b>因为父标签是<a>(属于我们的目标标签列表),就会被排除,只会保留外层的<a><b>test</b></a>。

问题2:同一website_id下去重并整理到clean_html

要实现去重并保持原格式,我们需要先把处理后的HTML片段转成字符串(因为BeautifulSoup的ResultSet对象没法直接去重),再对每个website_id的内容去重,最后存入clean_html。

调整后的完整代码如下:

from bs4 import BeautifulSoup

html_dict = {"l0000": ["<a href='some url'>test</a>", "lol", "<a><b>test</b></a>"], "l0001":["<p>this is html</p>", "<p>this is html</p>"]}
clean_html = {}

for website_id, raw_html_list in html_dict.items():
    processed_items = []
    for raw_html in raw_html_list:
        soup = BeautifulSoup(raw_html, 'html.parser')
        # 移除目标标签的所有属性
        for tag in soup.find_all(["a", "b", "p"]):
            tag.attrs = {}
        # 筛选顶级目标标签(解决问题1)
        top_level_tags = [tag for tag in soup.find_all(["a", "b", "p"]) if tag.parent.name not in ["a", "b", "p"]]
        # 处理内容:有目标标签就转成字符串,纯文本直接保留
        if top_level_tags:
            processed_str = ''.join(str(tag) for tag in top_level_tags)
        else:
            processed_str = raw_html
        processed_items.append(processed_str)
    
    # 去重并保留原顺序
    unique_processed = []
    seen = set()
    for item in processed_items:
        if item not in seen:
            seen.add(item)
            unique_processed.append(item)
    
    # 存入clean_html,保持和原字典一致的格式
    clean_html[website_id] = unique_processed

# 验证结果
for website_id, content in clean_html.items():
    print(website_id, content)

代码说明:

  • 先逐个处理每个raw_html片段,清理标签属性后筛选出顶级标签,转成字符串;如果是纯文本(比如"lol")就直接保留原内容。
  • 去重用了保留原顺序的方法(集合会打乱顺序,所以用遍历+已见集合的方式)。
  • 最终clean_html的格式和原html_dict完全一致。

运行这段代码的输出会是:

l0000 ['<a>test</a>', 'lol', '<a><b>test</b></a>']
l0001 ['<p>this is html</p>']

内容的提问来源于stack exchange,提问作者user47467

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:40:45