Python2中BeautifulSoup处理子标签及去重技术问询
解决BeautifulSoup清理HTML的两个问题
我来帮你搞定这两个问题,咱们一步步拆解:
问题1:嵌套标签重复输出的处理
你观察得没错,最后一条输出里的重复是因为find_all(["a", "b", "p"])会把所有匹配的标签都捞出来——包括嵌套在目标标签里的子标签(比如<a>里面的<b>)。要解决这个,我们只需要保留顶级的目标标签,也就是那些父标签不属于我们要处理的标签列表的元素。
把你代码里的scrape_selected_tags那行替换成这个列表推导式:
scrape_selected_tags = [tag for tag in soup.find_all(["a", "b", "p"]) if tag.parent.name not in ["a", "b", "p"]]
这样一来,嵌套在<a>里的<b>因为父标签是<a>(属于我们的目标标签列表),就会被排除,只会保留外层的<a><b>test</b></a>。
问题2:同一website_id下去重并整理到clean_html
要实现去重并保持原格式,我们需要先把处理后的HTML片段转成字符串(因为BeautifulSoup的ResultSet对象没法直接去重),再对每个website_id的内容去重,最后存入clean_html。
调整后的完整代码如下:
from bs4 import BeautifulSoup html_dict = {"l0000": ["<a href='some url'>test</a>", "lol", "<a><b>test</b></a>"], "l0001":["<p>this is html</p>", "<p>this is html</p>"]} clean_html = {} for website_id, raw_html_list in html_dict.items(): processed_items = [] for raw_html in raw_html_list: soup = BeautifulSoup(raw_html, 'html.parser') # 移除目标标签的所有属性 for tag in soup.find_all(["a", "b", "p"]): tag.attrs = {} # 筛选顶级目标标签(解决问题1) top_level_tags = [tag for tag in soup.find_all(["a", "b", "p"]) if tag.parent.name not in ["a", "b", "p"]] # 处理内容:有目标标签就转成字符串,纯文本直接保留 if top_level_tags: processed_str = ''.join(str(tag) for tag in top_level_tags) else: processed_str = raw_html processed_items.append(processed_str) # 去重并保留原顺序 unique_processed = [] seen = set() for item in processed_items: if item not in seen: seen.add(item) unique_processed.append(item) # 存入clean_html,保持和原字典一致的格式 clean_html[website_id] = unique_processed # 验证结果 for website_id, content in clean_html.items(): print(website_id, content)
代码说明:
- 先逐个处理每个raw_html片段,清理标签属性后筛选出顶级标签,转成字符串;如果是纯文本(比如"lol")就直接保留原内容。
- 去重用了保留原顺序的方法(集合会打乱顺序,所以用遍历+已见集合的方式)。
- 最终
clean_html的格式和原html_dict完全一致。
运行这段代码的输出会是:
l0000 ['<a>test</a>', 'lol', '<a><b>test</b></a>'] l0001 ['<p>this is html</p>']
内容的提问来源于stack exchange,提问作者user47467
相关产品推荐
相关产品推荐

