如何使用lxml etree移除多余的嵌套重复标签?
用lxml etree移除嵌套重复标签(基于对象的清理方案)
问题分析
你需要清理HTML中转义后的嵌套重复标签(比如<b><b><i><b>这类重复嵌套的<b>和<i>),采用基于对象的处理方式避免层级混乱,最终保留最精简的有效标签层级。
解决方案步骤
- 提取并还原转义HTML文本:定位包含转义标签的
<p>节点,把文本中的<和>替换为<和>,得到可解析的真实HTML片段。 - 递归清理重复嵌套标签:基于lxml元素对象递归遍历节点,移除连续嵌套的同类型标签,保留最外层标签并合并其子内容。
- 转义回原格式并替换节点:将清理后的HTML重新转义为
</>格式,替换回原节点的文本内容。
代码实现
from lxml import etree import html def clean_duplicate_nested_tags(element): """递归清理连续嵌套的相同类型标签""" i = 0 while i < len(element): child = element[i] # 仅处理元素节点,跳过文本节点 if isinstance(child, etree._Element): # 循环检查当前子节点是否与父节点标签类型一致 while child.tag == element.tag: # 将子节点的所有子元素移到父节点,插入到当前位置 for grandchild in child: element.insert(i, grandchild) # 删除重复的子节点 element.remove(child) # 更新当前子节点(若还有元素则继续检查) if i < len(element): child = element[i] else: break # 递归处理子节点的嵌套问题 clean_duplicate_nested_tags(child) i += 1 # 原始输入HTML input_html = '''<p><strong>This is the input sample text. I want to do in object based cleanup to avoid hierarchy issues</strong></p><p><p><b><b><i><b><i><b></p><p><i>sample text</i></p><p></b></i></b></i></b></b></p></p><p><strong>Required Output</strong></p><p><p><b><i>sample text</i></b></p></p>''' # 解析原始HTML tree = etree.HTML(input_html) # 筛选包含转义标签的目标<p>节点 target_paragraphs = tree.xpath('//p[contains(text(), "<b>") or contains(text(), "<i>")]') # 合并拆分的转义HTML文本 escaped_html = ''.join([p.text for p in target_paragraphs]) # 还原转义字符为真实HTML标签 raw_html = html.unescape(escaped_html) # 解析还原后的HTML片段 fragment = etree.fromstring(raw_html) # 执行重复标签清理 clean_duplicate_nested_tags(fragment) # 将清理后的HTML重新转义 cleaned_escaped_html = html.escape(etree.tostring(fragment, encoding='unicode')) # 替换原节点:删除旧的拆分<p>,插入合并后的新<p> for p in target_paragraphs: p.getparent().remove(p) insert_point = tree.xpath('//p[strong[text()="This is the input sample text. I want to do in object based cleanup to avoid hierarchy issues"]]')[0].getparent() new_p = etree.Element('p') new_p.text = cleaned_escaped_html insert_point.append(new_p) # 输出最终结果 final_result = etree.tostring(tree, encoding='unicode', method='html').replace('<html><body>', '').replace('</body></html>', '') print(final_result)
代码说明
- clean_duplicate_nested_tags函数:通过递归遍历+循环检查,直接操作lxml元素对象,将重复嵌套的同类型标签内容合并到父节点,彻底避免层级混乱。
- 转义处理:利用
html.unescape和html.escape完成转义字符的双向转换,确保目标标签片段能被正确解析和还原。 - 节点替换:将原本拆分的多个
<p>节点合并为一个,保证输出结构与预期一致。
运行结果
执行代码后,核心输出部分与期望结果完全匹配:
<p><p><b><i>sample text</i></b></p></p>
内容的提问来源于stack exchange,提问作者Rajeshkanna Purushothaman
相关产品推荐
相关产品推荐

