使用BeautifulSoup按H2标签拆分扁平化HTML为section区块的实现方法
扁平HTML按H2标签拆分section区块的实现方案
下面提供两种主流实现方式,均无需手动全量线性遍历DOM树维护状态,逻辑简洁易维护。
BS4 实现方案
利用BS4原生的节点兄弟关系属性处理,是最适配需求的优雅方案:
from bs4 import BeautifulSoup # 输入你的原始HTML raw_html = """ <h2>section 1</h2> <p>paragraph 1 in section 1</p> <ul> <li>list item 1</li> <li>list item 2</li> </ul> <h2>section 2</h2> <p>paragraph 1 in section 2</p> <div>div content in section 2</div> """ # 解析HTML soup = BeautifulSoup(raw_html, 'html.parser') # 获取所有顶层H2标签 all_h2 = soup.find_all('h2') for h2 in all_h2: # 创建新的section节点 section = soup.new_tag('section') # 将section插入到当前h2的前面 h2.insert_before(section) # 把当前h2移入section section.append(h2) # 遍历h2后续所有兄弟节点,直到碰到下一个h2停止 current_sibling = h2.next_sibling while current_sibling: # 提前存下下一个节点,避免移动节点后引用丢失 next_node = current_sibling.next_sibling # 碰到下一个h2就终止遍历 if current_sibling.name == 'h2': break # 把当前节点移入section section.append(current_sibling) current_sibling = next_node # 输出格式化后的结果 print(soup.prettify())
运行后输出的结构完全符合预期,支持任意数量的H2区块,也能正常处理H2之间的纯文本节点、任意类型的元素节点。
XPath 实现方案(基于lxml)
如果偏好使用XPath,可以用preceding-sibling轴做节点分组:
from lxml import etree raw_html = """ <h2>section 1</h2> <p>paragraph 1 in section 1</p> <ul> <li>list item 1</li> <li>list item 2</li> </ul> <h2>section 2</h2> <p>paragraph 1 in section 2</p> <div>div content in section 2</div> """ root = etree.HTML(raw_html) body = root.xpath('./body')[0] all_h2 = root.xpath('//h2') for idx, h2 in enumerate(all_h2): # 选取当前h2 + 后续直到下一个h2之前的所有节点 group_nodes = root.xpath(f'//node()[count(preceding-sibling::h2) = {idx}]') # 创建section节点 section = etree.Element('section') # 替换原有节点位置 h2.addprevious(section) for node in group_nodes: section.append(node) # 输出结果 print(etree.tostring(body, encoding='unicode', pretty_print=True))
内容的提问来源于stack exchange,提问作者masroore
相关产品推荐
相关产品推荐

