使用BeautifulSoup合并连续标签:处理拆分单词的混乱HTML
解决BeautifulSoup合并拆分单词的连续标签问题
这问题我之前处理过类似的,这种把单个单词拆成多个独立标签的情况,大多是Word转HTML的产物,确实挺头疼。用BeautifulSoup可以通过遍历并合并连续的文本标签来解决,给你两种方案,按需选择:
方案1:针对指定标签的批量合并
适合你已经明确知道哪些标签被拆分(比如示例里的<b>和<span>),逻辑简单直接:
from bs4 import BeautifulSoup, NavigableString # 示例混乱HTML html = """ <b style="mso-bidi-font-weight:normal"><span style='font-size:14.0pt;mso-bidi-font-size:11.0pt;line-height:107%;font-family:"Times New Roman",serif;mso-fareast-font-family:"Times New Roman"'>I</span></b><b style="mso-bidi-font-weight:normal"><span style='font-family:"Times New Roman",serif;mso-fareast-font-family:"Times New Roman"'>NTRODUCT</span></b><b style="mso-bidi-font-weight:normal"><span style='font-family:"Times New Roman",serif;mso-fareast-font-family:"Times New Roman"'>ION</span></b> """ soup = BeautifulSoup(html, 'html.parser') def merge_consecutive_text_tags(soup): # 这里可以添加你需要处理的内联标签,比如i、em等 target_tags = ['b', 'span'] for tag_name in target_tags: tags = soup.find_all(tag_name) i = 0 while i < len(tags) - 1: current_tag = tags[i] next_tag = tags[i+1] # 跳过标签之间的空白文本节点(比如换行、空格) next_sibling = current_tag.next_sibling while next_sibling and isinstance(next_sibling, NavigableString) and next_sibling.strip() == '': next_sibling = next_sibling.next_sibling # 如果两个标签是相邻的兄弟节点,就合并 if next_sibling == next_tag: # 保留原文本格式(比如空格),拼接两个标签的文本 current_tag.string = current_tag.get_text(strip=False) + next_tag.get_text(strip=False) # 删除多余的标签 next_tag.decompose() # 重新获取标签列表,因为已经删除了一个元素 tags = soup.find_all(tag_name) else: i += 1 return soup # 执行合并 cleaned_soup = merge_consecutive_text_tags(soup) print(cleaned_soup.prettify())
方案1说明:
- 我们遍历指定的标签类型,检查每个标签的下一个兄弟节点是否是同类型标签(中间只允许有空白文本)
- 合并后会保留第一个标签的样式属性,符合大多数场景的需求
- 代码逻辑清晰,适合结构相对简单的HTML
方案2:递归处理嵌套标签
如果你的HTML有多层嵌套(比如<b>里套<span>,<span>里又套<span>),用递归函数可以更全面地处理所有层级的拆分标签:
from bs4 import BeautifulSoup, NavigableString def merge_consecutive_tags_recursive(element): children = list(element.children) i = 0 while i < len(children) - 1: current = children[i] next_child = children[i+1] # 先清理空白文本节点 if isinstance(current, NavigableString) and current.strip() == '': current.decompose() children = list(element.children) continue if isinstance(next_child, NavigableString) and next_child.strip() == '': next_child.decompose() children = list(element.children) continue # 检查当前和下一个节点是否是同类型标签,且仅包含文本内容 if (isinstance(current, BeautifulSoup.Tag) and isinstance(next_child, BeautifulSoup.Tag) and current.name == next_child.name and len(current.contents) == 1 and isinstance(current.contents[0], NavigableString) and len(next_child.contents) == 1 and isinstance(next_child.contents[0], NavigableString)): # 合并文本 current.string = current.get_text(strip=False) + next_child.get_text(strip=False) next_child.decompose() children = list(element.children) else: # 递归处理子标签 if isinstance(current, BeautifulSoup.Tag): merge_consecutive_tags_recursive(current) i += 1 # 处理最后一个子节点的嵌套结构 if children and isinstance(children[-1], BeautifulSoup.Tag): merge_consecutive_tags_recursive(children[-1]) # 加载并处理HTML soup = BeautifulSoup(html, 'html.parser') merge_consecutive_tags_recursive(soup) print(soup.prettify())
方案2说明:
- 递归遍历所有层级的标签,自动处理嵌套结构
- 自动清理无意义的空白文本节点,让结构更整洁
- 只合并同类型、仅包含文本的连续标签,避免误合并其他有子元素的标签
注意事项
- 如果拆分的标签样式差异较大(比如一个是红色,一个是蓝色),合并后会保留第一个标签的样式。如果需要合并样式,你可以额外处理
style属性,比如提取两个标签的共同样式并合并,但这需要更复杂的逻辑,一般场景下不需要。 - 建议先用
prettify()查看处理后的HTML结构,确保合并结果符合预期。
内容的提问来源于stack exchange,提问作者jma
相关产品推荐
相关产品推荐

