You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup合并连续标签:处理拆分单词的混乱HTML

解决BeautifulSoup合并拆分单词的连续标签问题

这问题我之前处理过类似的,这种把单个单词拆成多个独立标签的情况,大多是Word转HTML的产物,确实挺头疼。用BeautifulSoup可以通过遍历并合并连续的文本标签来解决,给你两种方案,按需选择:

方案1:针对指定标签的批量合并

适合你已经明确知道哪些标签被拆分(比如示例里的<b>和<span>),逻辑简单直接:

from bs4 import BeautifulSoup, NavigableString

# 示例混乱HTML
html = """
<b style="mso-bidi-font-weight:normal"><span style='font-size:14.0pt;mso-bidi-font-size:11.0pt;line-height:107%;font-family:"Times New Roman",serif;mso-fareast-font-family:"Times New Roman"'>I</span></b><b style="mso-bidi-font-weight:normal"><span style='font-family:"Times New Roman",serif;mso-fareast-font-family:"Times New Roman"'>NTRODUCT</span></b><b style="mso-bidi-font-weight:normal"><span style='font-family:"Times New Roman",serif;mso-fareast-font-family:"Times New Roman"'>ION</span></b>
"""
soup = BeautifulSoup(html, 'html.parser')

def merge_consecutive_text_tags(soup):
    # 这里可以添加你需要处理的内联标签,比如i、em等
    target_tags = ['b', 'span']
    for tag_name in target_tags:
        tags = soup.find_all(tag_name)
        i = 0
        while i < len(tags) - 1:
            current_tag = tags[i]
            next_tag = tags[i+1]
            
            # 跳过标签之间的空白文本节点(比如换行、空格)
            next_sibling = current_tag.next_sibling
            while next_sibling and isinstance(next_sibling, NavigableString) and next_sibling.strip() == '':
                next_sibling = next_sibling.next_sibling
            
            # 如果两个标签是相邻的兄弟节点,就合并
            if next_sibling == next_tag:
                # 保留原文本格式(比如空格),拼接两个标签的文本
                current_tag.string = current_tag.get_text(strip=False) + next_tag.get_text(strip=False)
                # 删除多余的标签
                next_tag.decompose()
                # 重新获取标签列表,因为已经删除了一个元素
                tags = soup.find_all(tag_name)
            else:
                i += 1
    return soup

# 执行合并
cleaned_soup = merge_consecutive_text_tags(soup)
print(cleaned_soup.prettify())

方案1说明:

  • 我们遍历指定的标签类型,检查每个标签的下一个兄弟节点是否是同类型标签(中间只允许有空白文本)
  • 合并后会保留第一个标签的样式属性,符合大多数场景的需求
  • 代码逻辑清晰,适合结构相对简单的HTML

方案2:递归处理嵌套标签

如果你的HTML有多层嵌套(比如<b>里套<span>,<span>里又套<span>),用递归函数可以更全面地处理所有层级的拆分标签:

from bs4 import BeautifulSoup, NavigableString

def merge_consecutive_tags_recursive(element):
    children = list(element.children)
    i = 0
    while i < len(children) - 1:
        current = children[i]
        next_child = children[i+1]
        
        # 先清理空白文本节点
        if isinstance(current, NavigableString) and current.strip() == '':
            current.decompose()
            children = list(element.children)
            continue
        if isinstance(next_child, NavigableString) and next_child.strip() == '':
            next_child.decompose()
            children = list(element.children)
            continue
        
        # 检查当前和下一个节点是否是同类型标签,且仅包含文本内容
        if (isinstance(current, BeautifulSoup.Tag) and isinstance(next_child, BeautifulSoup.Tag)
            and current.name == next_child.name
            and len(current.contents) == 1 and isinstance(current.contents[0], NavigableString)
            and len(next_child.contents) == 1 and isinstance(next_child.contents[0], NavigableString)):
            # 合并文本
            current.string = current.get_text(strip=False) + next_child.get_text(strip=False)
            next_child.decompose()
            children = list(element.children)
        else:
            # 递归处理子标签
            if isinstance(current, BeautifulSoup.Tag):
                merge_consecutive_tags_recursive(current)
            i += 1
    # 处理最后一个子节点的嵌套结构
    if children and isinstance(children[-1], BeautifulSoup.Tag):
        merge_consecutive_tags_recursive(children[-1])

# 加载并处理HTML
soup = BeautifulSoup(html, 'html.parser')
merge_consecutive_tags_recursive(soup)
print(soup.prettify())

方案2说明:

  • 递归遍历所有层级的标签,自动处理嵌套结构
  • 自动清理无意义的空白文本节点,让结构更整洁
  • 只合并同类型、仅包含文本的连续标签,避免误合并其他有子元素的标签

注意事项

  • 如果拆分的标签样式差异较大(比如一个是红色,一个是蓝色),合并后会保留第一个标签的样式。如果需要合并样式,你可以额外处理style属性,比如提取两个标签的共同样式并合并,但这需要更复杂的逻辑,一般场景下不需要。
  • 建议先用prettify()查看处理后的HTML结构,确保合并结果符合预期。

内容的提问来源于stack exchange,提问作者jma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:31:45