You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python如何从HTML字符串提取各文本片段的对应样式属性

解决方案

你之前的实现问题在于从外层标签向内遍历,依赖删除节点的操作提取内容,天然会漏掉根节点下的直接文本,也无法正确处理多层嵌套的样式优先级。
正确的实现思路是从文本节点反向回溯:所有独立的文本片段在BeautifulSoup里都是NavigableString类型的节点,不管它在标签外还是嵌套在多少层span里,都可以直接遍历拿到;对每个文本节点,逐层向上遍历所有父级标签,收集父级上的样式,内层标签的样式优先级高于外层,未被任何标签设置的属性保持默认None即可。

可运行代码

from bs4 import BeautifulSoup, NavigableString

def extract_text_with_styles(html_content):
    # 定义默认样式,所有属性初始为None
    base_style = {
        "color": None,
        "font-style": None,
        "font-weight": None,
        "text-decoration": None
    }
    soup = BeautifulSoup(html_content, "html.parser")
    res = []

    # 遍历DOM中所有文本节点
    for text_node in soup.find_all(string=True):
        current_style = base_style.copy()
        # 从当前文本的直接父节点开始向上回溯
        parent = text_node.parent
        while parent and parent.name != "html":
            # 只处理带style属性的span标签
            if parent.name == "span" and parent.has_attr("style"):
                # 解析style字符串里的所有属性
                for decl in parent["style"].split(";"):
                    decl = decl.strip()
                    if not decl:
                        continue
                    prop, val = decl.split(":", 1)
                    prop = prop.strip()
                    val = val.strip()
                    # 仅收集需要的4种属性,内层已设置的属性不会被外层覆盖
                    if prop in current_style and current_style[prop] is None:
                        current_style[prop] = val
            parent = parent.parent
        res.append({
            "text": str(text_node),
            "styles": current_style
        })
    return res

# 测试
test_html = 'Normal<span style="font-weight: bold;">Bold <span style="font-style: italic;">BoldAndItalic</span></span><span style="font-style: italic;">Italic</span>'
print(extract_text_with_styles(test_html))

逻辑说明

  • 不需要修改/删除原DOM节点,避免结构破坏导致的内容遗漏
  • 自动适配任意层级的span嵌套,样式继承逻辑完全符合浏览器渲染规则:内层span的样式会覆盖外层同名样式,未设置的样式从外层继承
  • 所有文本片段都会被捕获,包括span标签外的普通文本、多层嵌套span内的文本
  • 自动过滤不需要的样式属性,只保留要求的四类属性
    如果需要过滤完全为空的文本节点(比如HTML换行产生的空白字符),加一个判断跳过长度为0或者全为换行符的文本节点即可。

内容的提问来源于stack exchange,提问作者Tejasisamazing

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 16:32:30