You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用bleach库清理HTML时,已允许dir属性仍被剥离是什么原因?

问题原因

你的allowed_attributes配置格式不符合bleach的要求。bleach的allowed_attributes字典规则是:

  • 字典的键为标签名称,值为该标签允许的属性配置
    你当前把属性名dir作为了字典的键,所以规则不会匹配到<p>标签上的dir属性,导致属性被过滤。
修正方案

方案1:给指定标签开放dir属性

如果仅需要部分标签(如p、span)支持dir属性,写法如下:
不需要限制dir取值的写法:

allowed_attributes = {
    'a': ['href', 'title'],
    'p': ['dir'],
    'span': ['dir']
}

如果需要限制dir的取值只能为rtl或ltr,用嵌套字典配置校验规则:

allowed_attributes = {
    'a': ['href', 'title'],
    'p': {
        'dir': ['rtl', 'ltr']
    },
    'span': {
        'dir': ['rtl', 'ltr']
    }
}

方案2:给所有允许的标签开放dir属性

如果需要所有在allowed_tags里的标签都支持dir属性,可以用通配符*匹配所有标签:
不需要限制取值的写法:

allowed_attributes = {
    'a': ['href', 'title'],
    '*': ['dir']
}

限制dir取值的写法:

allowed_attributes = {
    'a': ['href', 'title'],
    '*': {
        'dir': ['rtl', 'ltr']
    }
}
修正后完整运行代码
import bleach

string = """<p dir="rtl">asdasdasd <span>asdasdasd</span> asdsadasdsad .<br data-mce-bogus="1"></p>"""


def strip_invalid_html(html):
    """ strips invalid tags/attributes """

    allowed_tags = [
        'p', 'a', 'blockquote',
        'h1', 'h2', 'h3', 'h4', 'h5',
        'strong', 'em',
        'br',
        'span',
    ]

    allowed_attributes = {
        'a': ['href', 'title'],
        '*': {
            'dir': ['rtl', 'ltr']
        }
    }

    cleaned_html = bleach.clean(
        html,
        attributes=allowed_attributes,
        strip=True,
        tags=allowed_tags
    )

    print(cleaned_html)

strip_invalid_html(string)

运行后输出结果:

<p dir="rtl">asdasdasd <span>asdasdasd</span> asdsadasdsad .<br></p>

内容的提问来源于stack exchange,提问作者user12758446

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 12:18:03