You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何防止BeautifulSoup自动替换HTML转义实体 保留原始字符

需求背景

需要BeautifulSoup处理HTML时保留原始转义字符,不自动替换HTML实体为对应原生字符,测试代码如下:

soup1 = BeautifulSoup(
    "APOLLOE4: Early Alzheimer's disease study",
    "html.parser",
)
html1 = str(soup1)
print(html1)

实际运行输出

"APOLLOE4: Early Alzheimer's disease study"

预期输出

"APOLLOE4: Early Alzheimer's disease study"

仅调整默认输出格式化器参数无法实现效果,因为解析阶段html.parser已经把'解码为单引号,默认输出规则不会主动把单引号转回实体格式。

实现方案

方案1:自定义输出格式化器(推荐,无全局副作用)

继承BS4内置的HTMLFormatter类,重写转义方法,手动指定需要保留的实体映射,仅在输出阶段做转换,不影响解析逻辑:

from bs4 import BeautifulSoup
from bs4.formatter import HTMLFormatter

class PreserveEntityFormatter(HTMLFormatter):
    def escape(self, content):
        base_escaped = super().escape(content)
        # 按需添加需要保留的实体替换规则
        return base_escaped.replace("'", "'")

soup = BeautifulSoup(
    "APOLLOE4: Early Alzheimer's disease study",
    "html.parser"
)
# 输出时传入自定义格式化器
result = soup.decode(formatter=PreserveEntityFormatter(encoding="utf-8"))
print(result)

运行后输出完全匹配预期,你可以根据业务需要,在escape方法里追加其他字符到实体的替换规则。

方案2:修改解析阶段实体处理逻辑(适合全量保留实体场景)

直接替换html.parser构建器的实体处理函数,让解析阶段就不做实体解码,原样保留所有实体字符串:

from bs4 import BeautifulSoup
import bs4.builder._htmlparser as bs4_parser
from html.entities import name2codepoint

def keep_raw_entity(name):
    if name.startswith("#"):
        return f"&{name};"
    return f"&{name};" if name in name2codepoint else f"&{name};"

# 替换实体处理钩子
bs4_parser.handle_entityref = lambda ent: keep_raw_entity(ent)
bs4_parser.handle_charref = lambda ent: keep_raw_entity(f"#{ent}")

soup = BeautifulSoup(
    "APOLLOE4: Early Alzheimer's disease study",
    "html.parser"
)
print(str(soup))

注意:该方案会全局修改html.parser的实体解析逻辑,如果同一段代码里有其他需要正常解码HTML实体的场景,不要使用该方案。

内容的提问来源于stack exchange,提问作者Przemek Baj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 05:33:16