如何防止BeautifulSoup自动替换HTML转义实体 保留原始字符
需求背景
需要BeautifulSoup处理HTML时保留原始转义字符,不自动替换HTML实体为对应原生字符,测试代码如下:
soup1 = BeautifulSoup( "APOLLOE4: Early Alzheimer's disease study", "html.parser", ) html1 = str(soup1) print(html1)
实际运行输出
"APOLLOE4: Early Alzheimer's disease study"
预期输出
"APOLLOE4: Early Alzheimer's disease study"
仅调整默认输出格式化器参数无法实现效果,因为解析阶段html.parser已经把'解码为单引号,默认输出规则不会主动把单引号转回实体格式。
实现方案
方案1:自定义输出格式化器(推荐,无全局副作用)
继承BS4内置的HTMLFormatter类,重写转义方法,手动指定需要保留的实体映射,仅在输出阶段做转换,不影响解析逻辑:
from bs4 import BeautifulSoup from bs4.formatter import HTMLFormatter class PreserveEntityFormatter(HTMLFormatter): def escape(self, content): base_escaped = super().escape(content) # 按需添加需要保留的实体替换规则 return base_escaped.replace("'", "'") soup = BeautifulSoup( "APOLLOE4: Early Alzheimer's disease study", "html.parser" ) # 输出时传入自定义格式化器 result = soup.decode(formatter=PreserveEntityFormatter(encoding="utf-8")) print(result)
运行后输出完全匹配预期,你可以根据业务需要,在escape方法里追加其他字符到实体的替换规则。
方案2:修改解析阶段实体处理逻辑(适合全量保留实体场景)
直接替换html.parser构建器的实体处理函数,让解析阶段就不做实体解码,原样保留所有实体字符串:
from bs4 import BeautifulSoup import bs4.builder._htmlparser as bs4_parser from html.entities import name2codepoint def keep_raw_entity(name): if name.startswith("#"): return f"&{name};" return f"&{name};" if name in name2codepoint else f"&{name};" # 替换实体处理钩子 bs4_parser.handle_entityref = lambda ent: keep_raw_entity(ent) bs4_parser.handle_charref = lambda ent: keep_raw_entity(f"#{ent}") soup = BeautifulSoup( "APOLLOE4: Early Alzheimer's disease study", "html.parser" ) print(str(soup))
注意:该方案会全局修改html.parser的实体解析逻辑,如果同一段代码里有其他需要正常解码HTML实体的场景,不要使用该方案。
内容的提问来源于stack exchange,提问作者Przemek Baj
相关产品推荐
相关产品推荐

