如何禁用Python html.parser中的实体引用解码功能?
解决html.parser.HTMLParser属性实体自动解码的问题
标准的html.parser.HTMLParser没有直接禁用属性实体解码的开关——属性值的实体解码是解析器内置的逻辑,且这部分逻辑未暴露可配置选项。即使设置convert_charrefs=False,也仅影响文本节点的实体处理,属性里的实体解码仍会自动完成,且不会触发handle_entityref或handle_charref方法,导致无法干预该行为。
针对你遇到的URL解析错误(比如¶ms被误解析为¶ms),可以采用以下两种方案解决:
方案一:预处理HTML字符串,临时替换属性中的&
先把HTML中属性值里的&替换为临时占位符,避免解析器将其识别为实体开头,待解析完成后再替换回&。
示例代码:
import re from html.parser import HTMLParser from html.entities import name2codepoint # 预处理HTML:替换属性值中可能被误解析的& def preprocess_html(html): # 匹配属性值内的独立&(排除已规范编码的实体开头) return re.sub(r'(?<=["\'])&(?!amp;|lt;|gt;|quot;|apos;)', '&temp_amp;', html) class MyHTMLParser(HTMLParser): def handle_starttag(self, tag, attrs): print("Start tag:", tag) processed_attrs = [] for attr_name, attr_value in attrs: # 还原临时占位符为& fixed_value = attr_value.replace('&temp_amp;', '&') processed_attrs.append((attr_name, fixed_value)) print(" attr:", (attr_name, fixed_value)) def handle_endtag(self, tag): print("End tag :", tag) def handle_data(self, data): if data.strip(): print("Data :", data) def handle_entityref(self, name): c = chr(name2codepoint[name]) print("Named ent:", c) def handle_charref(self, name): if name.startswith('x'): c = chr(int(name[1:], 16)) else: c = chr(int(name)) print("Num ent :", c) html = """ <ul> <li>¶ & ¶</a> <li><a href="index.php?option=view&id=108">First link</a> <li><a href="index.php?option=view¶ms=28">Second link</a> </ul> """ # 预处理后再解析 processed_html = preprocess_html(html) parser = MyHTMLParser(convert_charrefs=False) parser.feed(processed_html)
执行后输出(URL被正确保留):
Start tag: ul Start tag: li Named ent: ¶ Named ent: & Num ent : ¶ End tag : a Start tag: li Start tag: a attr: ('href', 'index.php?option=view&id=108') Data : First link End tag : a Start tag: li Start tag: a attr: ('href', 'index.php?option=view¶ms=28') Data : Second link End tag : a End tag : ul
方案二:使用第三方解析库
如果不想自行处理预处理逻辑,可以使用第三方库(如BeautifulSoup)配合html.parser解析器,它会自动兼容不规范的HTML,避免实体误解析:
from bs4 import BeautifulSoup html = """ <ul> <li>¶ & ¶</a> <li><a href="index.php?option=view&id=108">First link</a> <li><a href="index.php?option=view¶ms=28">Second link</a> </ul> """ soup = BeautifulSoup(html, 'html.parser') for a_tag in soup.find_all('a'): print("href:", a_tag['href'])
输出:
href: index.php?option=view&id=108 href: index.php?option=view¶ms=28
内容的提问来源于stack exchange,提问作者MacFreek
相关产品推荐
相关产品推荐

