You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何禁用Python html.parser中的实体引用解码功能?

解决html.parser.HTMLParser属性实体自动解码的问题

标准的html.parser.HTMLParser没有直接禁用属性实体解码的开关——属性值的实体解码是解析器内置的逻辑,且这部分逻辑未暴露可配置选项。即使设置convert_charrefs=False,也仅影响文本节点的实体处理,属性里的实体解码仍会自动完成,且不会触发handle_entityref或handle_charref方法,导致无法干预该行为。

针对你遇到的URL解析错误(比如&params被误解析为¶ms),可以采用以下两种方案解决:

方案一:预处理HTML字符串,临时替换属性中的&

先把HTML中属性值里的&替换为临时占位符,避免解析器将其识别为实体开头,待解析完成后再替换回&。

示例代码:

import re
from html.parser import HTMLParser
from html.entities import name2codepoint

# 预处理HTML:替换属性值中可能被误解析的&
def preprocess_html(html):
    # 匹配属性值内的独立&(排除已规范编码的实体开头)
    return re.sub(r'(?<=["\'])&(?!amp;|lt;|gt;|quot;|apos;)', '&temp_amp;', html)

class MyHTMLParser(HTMLParser):
    def handle_starttag(self, tag, attrs):
        print("Start tag:", tag)
        processed_attrs = []
        for attr_name, attr_value in attrs:
            # 还原临时占位符为&
            fixed_value = attr_value.replace('&temp_amp;', '&')
            processed_attrs.append((attr_name, fixed_value))
            print("     attr:", (attr_name, fixed_value))
    def handle_endtag(self, tag):
        print("End tag  :", tag)
    def handle_data(self, data):
        if data.strip():
            print("Data     :", data)
    def handle_entityref(self, name):
        c = chr(name2codepoint[name])
        print("Named ent:", c)
    def handle_charref(self, name):
        if name.startswith('x'):
            c = chr(int(name[1:], 16))
        else:
            c = chr(int(name))
        print("Num ent  :", c)

html = """
<ul>
  <li>&para; &amp; &#xb6;</a>
  <li><a href="index.php?option=view&id=108">First link</a>
  <li><a href="index.php?option=view&params=28">Second link</a>
</ul>
"""

# 预处理后再解析
processed_html = preprocess_html(html)
parser = MyHTMLParser(convert_charrefs=False)
parser.feed(processed_html)

执行后输出(URL被正确保留):

Start tag: ul
Start tag: li
Named ent: ¶
Named ent: &
Num ent  : ¶
End tag  : a
Start tag: li
Start tag: a
     attr: ('href', 'index.php?option=view&id=108')
Data     : First link
End tag  : a
Start tag: li
Start tag: a
     attr: ('href', 'index.php?option=view&params=28')
Data     : Second link
End tag  : a
End tag  : ul

方案二:使用第三方解析库

如果不想自行处理预处理逻辑,可以使用第三方库(如BeautifulSoup)配合html.parser解析器,它会自动兼容不规范的HTML,避免实体误解析:

from bs4 import BeautifulSoup

html = """
<ul>
  <li>&para; &amp; &#xb6;</a>
  <li><a href="index.php?option=view&id=108">First link</a>
  <li><a href="index.php?option=view&params=28">Second link</a>
</ul>
"""

soup = BeautifulSoup(html, 'html.parser')
for a_tag in soup.find_all('a'):
    print("href:", a_tag['href'])

输出:

href: index.php?option=view&id=108
href: index.php?option=view&params=28

内容的提问来源于stack exchange,提问作者MacFreek

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 04:57:31