Python读取XML值报错:不支持带编码声明的Unicode字符串
解决XML解析时的"Unicode strings with encoding declaration are not supported"错误
我之前踩过一模一样的坑!这个错误的核心原因很明确:你传给XML解析器的是已经被解码成Unicode的字符串,但字符串里还保留着<?xml version="1.0" encoding="UTF-8"?>这类编码声明——解析器拿到Unicode内容后,不需要再从声明里获取编码信息,所以会直接抛出冲突错误。
结合你要提取最后一个file值的需求,给你两个靠谱的解决方案:
方案一:直接传入字节流(推荐)
这是最规范的做法,让解析器自己处理编码声明。不管是读取本地文件还是获取网络响应,都用字节形式传入:
import xml.etree.ElementTree as ET # 本地文件:用二进制模式打开 with open("your_target.xml", "rb") as xml_file: tree = ET.parse(xml_file) root = tree.getroot() # 提取最后一个file节点的值(假设XML结构是类似<files><file>xxx</file>...</files>) # 用XPath定位所有file节点,取最后一个的文本 last_file_node = root.findall(".//file")[-1] print(last_file_node.text)
如果是从网络请求获取XML(比如用requests),直接用response.content(字节形式)而不是response.text:
import requests import xml.etree.ElementTree as ET response = requests.get("your_xml_url") tree = ET.fromstring(response.content) last_file = tree.findall(".//file")[-1].text
方案二:移除编码声明后解析Unicode字符串
如果已经拿到了带编码声明的Unicode字符串,手动去掉开头的声明部分再解析:
import xml.etree.ElementTree as ET # 假设这是你拿到的带编码声明的Unicode字符串 xml_content = '''<?xml version="1.0" encoding="UTF-8"?> <root> <file_list> <file>doc1.pdf</file> <file>doc2.pdf</file> <file>target_doc.pdf</file> </file_list> </root>''' # 去掉开头的XML编码声明 if xml_content.strip().startswith('<?xml'): xml_content = xml_content.split('?>', 1)[1].strip() # 解析处理后的字符串 root = ET.fromstring(xml_content) last_file = root.findall(".//file")[-1].text print(last_file)
为什么会出现这个错误?
简单来说:当你把XML内容转换成Unicode字符串时,Python已经完成了「字节→Unicode」的解码过程,此时XML里的encoding声明就成了无效信息——解析器会觉得“我已经拿到Unicode了,你又告诉我要按UTF-8解码,这矛盾啊”,所以就抛出了这个错误。
内容的提问来源于stack exchange,提问作者Dan
相关产品推荐
相关产品推荐

