Python JSON转XML程序遇UnicodeDecodeError错误求助
解决JSON转XML时的UnicodeDecodeError问题
嗨,看你的情况,前560个文件都顺顺利利处理完,突然栽在UnicodeDecodeError上,这大概率是后面的JSON文件编码和你默认读取的编码不匹配导致的。咱们一步步拆解问题和解决办法:
问题根源
这个错误提示'charmap' codec can't decode byte 0x9d,说明Python在读取文件时用了系统默认的charmap编码(比如Windows上的cp1252),而这个编码压根不认识文件里的0x9d字节。前560个文件要么是标准UTF-8编码,要么刚好能被系统编码解析,后面的文件可能换了编码,或者包含了特殊字符。
你的代码片段(方便参考)
# -*- coding: utf-8 -*- ##IMPORT import codecs import string import sys from src.json2xml import Json2xml import unicodedata import os ##FUNCTION def fn_conversione(f): data = Json2xml.fromjsonfile('json//' + f).data data_object =...
解决方案
1. 强制指定UTF-8编码读取文件
第三方库json2xml的fromjsonfile方法可能默认用系统编码读取文件,咱们可以手动接管文件读取,指定编码后再传给库处理:
def fn_conversione(f): # 手动读取JSON文件,指定UTF-8编码,同时处理解码错误 with open('json//' + f, 'r', encoding='utf-8', errors='replace') as json_file: json_content = json_file.read() # 用fromstring方法传入读取好的内容 data = Json2xml.fromstring(json_content).data # 后续你的data_object处理逻辑...
这里的errors='replace'会把无法解码的字符替换成�,你也可以根据需求换成errors='ignore'(直接跳过错误字符),如果想严格校验就保留默认的strict(但会继续报错)。
2. 先检测文件实际编码(可选)
如果不确定文件的编码,可以用chardet库先检测:
# 先安装chardet:pip install chardet import chardet def fn_conversione(f): # 检测文件编码 with open('json//' + f, 'rb') as file: detect_result = chardet.detect(file.read()) file_encoding = detect_result.get('encoding', 'utf-8') # 没检测到就用UTF-8兜底 # 用检测到的编码读取文件 with open('json//' + f, 'r', encoding=file_encoding, errors='replace') as json_file: json_content = json_file.read() data = Json2xml.fromstring(json_content).data # 后续处理...
3. 修改后的完整示例
把这些整合到你的代码里,大概是这样:
# -*- coding: utf-8 -*- ##IMPORT import codecs import string import sys from src.json2xml import Json2xml import unicodedata import os import chardet ##FUNCTION def fn_conversione(f): try: # 检测文件编码 with open('json//' + f, 'rb') as file: detect_result = chardet.detect(file.read()) file_encoding = detect_result.get('encoding', 'utf-8') # 读取并转换 with open('json//' + f, 'r', encoding=file_encoding, errors='replace') as json_file: json_content = json_file.read() data = Json2xml.fromstring(json_content).data # 这里写你处理data_object的逻辑,比如生成XML xml_output = Json2xml(data).to_xml() # 保存XML文件时也指定UTF-8编码,避免后续出问题 xml_filename = f.replace('.json', '.xml') with open('xml//' + xml_filename, 'w', encoding='utf-8') as xml_file: xml_file.write(xml_output) print(f"✅ 成功处理文件: {f}") except Exception as e: print(f"❌ 处理文件 {f} 失败: {str(e)}") # 遍历json文件夹处理所有JSON文件 if __name__ == "__main__": if not os.path.exists('xml'): os.makedirs('xml') # 确保XML输出文件夹存在 for filename in os.listdir('json'): if filename.endswith('.json'): fn_conversione(filename)
小提示
- 保存XML文件时也记得指定
encoding='utf-8',避免后续读取XML时又踩编码的坑。 - 如果有个别文件编码特别奇怪,
chardet可能检测不准,这时候可以手动尝试utf-16、gbk等常见编码试试。
内容的提问来源于stack exchange,提问作者Gargantua
相关产品推荐
相关产品推荐

