You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python JSON转XML程序遇UnicodeDecodeError错误求助

解决JSON转XML时的UnicodeDecodeError问题

嗨,看你的情况,前560个文件都顺顺利利处理完,突然栽在UnicodeDecodeError上,这大概率是后面的JSON文件编码和你默认读取的编码不匹配导致的。咱们一步步拆解问题和解决办法:

问题根源

这个错误提示'charmap' codec can't decode byte 0x9d,说明Python在读取文件时用了系统默认的charmap编码(比如Windows上的cp1252),而这个编码压根不认识文件里的0x9d字节。前560个文件要么是标准UTF-8编码,要么刚好能被系统编码解析,后面的文件可能换了编码,或者包含了特殊字符。

你的代码片段(方便参考)

# -*- coding: utf-8 -*-
##IMPORT
import codecs
import string
import sys
from src.json2xml import Json2xml
import unicodedata
import os
##FUNCTION
def fn_conversione(f):
    data = Json2xml.fromjsonfile('json//' + f).data
    data_object =...

解决方案

1. 强制指定UTF-8编码读取文件

第三方库json2xml的fromjsonfile方法可能默认用系统编码读取文件,咱们可以手动接管文件读取,指定编码后再传给库处理:

def fn_conversione(f):
    # 手动读取JSON文件,指定UTF-8编码,同时处理解码错误
    with open('json//' + f, 'r', encoding='utf-8', errors='replace') as json_file:
        json_content = json_file.read()
    # 用fromstring方法传入读取好的内容
    data = Json2xml.fromstring(json_content).data
    # 后续你的data_object处理逻辑...

这里的errors='replace'会把无法解码的字符替换成�,你也可以根据需求换成errors='ignore'(直接跳过错误字符),如果想严格校验就保留默认的strict(但会继续报错)。

2. 先检测文件实际编码(可选)

如果不确定文件的编码,可以用chardet库先检测:

# 先安装chardet:pip install chardet
import chardet

def fn_conversione(f):
    # 检测文件编码
    with open('json//' + f, 'rb') as file:
        detect_result = chardet.detect(file.read())
        file_encoding = detect_result.get('encoding', 'utf-8')  # 没检测到就用UTF-8兜底
    
    # 用检测到的编码读取文件
    with open('json//' + f, 'r', encoding=file_encoding, errors='replace') as json_file:
        json_content = json_file.read()
    data = Json2xml.fromstring(json_content).data
    # 后续处理...

3. 修改后的完整示例

把这些整合到你的代码里,大概是这样:

# -*- coding: utf-8 -*-
##IMPORT
import codecs
import string
import sys
from src.json2xml import Json2xml
import unicodedata
import os
import chardet

##FUNCTION
def fn_conversione(f):
    try:
        # 检测文件编码
        with open('json//' + f, 'rb') as file:
            detect_result = chardet.detect(file.read())
            file_encoding = detect_result.get('encoding', 'utf-8')
        
        # 读取并转换
        with open('json//' + f, 'r', encoding=file_encoding, errors='replace') as json_file:
            json_content = json_file.read()
        data = Json2xml.fromstring(json_content).data
        
        # 这里写你处理data_object的逻辑,比如生成XML
        xml_output = Json2xml(data).to_xml()
        # 保存XML文件时也指定UTF-8编码,避免后续出问题
        xml_filename = f.replace('.json', '.xml')
        with open('xml//' + xml_filename, 'w', encoding='utf-8') as xml_file:
            xml_file.write(xml_output)
        
        print(f"✅ 成功处理文件: {f}")
    except Exception as e:
        print(f"❌ 处理文件 {f} 失败: {str(e)}")

# 遍历json文件夹处理所有JSON文件
if __name__ == "__main__":
    if not os.path.exists('xml'):
        os.makedirs('xml')  # 确保XML输出文件夹存在
    for filename in os.listdir('json'):
        if filename.endswith('.json'):
            fn_conversione(filename)

小提示

  • 保存XML文件时也记得指定encoding='utf-8',避免后续读取XML时又踩编码的坑。
  • 如果有个别文件编码特别奇怪,chardet可能检测不准,这时候可以手动尝试utf-16、gbk等常见编码试试。

内容的提问来源于stack exchange,提问作者Gargantua

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:25:10