Python中json.dumps处理含UTF-8编码元组的问题求助
服务器上的Python脚本从数据库获取元组格式的数据,其中包含德语名称(例如Weidmüller),获取到的数据显示为Weidm\xfcller(\xfc是ü的Latin-1编码,这一步本身正常)。但使用json.dumps()转换为JSON字符串时遇到以下问题:
- 使用
json.dumps(tableData, ensure_ascii=False)时,\xfc被转换为乱码� - 使用
json.dumps(tableData, ensure_ascii=True)时,抛出错误:UnicodeDecodeError: 'utf8' codec can't decode byte 0xfc in position 5: invalid start byte
需求:希望json.dumps能保留UTF-8编码的字符,直接传递\xfc给浏览器的JavaScript脚本解码,询问该方案是否可行,或当前处理思路是否有误。
完整代码
import MySQLdb ... # Open the data base and return a handle to it and its cursor dataBase, dbCursor = database.OpenDB() # Get data from the URL fieldStore = cgi.FieldStorage() selFieldName = selFieldValue = '' sqlQuery = 'SELECT * FROM %s' % (database.CompTableName) if ('fldName' in fieldStore) and ('fldValue' in fieldStore): fldName = fieldStore['fldName'].value fldValue = fieldStore['fldValue'].value sqlQuery += ' WHERE %s = \'%s\'' % (fldName,fldValue) if ('max' in fieldStore): maxRows = fieldStore['max'].value sqlQuery += ' LIMIT ' + maxRows # Get the selected data in the table as a list of lists rowsAffected = dbCursor.execute(sqlQuery) tableData = dbCursor.fetchall() # Close the database and return the results dataBase.close() jsonTableData = json.dumps(tableData,encoding='latin1',ensure_ascii=True) print jsonTableData
测试代码
tableData = (('item1', 'Jones',), ('item2', 'Weidm\xfcller')) jsonTableData = json.dumps(tableData,encoding='latin1',ensure_ascii=True) print jsonTableData
解决方案
问题根源
核心问题是字节串与Unicode字符串的编码混淆:从MySQLdb获取的Weidm\xfcller是Python 2中的str类型(即字节串),而json.dumps默认期望处理Unicode字符串。\xfc在UTF-8中是无效的单字节(UTF-8中ü的编码是双字节\xc3\xbc),但在Latin-1编码中正好对应ü字符,这也是测试代码中encoding='latin1'能临时生效的原因。
正确处理步骤
将数据库返回的字节串转为Unicode字符串
先把从数据库拿到的字节串按数据库实际使用的编码(这里是Latin-1)解码为Unicode,确保后续JSON处理的是标准字符串:# 替换原tableData = dbCursor.fetchall()的代码 tableData = [] for row in dbCursor.fetchall(): # 遍历元组中的每个元素,字节串转Unicode,非字符串类型保持原样 unicode_row = tuple(col.decode('latin1') if isinstance(col, str) else col for col in row) tableData.append(unicode_row)输出UTF-8编码的JSON
转为Unicode后,使用ensure_ascii=False让JSON保留原始字符,再编码为UTF-8输出:jsonTableData = json.dumps(tableData, ensure_ascii=False).encode('utf-8') print(jsonTableData)这样输出的JSON会包含ü的UTF-8编码
\xc3\xbc,浏览器的JavaScript会自动识别UTF-8编码并正确显示为ü。
关于“直接传递\xfc”的可行性
JSON规范要求字符串内容为Unicode编码,直接传递字节串\xfc不符合JSON标准。虽然JavaScript解析时可能会将\xfc当作Unicode转义字符(对应ü的Unicode码点U+00FC),但这种做法不规范,且容易引发其他编码问题。更可靠的方式是输出标准UTF-8编码的JSON,让浏览器自动处理解码。
测试代码验证
修改测试代码为Unicode字符串处理:
# 直接使用Unicode字符串 tableData = (('item1', u'Jones'), ('item2', u'Weidmüller')) # 或者从字节串转换 # tableData = (('item1', 'Jones'), ('item2', 'Weidm\xfcller'.decode('latin1'))) jsonTableData = json.dumps(tableData, ensure_ascii=False).encode('utf-8') print(jsonTableData)
输出结果为[["item1", "Jones"], ["item2", "Weidmüller"]],对应的UTF-8字节包含\xc3\xbc,浏览器可正确显示德语字符。
内容的提问来源于stack exchange,提问作者Davide Andrea

