使用Marshmallow序列化含变音符号的Python对象至JSON的编码问题
问题:Marshmallow序列化含变音符号的对象到JSON文件时,如何保留ä、ö、ü等原字符显示?
我想要将包含变音符号(如ä、ö、ü、Ä、Ö、Ü)的Python对象序列化为JSON文件,使用Marshmallow定义Schema。以下是代码片段:
from marshmallow import Schema, fields, post_load, post_dump import logging as log import os.path log.basicConfig(level=log.DEBUG) class Umlaut(): def __init__(self, name): self.name = name class UmlautSchema(Schema): name = fields.Str() @post_dump def post_dump(self, data, many=False): log.debug(data) # Umlauts are fine return data class Filehandling(): def write(self, u): pathToFile = os.path.abspath("/tmp/") schema = UmlautSchema() res = schema.dumps(u) log.debug(res) # Here we have 'Umlaut-\u00f6-\u00e4-\u00fc-\u00c4-\u00d6-\u00dc' file = os.path.join(pathToFile, "umlauts.json") with open(file, mode="w", encoding="UTF-8") as outfile: outfile.write(res) outfile.close() def run(): filehandling = Filehandling() u = Umlaut("Umlaut-ö-ä-ü-Ä-Ö-Ü") log.debug(u.name) # Umlauts are fine filehandling.write(u)
运行代码后,JSON文件中的变音符号显示为\u00f6这类转义字符,但post_dump方法的日志中变音符号显示正常,请问如何让JSON文件中正确显示ä、ö、ü等变音符号?
解决方案
问题核心是Marshmallow的dumps方法默认启用了JSON的ensure_ascii=True参数,该参数会把所有非ASCII字符转义为Unicode转义序列。要保留原变音符号,有两种修改方式:
方式1:调用dumps时直接传参
修改Filehandling类的write方法,在schema.dumps(u)中添加ensure_ascii=False参数:
res = schema.dumps(u, ensure_ascii=False)
修改后,序列化后的JSON字符串会直接保留ä、ö、ü等原字符,写入UTF-8编码的文件后即可正常显示。
方式2:在Schema类中设置全局默认参数
如果不想每次调用dumps都重复传参,可以在Schema的Meta类中配置默认的JSON序列化参数:
class UmlautSchema(Schema): name = fields.Str() @post_dump def post_dump(self, data, many=False): log.debug(data) return data class Meta: json_module_kwargs = {"ensure_ascii": False}
后续调用schema.dumps(u)时会自动应用ensure_ascii=False,同样能实现保留原变音符号的效果。
内容的提问来源于stack exchange,提问作者Stephan Mathys
相关产品推荐
相关产品推荐

