Python调用API获取含Unicode字符的球员姓氏显示异常排查
问题
调用返回足球运动员姓氏的API时,其中一名球员姓氏含字符“ć”,调用后输出的姓氏附带Unicode转义序列:
>>> last_name = (json.dumps(response["response"][2]["player"]["lastname"])) >>> print(last_name) "Mitrovi\u0107" >>> print(type(last_name)) <class 'str'>
但将该输出内容复制粘贴到变量中单独使用时,打印显示正常:
>>> print("Mitrovi\u0107") Mitrović >>> print(type("Mitrovi\u0107")) <class 'str'>
请问API接口调用及返回的字符串存在什么问题?
原因与解决方法
问题出在你错误使用了json.dumps()处理API返回的字符串。json.dumps()的作用是把Python对象序列化为JSON格式的字符串,它会自动将非ASCII字符(比如“ć”)转换为Unicode转义序列(\u0107),同时给字符串添加双引号,所以你得到的是一个JSON格式的字符串字面量,而非原始的姓氏字符串。
直接复制粘贴"Mitrovi\u0107"到代码里时,Python解释器会自动解析字符串中的Unicode转义序列,将其转换为对应的字符“ć”,因此打印显示正常。
正确的处理方式是直接提取API返回的原始字符串,无需使用json.dumps():
>>> last_name = response["response"][2]["player"]["lastname"] >>> print(last_name) Mitrović
如果因特定场景必须使用json.dumps(),后续可以用json.loads()将序列化后的字符串重新解析,还原为原始字符:
>>> last_name = json.dumps(response["response"][2]["player"]["lastname"]) >>> last_name = json.loads(last_name) >>> print(last_name) Mitrović
内容的提问来源于stack exchange,提问作者sq89
相关产品推荐
相关产品推荐

