Python读取Facebook导出JSON时的字符串与字节类型混淆问题
Facebook聊天记录JSON读取的类型混淆问题分析与解决
问题重现
读取Facebook导出的聊天记录JSON文件时,单独访问列表首个元素的content字段能正常输出,但遍历打印每个消息对象时报错:
import json jsonMessages = [] with open("Facebook/message_1.json", 'rb') as file1: json1 = json.load(file1) jsonMessages.extend(json1['messages']) jsonMessages[0]['content'] # 正常输出消息内容 for msg in jsonMessages: print(msg) # 触发类型错误
报错信息:
--------------------------------------------------------------------------- TypeError Traceback (most recent call last) Cell In[170], line 10 7 jsonMessages[0]['content'] 9 for msg in jsonMessages: ---> 10 print(msg) File <frozen codecs>:378, in write(self, object) File ~\AppData\Local\anaconda3\Lib\site-packages\ipykernel\iostream.py:622, in OutStream.write(self, string) 620 if not isinstance(string, str): 621 msg = f"write() argument must be str, not {type(string)}" --> 622 raise TypeError(msg) 624 if self.echo is not None: 625 try: TypeError: write() argument must be str, not <class 'bytes'>
尝试对消息对象调用decode()时,又触发:
AttributeError: 'dict' object has no attribute 'decode'
原因分析
- 文件打开模式不匹配:使用
'rb'二进制模式打开JSON文本文件,虽然json.load能自动解码部分内容,但Facebook导出的JSON中可能存在非标准编码字段,导致部分值被解析为bytes类型而非str。 - 消息结构存在差异:聊天记录中并非所有消息都有
content字段(比如系统通知类消息),这些消息的其他字段可能携带bytes类型数据。单独访问content时避开了这些字段,但打印整个消息对象时,IPython输出流会处理所有字段,遇到bytes就触发类型错误。 - IPython输出机制特性:IPython的输出流对数据类型校验更严格,不像标准Python环境会自动将
bytes转为字符串表示,因此直接打印包含bytes的字典会报错。
解决方案
方案1:用文本模式打开文件并指定编码
将文件打开模式改为'rt',显式指定utf-8编码,确保所有JSON字符串被正确解码为Python的str类型:
import json jsonMessages = [] # 文本模式打开,指定utf-8编码 with open("Facebook/message_1.json", 'rt', encoding='utf-8') as file1: json1 = json.load(file1) jsonMessages.extend(json1['messages']) for msg in jsonMessages: print(msg)
方案2:递归转换消息对象中的bytes类型
如果文件存在特殊编码情况,可遍历消息对象,将所有bytes类型的值转换为str:
import json def convert_bytes_to_str(obj): if isinstance(obj, bytes): return obj.decode('utf-8', errors='replace') # 替换无法解码的字符 elif isinstance(obj, dict): return {k: convert_bytes_to_str(v) for k, v in obj.items()} elif isinstance(obj, list): return [convert_bytes_to_str(item) for item in obj] else: return obj jsonMessages = [] with open("Facebook/message_1.json", 'rb') as file1: json1 = json.load(file1) # 转换所有bytes为str jsonMessages.extend(convert_bytes_to_str(json1['messages'])) for msg in jsonMessages: print(msg)
方案3:仅打印需要的字段
如果只关注消息文本内容,可直接打印content字段(处理无content的系统消息):
import json jsonMessages = [] with open("Facebook/message_1.json", 'rb') as file1: json1 = json.load(file1) jsonMessages.extend(json1['messages']) for msg in jsonMessages: # 存在content则打印,否则输出提示 print(msg.get('content', '【无文本内容的系统消息】'))
内容的提问来源于stack exchange,提问作者jared
相关产品推荐
相关产品推荐

