You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python读取Facebook导出JSON时的字符串与字节类型混淆问题

Facebook聊天记录JSON读取的类型混淆问题分析与解决

问题重现

读取Facebook导出的聊天记录JSON文件时,单独访问列表首个元素的content字段能正常输出,但遍历打印每个消息对象时报错:

import json

jsonMessages = []

with open("Facebook/message_1.json", 'rb') as file1:
    json1 = json.load(file1)
    jsonMessages.extend(json1['messages'])

jsonMessages[0]['content']  # 正常输出消息内容

for msg in jsonMessages:
    print(msg)  # 触发类型错误

报错信息:

---------------------------------------------------------------------------
TypeError                                 Traceback (most recent call last)
Cell In[170], line 10
  7 jsonMessages[0]['content']
  9 for msg in jsonMessages:
---> 10     print(msg)

File <frozen codecs>:378, in write(self, object)

File ~\AppData\Local\anaconda3\Lib\site-packages\ipykernel\iostream.py:622, in OutStream.write(self, string)
620 if not isinstance(string, str):
621     msg = f"write() argument must be str, not {type(string)}"
--> 622     raise TypeError(msg)
624 if self.echo is not None:
625     try:

TypeError: write() argument must be str, not <class 'bytes'>

尝试对消息对象调用decode()时,又触发:

AttributeError: 'dict' object has no attribute 'decode'

原因分析

  1. 文件打开模式不匹配:使用'rb'二进制模式打开JSON文本文件,虽然json.load能自动解码部分内容,但Facebook导出的JSON中可能存在非标准编码字段,导致部分值被解析为bytes类型而非str。
  2. 消息结构存在差异:聊天记录中并非所有消息都有content字段(比如系统通知类消息),这些消息的其他字段可能携带bytes类型数据。单独访问content时避开了这些字段,但打印整个消息对象时,IPython输出流会处理所有字段,遇到bytes就触发类型错误。
  3. IPython输出机制特性:IPython的输出流对数据类型校验更严格,不像标准Python环境会自动将bytes转为字符串表示,因此直接打印包含bytes的字典会报错。

解决方案

方案1:用文本模式打开文件并指定编码

将文件打开模式改为'rt',显式指定utf-8编码,确保所有JSON字符串被正确解码为Python的str类型:

import json

jsonMessages = []

# 文本模式打开,指定utf-8编码
with open("Facebook/message_1.json", 'rt', encoding='utf-8') as file1:
    json1 = json.load(file1)
    jsonMessages.extend(json1['messages'])

for msg in jsonMessages:
    print(msg)

方案2:递归转换消息对象中的bytes类型

如果文件存在特殊编码情况,可遍历消息对象,将所有bytes类型的值转换为str:

import json

def convert_bytes_to_str(obj):
    if isinstance(obj, bytes):
        return obj.decode('utf-8', errors='replace')  # 替换无法解码的字符
    elif isinstance(obj, dict):
        return {k: convert_bytes_to_str(v) for k, v in obj.items()}
    elif isinstance(obj, list):
        return [convert_bytes_to_str(item) for item in obj]
    else:
        return obj

jsonMessages = []

with open("Facebook/message_1.json", 'rb') as file1:
    json1 = json.load(file1)
    # 转换所有bytes为str
    jsonMessages.extend(convert_bytes_to_str(json1['messages']))

for msg in jsonMessages:
    print(msg)

方案3:仅打印需要的字段

如果只关注消息文本内容,可直接打印content字段(处理无content的系统消息):

import json

jsonMessages = []

with open("Facebook/message_1.json", 'rb') as file1:
    json1 = json.load(file1)
    jsonMessages.extend(json1['messages'])

for msg in jsonMessages:
    # 存在content则打印,否则输出提示
    print(msg.get('content', '【无文本内容的系统消息】'))

内容的提问来源于stack exchange,提问作者jared

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 16:33:15