如何在Python中从HBase记录获取JSON?解决bytes序列化问题
解决包含bytes类型的Python字典JSON序列化问题
针对你遇到的bytes类型无法JSON序列化的问题,不需要手动遍历每个字段,有两种简便的解决方式:
方法1:自定义JSON编码器类
通过继承json.JSONEncoder扩展序列化逻辑,让JSON自动处理bytes类型:
import json class BytesJSONEncoder(json.JSONEncoder): def default(self, obj): if isinstance(obj, bytes): # 按实际编码调整,比如utf-8,可添加错误处理避免解码失败 return obj.decode('utf-8', errors='replace') # 交给父类处理其他默认支持的类型 return super().default(obj)
使用方式:
# 生成格式化JSON字符串 json_result = json.dumps(row, indent=4, cls=BytesJSONEncoder) print(json_result)
方法2:直接使用default参数快速实现
如果不想定义类,可直接在json.dumps中传入lambda函数处理bytes:
import json json_result = json.dumps( row, indent=4, default=lambda x: x.decode('utf-8', errors='replace') if isinstance(x, bytes) else x ) print(json_result)
额外:优雅打印处理后的字典
如果只是需要可读性高的打印输出,不需要生成JSON,可递归转换所有bytes为字符串后用pprint:
import pprint def convert_all_bytes(obj): if isinstance(obj, bytes): return obj.decode('utf-8', errors='replace') elif isinstance(obj, dict): return {k: convert_all_bytes(v) for k, v in obj.items()} elif isinstance(obj, list): return [convert_all_bytes(item) for item in obj] return obj processed_row = convert_all_bytes(row) pprint.pprint(processed_row, indent=4)
注意:如果你的bytes数据不是UTF-8编码,需要把decode里的编码参数改成对应格式(比如gbk),避免乱码或解码错误。
内容的提问来源于stack exchange,提问作者peter.petrov
相关产品推荐
相关产品推荐

