序列化含bytes类型属性对象时的JSON序列化错误问题
问题描述
需将Avro记录发布至消息编码类型为JSON的Pub/Sub主题,Avro Schema中qty与cur字段为bytes类型,订阅类型为推送到BigQuery。待处理的Avro数据如下:
{'id': 1830170602, 'qty': b"'\x10", 'cur': b'J\xc4\xa0'}
读取Avro文件数据并写入主题的代码如下:
reader = DataFileReader(open(file_name, "rb"), DatumReader()) for x in reader: # Get the topic encoding type. topic = publisher_client.get_topic(request={"topic": topic_path}) encoding = topic.schema_settings.encoding # Encode the data according to the message serialization type. if encoding == Encoding.JSON: data_str = json.dumps(x) print(f"Preparing a JSON-encoded message:\n{data_str}") data = data_str.encode("utf-8")
执行至json.dumps()时出现错误:
TypeError: Object of type bytes is not JSON serializable
尝试将主题消息编码改为Binary仍未解决,咨询是否需要将字节字段值用Base64编码?
解决方案
- 必须对bytes字段做Base64编码,JSON规范不支持原生序列化bytes类型,这是报错的直接原因。
- 修改代码可通过两种方式实现:
- 手动转换bytes字段为Base64字符串:
import base64 import json reader = DataFileReader(open(file_name, "rb"), DatumReader()) for x in reader: topic = publisher_client.get_topic(request={"topic": topic_path}) encoding = topic.schema_settings.encoding if encoding == Encoding.JSON: # 转换bytes字段为Base64字符串 x['qty'] = base64.b64encode(x['qty']).decode('utf-8') x['cur'] = base64.b64encode(x['cur']).decode('utf-8') data_str = json.dumps(x) print(f"Preparing a JSON-encoded message:\n{data_str}") data = data_str.encode("utf-8") - 自定义JSON编码器(适合字段较多的场景):
import base64 import json class BytesEncoder(json.JSONEncoder): def default(self, obj): if isinstance(obj, bytes): return base64.b64encode(obj).decode('utf-8') return super().default(obj) reader = DataFileReader(open(file_name, "rb"), DatumReader()) for x in reader: topic = publisher_client.get_topic(request={"topic": topic_path}) encoding = topic.schema_settings.encoding if encoding == Encoding.JSON: data_str = json.dumps(x, cls=BytesEncoder) print(f"Preparing a JSON-encoded message:\n{data_str}") data = data_str.encode("utf-8")
- 手动转换bytes字段为Base64字符串:
- 关于改为Binary编码未解决的问题:Binary编码下Pub/Sub传递的是原始Avro二进制数据,但订阅推送到BigQuery时,需要确保订阅配置了与Avro Schema匹配的解析规则,否则BigQuery无法正确识别
qty和cur的bytes类型。而你的核心需求是使用JSON编码主题,所以用Base64处理bytes字段适配JSON的方案更直接。 - BigQuery侧注意事项:如果BigQuery表中
qty和cur字段类型是STRING,Base64编码后的字符串可直接写入;如果需要还原为二进制数据,可在BigQuery中使用FROM_BASE64()函数转换,例如:SELECT FROM_BASE64(qty) AS qty_bytes FROM your_table。
内容的提问来源于stack exchange,提问作者mehere
相关产品推荐
相关产品推荐

