从Azure Blob获取的bytearray如何转为JSON可序列化字符串?
解决Azure Blob JSON数据的bytearray与字符串转换问题
核心问题解析
你当前代码里,blob_client.download_blob().readall()返回的本身就是bytes类型,直接用json.loads()就能解析,没必要先转bytearray。如果确实需要用bytearray处理数据,后续转回字符串只需要通过编码解码操作即可。
具体解决方案
读取Blob数据并解析JSON(推荐方式)
跳过不必要的bytearray转换,直接读取解析更高效:# 直接读取Blob内容为bytes raw_bytes = blob_client.download_blob().readall() # json.loads支持直接解析utf-8编码的bytes process = json.loads(raw_bytes)bytearray转回字符串(兼容现有逻辑)
如果已经将数据转为bytearray,调用decode()方法即可转回字符串(JSON文件默认用utf-8编码):data = bytearray(blob_client.download_blob().readall()) # bytearray转JSON字符串 json_str = data.decode('utf-8') process = json.loads(json_str)处理后的数据转回JSON字符串(序列化)
完成批量数据处理后,使用json.dumps()将数据序列化为JSON字符串:for batch in [process[i:i+batch_size] for i in range(0, len(process), batch_size)]: # 示例处理逻辑:给每个数据项添加处理标记 for item in batch: item["processed"] = True # 将处理后的batch转为JSON字符串 serialized_batch = json.dumps(batch, ensure_ascii=False) # 若需要再转成bytes/bytearray,可调用encode # serialized_bytes = serialized_batch.encode('utf-8') # serialized_bytearray = bytearray(serialized_bytes)
完整修改后的代码
import json from azure.storage.blob import BlobServiceClient from azure.identity import DefaultAzureCredential account = "myAccount" container = "myContainer" blob_name = "myBlob.json" default_credential = DefaultAzureCredential() blob_service_client = BlobServiceClient(account, credential=default_credential) container_client = blob_service_client.get_container_client(container) blob_client = container_client.get_blob_client(blob_name) # 直接读取解析JSON raw_bytes = blob_client.download_blob().readall() process = json.loads(raw_bytes) batch_size = 1000 for batch in [process[i:i+batch_size] for i in range(0, len(process), batch_size)]: # 自定义数据处理逻辑 for item in batch: item["status"] = "processed" # 序列化处理后的batch为JSON字符串 batch_json_str = json.dumps(batch, ensure_ascii=False) # 后续可根据需求使用该字符串(如写入Blob、调用API等) print(f"已完成batch序列化,字符串长度:{len(batch_json_str)}")
内容的提问来源于stack exchange,提问作者paone
相关产品推荐
相关产品推荐

