You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从Azure Blob获取的bytearray如何转为JSON可序列化字符串?

解决Azure Blob JSON数据的bytearray与字符串转换问题

核心问题解析

你当前代码里,blob_client.download_blob().readall()返回的本身就是bytes类型,直接用json.loads()就能解析,没必要先转bytearray。如果确实需要用bytearray处理数据,后续转回字符串只需要通过编码解码操作即可。

具体解决方案

  1. 读取Blob数据并解析JSON(推荐方式)
    跳过不必要的bytearray转换,直接读取解析更高效:

    # 直接读取Blob内容为bytes
    raw_bytes = blob_client.download_blob().readall()
    # json.loads支持直接解析utf-8编码的bytes
    process = json.loads(raw_bytes)
    
  2. bytearray转回字符串(兼容现有逻辑)
    如果已经将数据转为bytearray,调用decode()方法即可转回字符串(JSON文件默认用utf-8编码):

    data = bytearray(blob_client.download_blob().readall())
    # bytearray转JSON字符串
    json_str = data.decode('utf-8')
    process = json.loads(json_str)
    
  3. 处理后的数据转回JSON字符串(序列化)
    完成批量数据处理后,使用json.dumps()将数据序列化为JSON字符串:

    for batch in [process[i:i+batch_size] for i in range(0, len(process), batch_size)]:        
        # 示例处理逻辑:给每个数据项添加处理标记
        for item in batch:
            item["processed"] = True
        
        # 将处理后的batch转为JSON字符串
        serialized_batch = json.dumps(batch, ensure_ascii=False)
        # 若需要再转成bytes/bytearray,可调用encode
        # serialized_bytes = serialized_batch.encode('utf-8')
        # serialized_bytearray = bytearray(serialized_bytes)
    

完整修改后的代码

import json
from azure.storage.blob import BlobServiceClient
from azure.identity import DefaultAzureCredential

account = "myAccount"
container = "myContainer"
blob_name = "myBlob.json"

default_credential = DefaultAzureCredential()
blob_service_client = BlobServiceClient(account, credential=default_credential)
container_client = blob_service_client.get_container_client(container)
blob_client = container_client.get_blob_client(blob_name)

# 直接读取解析JSON
raw_bytes = blob_client.download_blob().readall()
process = json.loads(raw_bytes)

batch_size = 1000

for batch in [process[i:i+batch_size] for i in range(0, len(process), batch_size)]:        
    # 自定义数据处理逻辑
    for item in batch:
        item["status"] = "processed"
    
    # 序列化处理后的batch为JSON字符串
    batch_json_str = json.dumps(batch, ensure_ascii=False)
    # 后续可根据需求使用该字符串(如写入Blob、调用API等)
    print(f"已完成batch序列化,字符串长度:{len(batch_json_str)}")

内容的提问来源于stack exchange,提问作者paone

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 13:50:23