You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何存在的Azure Blob无法被追踪识别?部分容器Blob异常

问题描述

我的代码能正常处理容器A及其内的Blob A-1,但无法识别其他容器及其中的Blob。代码功能是从JSON文件提取指定列并转换为CSV保存到本地磁盘。所有无法识别的文件均为JSON格式,大小不超过16MB,不清楚问题出在哪里。

原代码
import csv
from azure.storage.blob import BlobServiceClient
import json

# Azure Blob storage connection string
    connection_string = 'my_connection_string'

    # Blob container and file information
    container_name = 'my_container_name'
    blob_name = 'My_jsonfile.json'

    # Create BlobServiceClient instance
    blob_service_client = BlobServiceClient.from_connection_string(connection_string)

    # Get BlobClient for the JSON file
    blob_client = blob_service_client.get_blob_client(container=container_name, blob=blob_name)

    # Download and load the JSON data
    blob_data = blob_client.download_blob().content_as_text()
    json_data = json.loads(blob_data)

    # Define the local CSV file path
    csv_file_path = 'local_path'

    # Define the fieldnames for the CSV file
    fieldnames = ['id', 'type', 'actor/login', 'payload/issue/labels/0/name', 'payload/issue/body', 'payload/issue/title',]
    # Open the local CSV file in write mode
    with open(csv_file_path, 'w', newline='', encoding='utf-8') as csv_file:
    writer = csv.DictWriter(csv_file, fieldnames=fieldnames)

    writer.writeheader()

    
    # Write the data rows
    for event in json_data:
        try:
            id = event['id']
            type = event['type']
            actor = event['actor']['login']
            name = event['payload']['issue']['labels'][0]['name']
            body = event['payload']['issue']['body']
            title = event['payload']['issue']['title']

            writer.writerow({'id': id, 'type': type, 'actor/login': actor,     'payload/issue/labels/0/name': name, 'payload/issue/body': body, 'payload/issue/title': title})
        except (KeyError, IndexError):
            # Handle missing keys or empty lists gracefully
            pass

       
            print(' ')
            print()  # Add an empty line for separation

print('CSV file created successfully! File path:', csv_file_path)
问题分析与解决方案
  • 硬编码的容器/Blob名称:原代码中container_name和blob_name是固定值,只能处理指定的单个容器和Blob。要处理所有容器及Blob,需要遍历资源:
    1. 使用blob_service_client.list_containers()获取所有容器
    2. 对每个容器,使用container_client.list_blobs()获取容器内的所有Blob
  • 权限不足:确认连接字符串对应的账号拥有容器列表权限(列出所有容器)和Blob读取权限。如果用SAS令牌,需确保包含list和read权限;如果用账号密钥,需确保账号被分配了Storage Blob Data Contributor或类似角色。
  • 容器/Blob名称问题:检查其他容器/Blob名称是否存在大小写敏感、特殊字符(如空格、斜杠)的情况,确保代码中获取BlobClient时名称完全匹配。
  • JSON结构差异:其他Blob的JSON结构可能与A-1不同,导致代码触发KeyError或IndexError后被pass,看起来像未识别。建议在异常处理中添加日志,记录跳过的事件或Blob信息,排查数据结构差异。
修改后的示例代码(支持遍历所有容器和Blob)
import csv
from azure.storage.blob import BlobServiceClient
import json
import os

# Azure Blob storage connection string
connection_string = 'my_connection_string'
# 本地保存CSV的根目录
local_root_path = './csv_output'
os.makedirs(local_root_path, exist_ok=True)

# 定义CSV字段名
fieldnames = ['id', 'type', 'actor/login', 'payload/issue/labels/0/name', 'payload/issue/body', 'payload/issue/title']

# 创建BlobServiceClient实例
blob_service_client = BlobServiceClient.from_connection_string(connection_string)

# 遍历所有容器
for container in blob_service_client.list_containers():
    container_name = container.name
    print(f"处理容器: {container_name}")
    container_client = blob_service_client.get_container_client(container_name)
    
    # 遍历容器内的所有Blob(仅处理JSON文件)
    for blob in container_client.list_blobs():
        blob_name = blob.name
        if not blob_name.endswith('.json'):
            continue
        print(f"  处理Blob: {blob_name}")
        
        # 获取BlobClient并下载数据
        blob_client = container_client.get_blob_client(blob_name)
        try:
            blob_data = blob_client.download_blob().content_as_text()
            json_data = json.loads(blob_data)
        except Exception as e:
            print(f"    下载/解析Blob失败: {str(e)}")
            continue
        
        # 生成本地CSV路径(按容器名创建子目录)
        container_dir = os.path.join(local_root_path, container_name)
        os.makedirs(container_dir, exist_ok=True)
        csv_file_name = os.path.splitext(blob_name)[0] + '.csv'
        csv_file_path = os.path.join(container_dir, csv_file_name)
        
        # 写入CSV文件
        with open(csv_file_path, 'w', newline='', encoding='utf-8') as csv_file:
            writer = csv.DictWriter(csv_file, fieldnames=fieldnames)
            writer.writeheader()
            
            skipped_count = 0
            for event in json_data:
                try:
                    row_data = {
                        'id': event['id'],
                        'type': event['type'],
                        'actor/login': event['actor']['login'],
                        'payload/issue/labels/0/name': event['payload']['issue']['labels'][0]['name'],
                        'payload/issue/body': event['payload']['issue']['body'],
                        'payload/issue/title': event['payload']['issue']['title']
                    }
                    writer.writerow(row_data)
                except (KeyError, IndexError) as e:
                    skipped_count += 1
                    # 记录跳过的事件ID(如果有),方便排查
                    event_id = event.get('id', '未知ID')
                    print(f"    跳过事件 {event_id}: 缺失字段或索引错误 - {str(e)}")
            
            print(f"    CSV文件已生成: {csv_file_path},共跳过 {skipped_count} 条事件")

print("所有容器和Blob处理完成!")

内容的提问来源于stack exchange,提问作者user21214592

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 01:58:20