为何存在的Azure Blob无法被追踪识别?部分容器Blob异常
问题描述
我的代码能正常处理容器A及其内的Blob A-1,但无法识别其他容器及其中的Blob。代码功能是从JSON文件提取指定列并转换为CSV保存到本地磁盘。所有无法识别的文件均为JSON格式,大小不超过16MB,不清楚问题出在哪里。
原代码
import csv from azure.storage.blob import BlobServiceClient import json # Azure Blob storage connection string connection_string = 'my_connection_string' # Blob container and file information container_name = 'my_container_name' blob_name = 'My_jsonfile.json' # Create BlobServiceClient instance blob_service_client = BlobServiceClient.from_connection_string(connection_string) # Get BlobClient for the JSON file blob_client = blob_service_client.get_blob_client(container=container_name, blob=blob_name) # Download and load the JSON data blob_data = blob_client.download_blob().content_as_text() json_data = json.loads(blob_data) # Define the local CSV file path csv_file_path = 'local_path' # Define the fieldnames for the CSV file fieldnames = ['id', 'type', 'actor/login', 'payload/issue/labels/0/name', 'payload/issue/body', 'payload/issue/title',] # Open the local CSV file in write mode with open(csv_file_path, 'w', newline='', encoding='utf-8') as csv_file: writer = csv.DictWriter(csv_file, fieldnames=fieldnames) writer.writeheader() # Write the data rows for event in json_data: try: id = event['id'] type = event['type'] actor = event['actor']['login'] name = event['payload']['issue']['labels'][0]['name'] body = event['payload']['issue']['body'] title = event['payload']['issue']['title'] writer.writerow({'id': id, 'type': type, 'actor/login': actor, 'payload/issue/labels/0/name': name, 'payload/issue/body': body, 'payload/issue/title': title}) except (KeyError, IndexError): # Handle missing keys or empty lists gracefully pass print(' ') print() # Add an empty line for separation print('CSV file created successfully! File path:', csv_file_path)
问题分析与解决方案
- 硬编码的容器/Blob名称:原代码中
container_name和blob_name是固定值,只能处理指定的单个容器和Blob。要处理所有容器及Blob,需要遍历资源:- 使用
blob_service_client.list_containers()获取所有容器 - 对每个容器,使用
container_client.list_blobs()获取容器内的所有Blob
- 使用
- 权限不足:确认连接字符串对应的账号拥有容器列表权限(列出所有容器)和Blob读取权限。如果用SAS令牌,需确保包含
list和read权限;如果用账号密钥,需确保账号被分配了Storage Blob Data Contributor或类似角色。 - 容器/Blob名称问题:检查其他容器/Blob名称是否存在大小写敏感、特殊字符(如空格、斜杠)的情况,确保代码中获取BlobClient时名称完全匹配。
- JSON结构差异:其他Blob的JSON结构可能与A-1不同,导致代码触发
KeyError或IndexError后被pass,看起来像未识别。建议在异常处理中添加日志,记录跳过的事件或Blob信息,排查数据结构差异。
修改后的示例代码(支持遍历所有容器和Blob)
import csv from azure.storage.blob import BlobServiceClient import json import os # Azure Blob storage connection string connection_string = 'my_connection_string' # 本地保存CSV的根目录 local_root_path = './csv_output' os.makedirs(local_root_path, exist_ok=True) # 定义CSV字段名 fieldnames = ['id', 'type', 'actor/login', 'payload/issue/labels/0/name', 'payload/issue/body', 'payload/issue/title'] # 创建BlobServiceClient实例 blob_service_client = BlobServiceClient.from_connection_string(connection_string) # 遍历所有容器 for container in blob_service_client.list_containers(): container_name = container.name print(f"处理容器: {container_name}") container_client = blob_service_client.get_container_client(container_name) # 遍历容器内的所有Blob(仅处理JSON文件) for blob in container_client.list_blobs(): blob_name = blob.name if not blob_name.endswith('.json'): continue print(f" 处理Blob: {blob_name}") # 获取BlobClient并下载数据 blob_client = container_client.get_blob_client(blob_name) try: blob_data = blob_client.download_blob().content_as_text() json_data = json.loads(blob_data) except Exception as e: print(f" 下载/解析Blob失败: {str(e)}") continue # 生成本地CSV路径(按容器名创建子目录) container_dir = os.path.join(local_root_path, container_name) os.makedirs(container_dir, exist_ok=True) csv_file_name = os.path.splitext(blob_name)[0] + '.csv' csv_file_path = os.path.join(container_dir, csv_file_name) # 写入CSV文件 with open(csv_file_path, 'w', newline='', encoding='utf-8') as csv_file: writer = csv.DictWriter(csv_file, fieldnames=fieldnames) writer.writeheader() skipped_count = 0 for event in json_data: try: row_data = { 'id': event['id'], 'type': event['type'], 'actor/login': event['actor']['login'], 'payload/issue/labels/0/name': event['payload']['issue']['labels'][0]['name'], 'payload/issue/body': event['payload']['issue']['body'], 'payload/issue/title': event['payload']['issue']['title'] } writer.writerow(row_data) except (KeyError, IndexError) as e: skipped_count += 1 # 记录跳过的事件ID(如果有),方便排查 event_id = event.get('id', '未知ID') print(f" 跳过事件 {event_id}: 缺失字段或索引错误 - {str(e)}") print(f" CSV文件已生成: {csv_file_path},共跳过 {skipped_count} 条事件") print("所有容器和Blob处理完成!")
内容的提问来源于stack exchange,提问作者user21214592
相关产品推荐
相关产品推荐

