使用Python从Azure Blob存储下载指定文件的问题排查
问题与解决方案
问题描述
需要从Azure存储容器cont中,下载special/mmm和special/ppp子文件夹下、特定时间段内的zip格式文件(示例文件名:ppp 2023-01-09 11:00:00.zip),但当前代码仅从ddd文件夹下载了一个文件,现有代码如下:
import os from datetime import datetime from dateutil import tz from azure.storage.blob import BlobServiceClient start_time_str = input("Enter start time (format: yyyy-mm-dd hh:mm:ss): ") end_time_str = input("Enter end time (format: yyyy-mm-dd hh:mm:ss): ") start_time = datetime.fromisoformat(start_time_str).replace(tzinfo=tz.tzlocal()) end_time = datetime.fromisoformat(end_time_str).replace(tzinfo=tz.tzlocal()) print("Start time:", start_time) print("End time:", end_time) # Create a BlobServiceClient object connection_string = <"connection string"> blob_service_client = BlobServiceClient.from_connection_string(connection_string) container_name = "sp" container_client = blob_service_client.get_container_client(container_name) local_path = "C:/Users/aaa/Downloads/" for blob in container_client.list_blobs(): blob_client = container_client.get_blob_client(blob) blob_props = blob_client.get_blob_properties() last_modified = blob_props.last_modified.astimezone(tz.tzlocal()) if last_modified >= start_time and last_modified <= end_time: download_path = os.path.join(local_path, blob.name.split("/")[-1]) with open(download_path, "wb") as download_file: download_stream = blob_client.download_blob() download_file.write(download_stream.readall()) print(f"Downloaded blob: {blob.name}")
问题分析
现有代码存在以下关键问题:
- 容器名称错误:代码中使用
"sp",但实际目标容器是"cont" - 未限定目标路径:遍历容器内所有blob,未过滤
special/mmm和special/ppp子文件夹 - 未过滤文件格式:没有限制仅下载
.zip后缀的文件 - 未处理同名文件:直接取文件名可能导致不同文件夹下的同名文件被覆盖
修正后的代码
import os from datetime import datetime from dateutil import tz from azure.storage.blob import BlobServiceClient start_time_str = input("Enter start time (format: yyyy-mm-dd hh:mm:ss): ") end_time_str = input("Enter end time (format: yyyy-mm-dd hh:mm:ss): ") start_time = datetime.fromisoformat(start_time_str).replace(tzinfo=tz.tzlocal()) end_time = datetime.fromisoformat(end_time_str).replace(tzinfo=tz.tzlocal()) print("Start time:", start_time) print("End time:", end_time) # 修正:使用正确的容器名称 connection_string = "<your connection string>" blob_service_client = BlobServiceClient.from_connection_string(connection_string) container_name = "cont" container_client = blob_service_client.get_container_client(container_name) local_path = "C:/Users/aaa/Downloads/" # 定义需要下载的目标子文件夹路径 target_folders = ["special/mmm/", "special/ppp/"] for blob in container_client.list_blobs(): blob_path = blob.name # 过滤目标文件夹和zip格式文件 is_target_folder = any(blob_path.startswith(folder) for folder in target_folders) is_zip_file = blob_path.endswith(".zip") if not is_target_folder or not is_zip_file: continue blob_client = container_client.get_blob_client(blob) blob_props = blob_client.get_blob_properties() last_modified = blob_props.last_modified.astimezone(tz.tzlocal()) # 检查时间范围 if start_time <= last_modified <= end_time: # 保留原文件夹结构,避免同名文件覆盖 local_file_path = os.path.join(local_path, blob_path) # 创建本地文件夹(如果不存在) os.makedirs(os.path.dirname(local_file_path), exist_ok=True) with open(local_file_path, "wb") as download_file: download_stream = blob_client.download_blob() download_file.write(download_stream.readall()) print(f"Downloaded blob: {blob.name}")
关键修改说明
- 修正容器名称为
"cont" - 添加
target_folders变量,限定仅处理指定子文件夹下的文件 - 增加文件格式过滤,仅处理
.zip后缀文件 - 保留原文件夹结构创建本地文件,避免同名文件覆盖
- 提前过滤不符合条件的blob,减少不必要的属性查询操作
内容的提问来源于stack exchange,提问作者user20777609
相关产品推荐
相关产品推荐

