You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python从Azure Blob存储下载指定文件的问题排查

问题与解决方案

问题描述

需要从Azure存储容器cont中,下载special/mmm和special/ppp子文件夹下、特定时间段内的zip格式文件(示例文件名:ppp 2023-01-09 11:00:00.zip),但当前代码仅从ddd文件夹下载了一个文件,现有代码如下:

import os
from datetime import datetime
from dateutil import tz
from azure.storage.blob import BlobServiceClient

start_time_str = input("Enter start time (format: yyyy-mm-dd hh:mm:ss): ")
end_time_str = input("Enter end time (format: yyyy-mm-dd hh:mm:ss): ")

start_time = datetime.fromisoformat(start_time_str).replace(tzinfo=tz.tzlocal())
end_time = datetime.fromisoformat(end_time_str).replace(tzinfo=tz.tzlocal())

print("Start time:", start_time)
print("End time:", end_time)

# Create a BlobServiceClient object
connection_string = <"connection string">
blob_service_client = BlobServiceClient.from_connection_string(connection_string)
container_name = "sp"
container_client = blob_service_client.get_container_client(container_name)

local_path = "C:/Users/aaa/Downloads/"
for blob in container_client.list_blobs():
    blob_client = container_client.get_blob_client(blob)
    blob_props = blob_client.get_blob_properties()
    last_modified = blob_props.last_modified.astimezone(tz.tzlocal())
    if last_modified >= start_time and last_modified <= end_time:
        download_path = os.path.join(local_path, blob.name.split("/")[-1])
        with open(download_path, "wb") as download_file:
            download_stream = blob_client.download_blob()
            download_file.write(download_stream.readall())
        print(f"Downloaded blob: {blob.name}")

问题分析

现有代码存在以下关键问题:

  • 容器名称错误:代码中使用"sp",但实际目标容器是"cont"
  • 未限定目标路径:遍历容器内所有blob,未过滤special/mmm和special/ppp子文件夹
  • 未过滤文件格式:没有限制仅下载.zip后缀的文件
  • 未处理同名文件:直接取文件名可能导致不同文件夹下的同名文件被覆盖

修正后的代码

import os
from datetime import datetime
from dateutil import tz
from azure.storage.blob import BlobServiceClient

start_time_str = input("Enter start time (format: yyyy-mm-dd hh:mm:ss): ")
end_time_str = input("Enter end time (format: yyyy-mm-dd hh:mm:ss): ")

start_time = datetime.fromisoformat(start_time_str).replace(tzinfo=tz.tzlocal())
end_time = datetime.fromisoformat(end_time_str).replace(tzinfo=tz.tzlocal())

print("Start time:", start_time)
print("End time:", end_time)

# 修正:使用正确的容器名称
connection_string = "<your connection string>"
blob_service_client = BlobServiceClient.from_connection_string(connection_string)
container_name = "cont"
container_client = blob_service_client.get_container_client(container_name)

local_path = "C:/Users/aaa/Downloads/"
# 定义需要下载的目标子文件夹路径
target_folders = ["special/mmm/", "special/ppp/"]

for blob in container_client.list_blobs():
    blob_path = blob.name
    # 过滤目标文件夹和zip格式文件
    is_target_folder = any(blob_path.startswith(folder) for folder in target_folders)
    is_zip_file = blob_path.endswith(".zip")
    
    if not is_target_folder or not is_zip_file:
        continue
    
    blob_client = container_client.get_blob_client(blob)
    blob_props = blob_client.get_blob_properties()
    last_modified = blob_props.last_modified.astimezone(tz.tzlocal())
    
    # 检查时间范围
    if start_time <= last_modified <= end_time:
        # 保留原文件夹结构,避免同名文件覆盖
        local_file_path = os.path.join(local_path, blob_path)
        # 创建本地文件夹(如果不存在)
        os.makedirs(os.path.dirname(local_file_path), exist_ok=True)
        
        with open(local_file_path, "wb") as download_file:
            download_stream = blob_client.download_blob()
            download_file.write(download_stream.readall())
        print(f"Downloaded blob: {blob.name}")

关键修改说明

  • 修正容器名称为"cont"
  • 添加target_folders变量,限定仅处理指定子文件夹下的文件
  • 增加文件格式过滤,仅处理.zip后缀文件
  • 保留原文件夹结构创建本地文件,避免同名文件覆盖
  • 提前过滤不符合条件的blob,减少不必要的属性查询操作

内容的提问来源于stack exchange,提问作者user20777609

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 09:07:41