如何通过Python Azure API按模式查询Azure Blobs?
使用Python Azure SDK实现Unix风格Glob模式查询Blob文件
嘿,刚好对这个场景很熟悉!要在Azure Blob存储里用类似Unix glob的路径模式筛选文件(比如找出所有嵌套目录下reports文件夹里的PDF报告),用Python结合Azure SDK就能轻松实现,下面给你分两种常用SDK版本的方案:
方案一:适配旧版Azure Storage Blob SDK(对应你示例的写法)
如果你还在使用基于BlockBlobService的旧版SDK(azure-storage-blob < 12.x),可以这样实现:
首先确保初始化好服务客户端:
from azure.storage.blob import BlockBlobService import glob # 替换成你的Azure存储账户名和密钥 block_blob_service = BlockBlobService(account_name='your_account_name', account_key='your_account_key')
旧版的list_blobs没法直接识别**这类递归通配符,所以我们先列出容器内所有Blob,再用Python的glob.fnmatch工具来筛选符合模式的文件:
container_name = 'mycontainer' # 你的目标匹配模式 target_pattern = '**/reports/*.pdf' # 获取容器内所有Blob对象 all_blobs = block_blob_service.list_blobs(container_name) # 过滤出符合模式的Blob matching_blobs = [blob for blob in all_blobs if glob.fnmatch.fnmatch(blob.name, target_pattern)] # 输出结果示例 for blob in matching_blobs: print(f"找到匹配的报告文件: {blob.name}")
方案二:使用新版Azure Storage Blob SDK(推荐)
新版SDK(12.x+)是官方主推的版本,API设计更直观,功能也更全面,推荐优先使用这个方案:
首先安装新版SDK:
pip install azure-storage-blob
然后编写代码实现:
from azure.storage.blob import BlobServiceClient import glob # 替换成你的Azure存储连接字符串 connection_string = "your_storage_connection_string" blob_service_client = BlobServiceClient.from_connection_string(connection_string) # 获取目标容器的客户端 container_client = blob_service_client.get_container_client("mycontainer") target_pattern = '**/reports/*.pdf' # 遍历并过滤Blob matching_blobs = [] for blob in container_client.list_blobs(): if glob.fnmatch.fnmatch(blob.name, target_pattern): matching_blobs.append(blob) # 处理匹配到的Blob,比如打印路径或下载 for blob in matching_blobs: print(f"匹配的报告文件路径: {blob.name}") # 如果需要下载文件,可以取消下面的注释 # blob_client = container_client.get_blob_client(blob) # with open(f"./local_reports/{blob.name.split('/')[-1]}", "wb") as local_file: # blob_data = blob_client.download_blob() # blob_data.readinto(local_file)
进阶优化:大数据量场景下的分页查询
如果你的容器里有大量Blob,一次性列出所有Blob可能会占用过多内存。新版SDK支持分页查询,我们可以分批获取并过滤:
matching_blobs = [] continuation_token = None while True: # 每次获取1000个Blob(可根据需求调整) blobs_page = container_client.list_blobs(continuation_token=continuation_token, max_results=1000) for blob in blobs_page: if glob.fnmatch.fnmatch(blob.name, target_pattern): matching_blobs.append(blob) # 更新续传令牌,直到没有更多Blob continuation_token = blobs_page.continuation_token if not continuation_token: break
内容的提问来源于stack exchange,提问作者Joost Dübken
相关产品推荐
相关产品推荐

