You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python异步读取Azure Blob存储如何跳过无效路径异常

问题根因

两个逻辑错误导致异常无法正确捕获、结果丢失:

  • asyncio.gather 默认规则:任意一个传入协程抛出未捕获异常时,整个gather会立即抛出该异常,不会返回其他已完成协程的结果,这是加入不存在路径后拿不到任何有效数据的核心原因。
  • 循环版本的gather用法错误:asyncio.gather(*(read_blob(f))) 中read_blob(f)是单个协程对象,并非可迭代的协程列表,使用*解包会导致协程没有被正确传入gather、也没有被await,才会触发coroutine was never awaited警告,自然无法获取结果。
可落地方案

方案1:单任务包装异常捕获(推荐,异常粒度可控)

给每个blob读取逻辑单独封装异常捕获,遇到ResourceNotFoundError直接返回空标记,再把所有封装后的协程交给gather并发执行,最后过滤无效值即可,不会丢失有效路径的结果:

from azure.core.exceptions import ResourceNotFoundError
import asyncio

async def _safe_read(blob_path):
    try:
        return await read_blob(blob_path)
    except ResourceNotFoundError:
        # 不存在的路径直接返回None标记
        return None

# 并发执行所有读任务
raw_results = await asyncio.gather(*(_safe_read(f) for f in test_dirs))
# 过滤无效结果,得到所有存在路径的读取数据
res = [item for item in raw_results if item is not None]

该方式不会误吞非预期异常(比如网络错误、权限错误、文件格式错误),问题排查成本更低。

方案2:使用gather的return_exceptions参数

asyncio.gather原生支持return_exceptions=True参数,开启后协程抛出的异常会作为对应位置的结果返回,不会中断整体执行,执行完成后手动过滤目标异常即可:

from azure.core.exceptions import ResourceNotFoundError
import asyncio

raw_results = await asyncio.gather(
    *(read_blob(f) for f in test_dirs),
    return_exceptions=True
)
res = []
for item in raw_results:
    # 跳过不存在的blob异常
    if isinstance(item, ResourceNotFoundError):
        continue
    # 非预期异常正常抛出,避免静默失败
    if isinstance(item, Exception):
        raise item
    res.append(item)

补充:循环写法的正确版本(不推荐,无并发性能)

如果坚持用循环逐次处理,不需要嵌套gather,直接await单个协程即可,但该方式是串行执行请求,完全丢失异步并发的性能优势,仅做语法参考:

from azure.core.exceptions import ResourceNotFoundError

res = []
for f in test_dirs:
    try:
        data = await read_blob(f)
        res.append(data)
    except ResourceNotFoundError:
        continue

注意不要使用裸except:,会吞掉所有异常包括系统级中断信号,导致程序无法正常退出、问题无法排查。


内容的提问来源于stack exchange,提问作者user9532692

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 12:27:44