You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python上传至Azure Blob Storage速度过慢问题咨询

上传速度分析与优化建议

速度是否正常?

完全不正常。先计算带宽对应的理论传输能力:
你的本地带宽是110Mbps,转换成实际文件传输的MB/s(注意单位换算:1Byte=8bit):
110Mbps ÷ 8 = 13.75MB/s
4MB的分块理论上仅需约0.29秒就能完成上传,但实际耗时8秒,传输速度仅0.5MB/s,仅为理论带宽的3.6%;50MB文件理论传输耗时不到4秒,实际却用了2分钟,显然存在严重的性能瓶颈。

问题根源(从代码分析)

你的上传逻辑存在几个关键性能缺陷:

  • 串行分块上传:代码循环读取一个块、上传一个块,完全串行执行。Azure Blob Storage支持并行上传多个分块,串行模式彻底浪费了带宽资源。
  • 重复创建服务客户端:generate_blob_client方法每次都重新初始化BlobServiceClient,这个对象属于重量级资源(包含连接池等),重复创建会带来额外的初始化开销。
  • 未优化文件读取:直接从request.files['file'](Werkzeug的FileStorage对象)逐块读取,相比先写入本地临时文件再读取,可能存在IO效率问题。

优化方案

1. 改用并行分块上传

Azure Storage Blob SDK内置了并行分块上传逻辑,推荐直接使用,代码更简洁且性能更高:

def upload_chunks(self, blob_client: BlobClient, file):
    # 利用SDK内置并行能力上传
    blob_client.upload_blob(
        data=file,
        chunk_size=self.chunk_size,
        max_concurrency=8,  # 并行上传的块数,可根据带宽调整(建议4-16)
        overwrite=True
    )
    return blob_client.url

如果需要手动实现并行,可使用concurrent.futures.ThreadPoolExecutor批量调用stage_block,但SDK内置实现已做过优化,优先选择内置方法。

2. 复用BlobServiceClient

将BlobServiceClient的初始化放到类的构造函数中,避免每次创建BlobClient都重复初始化:

def __init__(self):
    self.blob_service_client = BlobServiceClient.from_connection_string(self.connection_string)
    self.container_client = self.blob_service_client.get_container_client(self.container_name)

def generate_blob_client(self, file_name: str):
    for _ in range(self.max_blob_name_tries):
        blob_name = self.generate_blob_name(file_name)
        blob_client = self.container_client.get_blob_client(blob_name)
        if not blob_client.exists():
            return blob_client
    raise Exception("无法创建唯一Blob名称")

3. 优化文件读取(可选)

上传大文件时,先将FileStorage写入本地临时文件,再从文件读取上传,能提升IO效率:

import tempfile

def upload_to_blob(self, file):
    file_name = file.filename
    file_type = file.content_type

    # 写入临时文件
    with tempfile.NamedTemporaryFile(delete=False) as tmp_file:
        tmp_file.write(file.read())
        tmp_file_path = tmp_file.name

    blob_client = self.generate_blob_client(file_name)
    try:
        with open(tmp_file_path, "rb") as f:
            blob_client.upload_blob(
                data=f,
                chunk_size=self.chunk_size,
                max_concurrency=8,
                overwrite=True
            )
        blob_url = blob_client.url
    finally:
        os.unlink(tmp_file_path)  # 上传完成后删除临时文件
    return blob_url, file_name, file_type

额外排查方向

  • 检查Azure Blob Storage的区域是否靠近你的服务器/用户,跨区域传输会带来明显延迟。
  • 开启SDK日志,查看是否存在请求重试、超时等网络层面的异常。

内容的提问来源于stack exchange,提问作者Mikhael Abdallah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 02:10:38