You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python上传200MB文件到ADLS时出现超时错误如何解决?

解决ADLS大文件上传超时问题(200MB文件)

你的问题根源在于一次性将整个200MB文件读入内存并使用upload_data单请求上传——这种方式不仅会占用大量内存,还会因为单次请求数据量过大、传输时间过长触发连接超时。针对大文件,Azure的ADLS Python SDK提供了更合适的分块上传方案,以下是具体解决方法:

核心解决方案:使用upload_file方法替代upload_data

upload_file方法会自动将大文件分割为多个小块(默认块大小为4MB)分批次上传,每个块独立传输,大幅降低超时风险,同时不需要手动处理分块逻辑。

修改后的代码示例

def upload_large_file_to_adls():
    try:
        file_system_client = service_client.get_file_system_client(file_system="system")
        directory_client = file_system_client.get_directory_client("my-directory")
        file_client = directory_client.get_file_client("uploaded-file.txt")

        # 用二进制模式打开文件(兼容文本/二进制文件,避免编码问题且读取效率更高)
        with open("C:\\file-to-upload.txt", 'rb') as local_file:
            # upload_file自动分块上传,可设置超时时间(单位:秒)
            file_client.upload_file(
                local_file,
                overwrite=True,
                timeout=300,  # 设置5分钟超时,根据网络情况调整
                chunk_size=4*1024*1024  # 可选:自定义分块大小,默认4MB
            )
            
        print("大文件上传完成")
    except Exception as e:
        print(f"上传失败: {str(e)}")

额外优化建议

  • 避免一次性读取文件:不要用read()把整个文件读进内存,用文件对象直接传给upload_file,SDK会自行处理分块读取。
  • 调整超时参数:如果网络环境较差,可以适当增大timeout值(比如设置为600秒即10分钟),或者在初始化service_client时配置全局超时:
    from azure.storage.filedatalake import DataLakeServiceClient
    
    # 初始化客户端时设置全局超时
    service_client = DataLakeServiceClient(
        account_url="https://<account-name>.dfs.core.windows.net",
        credential=<your-credential>,
        timeout=300
    )
    
  • 增加重试逻辑:如果网络波动频繁,可以结合tenacity库实现自动重试,避免单次网络抖动导致上传失败。

内容的提问来源于stack exchange,提问作者Gary

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 13:33:26