使用Python上传200MB文件到ADLS时出现超时错误如何解决?
解决ADLS大文件上传超时问题(200MB文件)
你的问题根源在于一次性将整个200MB文件读入内存并使用upload_data单请求上传——这种方式不仅会占用大量内存,还会因为单次请求数据量过大、传输时间过长触发连接超时。针对大文件,Azure的ADLS Python SDK提供了更合适的分块上传方案,以下是具体解决方法:
核心解决方案:使用upload_file方法替代upload_data
upload_file方法会自动将大文件分割为多个小块(默认块大小为4MB)分批次上传,每个块独立传输,大幅降低超时风险,同时不需要手动处理分块逻辑。
修改后的代码示例
def upload_large_file_to_adls(): try: file_system_client = service_client.get_file_system_client(file_system="system") directory_client = file_system_client.get_directory_client("my-directory") file_client = directory_client.get_file_client("uploaded-file.txt") # 用二进制模式打开文件(兼容文本/二进制文件,避免编码问题且读取效率更高) with open("C:\\file-to-upload.txt", 'rb') as local_file: # upload_file自动分块上传,可设置超时时间(单位:秒) file_client.upload_file( local_file, overwrite=True, timeout=300, # 设置5分钟超时,根据网络情况调整 chunk_size=4*1024*1024 # 可选:自定义分块大小,默认4MB ) print("大文件上传完成") except Exception as e: print(f"上传失败: {str(e)}")
额外优化建议
- 避免一次性读取文件:不要用
read()把整个文件读进内存,用文件对象直接传给upload_file,SDK会自行处理分块读取。 - 调整超时参数:如果网络环境较差,可以适当增大
timeout值(比如设置为600秒即10分钟),或者在初始化service_client时配置全局超时:from azure.storage.filedatalake import DataLakeServiceClient # 初始化客户端时设置全局超时 service_client = DataLakeServiceClient( account_url="https://<account-name>.dfs.core.windows.net", credential=<your-credential>, timeout=300 ) - 增加重试逻辑:如果网络波动频繁,可以结合
tenacity库实现自动重试,避免单次网络抖动导致上传失败。
内容的提问来源于stack exchange,提问作者Gary
相关产品推荐
相关产品推荐

