使用from_generator流式导入Azure数据存储图像到datasets遇性能问题求解
问题:AzureML中生成器方式流式传输图像到Dataset效率低下
我在AzureML Jupyter Notebook实例中,需要将Azure数据存储中的数千张图像流式传输到datasets对象。使用以下代码时,运行45分钟仍无结束迹象:
from azureml.fsspec import AzureMachineLearningFileSystem from datasets import Dataset fs = AzureMachineLearningFileSystem('azureml://subscriptions/<sub_id>/resourcegroups/<rg_name>/workspaces/<ws_name>/datastores/<ds_name') file_lst = fs.ls('datastore_subdir/')[1:] def process_image(): for file in file_lst: img = Image.open(fs.open(file)) yield {'image': img} image_ds = Dataset.from_generator(process_image)
我已通过其他方法实现需求,但希望了解是否能通过生成器方法实现高效运行。作为对比,以下代码仅需1.5分钟即可完成:
import io img_lst = [] for file in fs.ls('datastore_subdir/')[1:]: with fs.open(file) as f: print(counter) img = Image.open(io.BytesIO(f.read())) img_lst.append(img)
解答
原生成器代码低效的原因
- 未正确管理远程文件流:直接将
fs.open(file)返回的远程流传给Image.open,流不会被主动关闭,会导致远程连接资源泄漏,后续IO操作被阻塞。 - 远程流的低效读取:
Image.open操作远程流时,会分多次读取数据,频繁的远程IO交互大幅拖慢速度。 - 单线程处理:默认情况下
Dataset.from_generator用单线程执行生成器,无法利用多核资源加速。
优化后的生成器实现
将文件读取逻辑和高效代码对齐,同时启用并行处理,即可大幅提升效率:
from azureml.fsspec import AzureMachineLearningFileSystem from datasets import Dataset import io from PIL import Image fs = AzureMachineLearningFileSystem('azureml://subscriptions/<sub_id>/resourcegroups/<rg_name>/workspaces/<ws_name>/datastores/<ds_name') file_lst = fs.ls('datastore_subdir/')[1:] def process_image(): for file in file_lst: # 用with语句确保文件流及时关闭 with fs.open(file) as f: # 一次性读取文件字节到内存,避免远程流的多次IO img_bytes = f.read() img = Image.open(io.BytesIO(img_bytes)) yield {'image': img} # 启用并行处理,根据实例CPU核心数调整num_proc参数 image_ds = Dataset.from_generator(process_image, num_proc=4)
优化说明
- 文件流管理:
with语句会自动关闭文件流,释放远程连接资源,避免IO阻塞。 - 内存中处理图像:先将文件一次性读入内存再用
io.BytesIO包装,和高效代码的逻辑一致,消除远程流的低效读取问题。 - 并行加速:通过
num_proc参数启用多进程并行处理生成器,充分利用AzureML实例的多核资源,进一步缩短处理时间。
内容的提问来源于stack exchange,提问作者matsuo_basho
相关产品推荐
相关产品推荐

