You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch Dataset类多层嵌套目录图片路径获取提速咨询

优化多层文件夹图片路径收集的方案

核心优化思路

你的场景属于IO密集型操作(遍历文件系统),可以从简化遍历逻辑、减少计算开销、并行处理三个方向提速:


1. 简化目录遍历逻辑

原代码的三层嵌套glob可以直接用多层路径匹配替代,减少冗余循环层级:

p = Path(self.root_dir) / ("Training" if self.is_train else "Validation")
image_p = p / "01.원천데이터" / f"{'T' if self.is_train else 'V'}S_images"

# 直接匹配所有三层子目录(对应原sentence_dir -> true_false_dir -> posture_dir)
posture_dirs = list(image_p.glob("*/*/*"))

2. 替换排序操作,降低计算开销

原代码用sorted(list(...))[-1]取最后一张图,需要对目录下所有文件全量排序(时间复杂度O(n log n))。改用max()直接找到最大文件名(按Path自然排序,时间复杂度O(n)),大幅节省时间:

# 单线程优化版本
image_path_list = []
for dir_path in posture_dirs:
    # 只筛选文件,避免误取子文件夹
    last_image = max(f for f in dir_path.glob("*") if f.is_file())
    image_path_list.append(str(last_image))

3. 多线程并行处理(IO密集场景首选)

文件系统操作属于IO密集型,Python的GIL在IO等待时会释放,用多线程可同时处理多个目录的文件查找,进一步提升速度:

from concurrent.futures import ThreadPoolExecutor

def get_last_image(dir_path):
    # 封装单个目录的处理逻辑
    return str(max(f for f in dir_path.glob("*") if f.is_file()))

# 创建线程池并行处理所有目录
with ThreadPoolExecutor() as executor:
    image_path_list = list(executor.map(get_last_image, posture_dirs))

额外注意事项

  • 如果目录结构可能变化,glob("*/*/*")可换成image_p.rglob("*")后筛选目录深度,但固定层级的glob效率更高。
  • 线程池大小可手动指定(比如ThreadPoolExecutor(max_workers=8)),一般设为CPU核心数的2-4倍即可,过多线程会增加调度开销。

内容的提问来源于stack exchange,提问作者djmoon13

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 00:42:49