如何从大型图像数据集中筛选特定图像?基于索引选取指定数量图像的代码实现请求
当然可以!用索引筛选完全能实现你的需求
这种方法简单直接,能精准控制每个文件夹中选取的图像数量,不管是固定选前36张、按间隔索引选,还是随机选36张都能轻松做到。下面我给你两种常用场景的代码实现:
场景1:收集选中图像的路径(直接用于模型训练)
如果你不需要复制文件,只是想获取符合要求的图像路径来喂给神经网络,这段代码就够用:
import os # 替换成你的数据集根目录 dataset_root = "/path/to/your/daily_folders" selected_paths = [] # 遍历每个日期文件夹 for day_dir in os.listdir(dataset_root): day_dir_full = os.path.join(dataset_root, day_dir) # 跳过非文件夹的文件 if not os.path.isdir(day_dir_full): continue # 获取文件夹内所有图像文件(可根据你的图像后缀调整) image_files = [ f for f in os.listdir(day_dir_full) if f.lower().endswith(('.jpg', '.jpeg', '.png', '.bmp')) ] # 按文件名排序,保证索引的稳定性(避免每次运行结果不一致) image_files.sort() # 选取前36张(这里用索引切片,你也可以指定特定索引,比如[0,3,6,...]) selected_imgs = image_files[:36] # 收集完整路径 for img in selected_imgs: selected_paths.append(os.path.join(day_dir_full, img)) print(f"成功筛选出 {len(selected_paths)} 张图像,可直接用于数据加载")
场景2:将选中图像复制到新目录(方便后续管理)
如果想把筛选后的图像整理到一个新的数据集目录,方便后续复用,可以用shutil模块实现复制:
import os import shutil dataset_root = "/path/to/your/daily_folders" output_root = "/path/to/your/selected_dataset" # 创建输出根目录(如果不存在的话) os.makedirs(output_root, exist_ok=True) for day_dir in os.listdir(dataset_root): day_dir_full = os.path.join(dataset_root, day_dir) if not os.path.isdir(day_dir_full): continue # 创建对应日期的子文件夹 output_day_dir = os.path.join(output_root, day_dir) os.makedirs(output_day_dir, exist_ok=True) image_files = [ f for f in os.listdir(day_dir_full) if f.lower().endswith(('.jpg', '.jpeg', '.png', '.bmp')) ] image_files.sort() selected_imgs = image_files[:36] # 复制选中的图像 for img in selected_imgs: src = os.path.join(day_dir_full, img) dst = os.path.join(output_day_dir, img) shutil.copy2(src, dst) # copy2会保留文件元数据 print("图像筛选并复制完成,新数据集已生成!")
额外优化:随机选取36张(避免顺序偏差)
如果不想固定选前36张,而是随机从每个文件夹选36张(减少因图像顺序带来的潜在偏差),可以用random.sample替换索引切片:
import random # 替换场景中的selected_imgs行 selected_imgs = random.sample(image_files, 36)
注意事项
- 你提到每个文件夹有51-55张图像,所以无需处理图像数量不足36的情况;如果后续有数量不足的文件夹,可以加个判断跳过或者抛出提示。
- 建议对图像文件名排序,这样索引的逻辑更稳定,不会因为系统返回文件顺序不同导致结果变化。
内容的提问来源于stack exchange,提问作者Shivam Yadav
相关产品推荐
相关产品推荐

