You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python根据YAML标签分类单文件夹图片 报错排查与实现

图片按标签分类脚本报错排查与实现

需求说明

现有存储文件名与对应标签映射关系的YAML文件,格式如下:

- 1003520587_a35cd70a-22ab-4f89-92f8-e44884d6225d_boolean: no-checkbox
- 1003522794_3f24dcc8-60f1-4959-b0f4-d00a9f723a46_boolean: filled
- 1003522794_3f24dcc8-60f1-4959-b0f4-d00a9f723a46_boolean: empty

所有待处理图片存放在同一文件夹,需要通过Python将图片按标签移动到对应标签命名的文件夹下。先后编写两版代码均运行报错。


报错原因逐一排查

第一版(读取YAML)报FileNotFoundError

运行抛出错误:

FileNotFoundError: [Errno 2] No such file or directory

核心问题有3个:

  • 未导入os模块,代码中调用os.path.join、os.makedirs没有依赖支撑
  • 目标根路径to_folder_base配置错误,路径开头缺失根目录斜杠/,写为r'home/user/Downloads/sorted_image/abschluss',会被识别为相对路径,从脚本当前运行目录开始查找,自然无法匹配实际绝对路径
  • YAML示例中存在重复文件名对应不同标签的条目,第一次移动后文件就不在原路径了,后续匹配到同文件名的条目时会直接触发文件不存在错误。

第二版(读取CSV)报KeyError: 0

运行抛出错误:

KeyError: 0

核心问题有4个:

  • 未导入os、shutil模块,路径处理、文件移动的方法无依赖
  • Pandas DataFrame取值逻辑错误:meta_ham[index]是按列名取列的写法,不能直接传行索引取整行数据,不存在名为0的列就会触发KeyError
  • 标签取值使用了未定义的变量label,没有绑定实际CSV中的标签列
  • 目标文件夹路径写死为字符串'label',没有传入实际读取到的标签值,无法实现按标签分类的效果。

正确可运行代码

推荐:直接适配原有YAML格式的实现

代码增加了文件存在性校验,避免重复映射条目导致的运行中断:

import os
import yaml
import shutil

# 路径配置 注意核对所有路径的拼写与开头斜杠
from_folder =  r'/home/user/Downloads/by_field/abschluss/'
to_folder_base = r'/home/user/Downloads/sorted_image/abschluss'
yaml_path = '/home/user/Downloads/Yaml/abschluss.yaml'

# 提前创建目标根目录
os.makedirs(to_folder_base, exist_ok=True)

with open(yaml_path, 'r', encoding='utf-8') as stream:
    parsed_yaml = yaml.safe_load(stream)
    for item in parsed_yaml:
        for filename, tag in item.items():
            img_name = f"{filename}.png"
            old_img_path = os.path.join(from_folder, img_name)
            # 跳过已移动/不存在的文件
            if not os.path.exists(old_img_path):
                print(f"跳过不存在的文件:{old_img_path}")
                continue
            # 创建对应标签的分类文件夹
            target_folder = os.path.join(to_folder_base, tag)
            os.makedirs(target_folder, exist_ok=True)
            new_img_path = os.path.join(target_folder, img_name)
            shutil.move(old_img_path, new_img_path)
            print(f"移动完成:{img_name} -> {tag}文件夹")

CSV格式适配版本

使用前先确认CSV包含两列:一列存不带后缀的图片文件名,一列存对应标签,将代码中的列名替换为CSV实际的列名即可:

import os
import pandas as pd
import shutil

from_folder = "/home/user/Downloads/images/Leibteil/"
to_folder_base = "/home/user/Downloads/sorted_image/"
csv_path = '/home/user/Downloads/csv/Leibteil.csv'

os.makedirs(to_folder_base, exist_ok=True)

meta_df = pd.read_csv(csv_path)
# 按行遍历CSV数据
for _, row in meta_df.iterrows():
    # 替换单引号内的字段名为CSV实际列名
    img_name = f"{row['filename']}.png"
    tag = row['label']
    old_img_path = os.path.join(from_folder, img_name)
    if not os.path.exists(old_img_path):
        print(f"跳过不存在的文件:{old_img_path}")
        continue
    target_folder = os.path.join(to_folder_base, tag)
    os.makedirs(target_folder, exist_ok=True)
    new_img_path = os.path.join(target_folder, img_name)
    shutil.move(old_img_path, new_img_path)
    print(f"移动完成:{img_name} -> {tag}文件夹")

内容的提问来源于stack exchange,提问作者gyaha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 16:06:22