使用sklearn划分数据集时出现样本数量不匹配ValueError错误求助
解决sklearn train_test_split样本数不一致错误
从错误信息ValueError: Found input variables with inconsistent numbers of samples: [10001, 0]可以直接定位问题:你的images列表包含10001个元素,但annotations列表是空的,train_test_split要求传入的所有数组长度必须一致,因此触发报错。
问题根源
错误出在标注文件的筛选逻辑上:
annotations = [os.path.join('C:/Users/X3/1/text_files', x) for x in os.listdir('C:/Users/X3/1/text_files') if x[-3] == "txt"]
这里x[-3] == "txt"是取文件名的倒数第三个字符和字符串"txt"做比较,逻辑完全错误——单个字符不可能等于长度为3的字符串,导致筛选条件永远不成立,最终annotations为空列表。
修复方案
将筛选条件改为判断文件后缀是否为.txt,使用Python内置的endswith方法:
# 修复后的标注文件读取代码 image_dir = r"C:/Users/X3/pharmaceutical-drugs-and-vitamins-synthetic-images/ImageClassesCombinedWithCOCOAnnotations/images_raw" images = [os.path.join(image_dir, x) for x in os.listdir(image_dir)] # 修正筛选条件 annotations = [os.path.join('C:/Users/X3/1/text_files', x) for x in os.listdir('C:/Users/X3/1/text_files') if x.endswith(".txt")] images.sort() annotations.sort() # 先验证长度是否一致 print(len(images), len(annotations)) # 再执行数据集划分 train_images, val_images, train_annotations, val_annotations = train_test_split(images, annotations, test_size = 0.2, random_state = 1) val_images, test_images, val_annotations, test_annotations = train_test_split(val_images, val_annotations, test_size = 0.5, random_state = 1)
额外建议
如果你的图像和标注文件是一一对应的(比如图像名img001.jpg对应标注img001.txt),在排序后可以增加一层校验,确保每个图像都有匹配的标注:
# 校验文件名前缀是否匹配 for img_path, ann_path in zip(images, annotations): img_name = os.path.splitext(os.path.basename(img_path))[0] ann_name = os.path.splitext(os.path.basename(ann_path))[0] if img_name != ann_name: print(f"不匹配的文件对:{img_path} <-> {ann_path}")
内容的提问来源于stack exchange,提问作者02stream
相关产品推荐
相关产品推荐

