训练隐私预测模型遇ValueError:数据样本数不匹配问题排查
隐私预测模型训练时的数据维度不匹配问题分析
问题描述
我正尝试基于JSON文件中的标签训练一个隐私预测模型,但训练过程中遇到如下错误:
ValueError: Data cardinality is ambiguous. Make sure all arrays contain the same number of samples. 'x' sizes: 1988, 'y' sizes: 1992
相关代码
import os import json import numpy as np from PIL import Image from sklearn.model_selection import train_test_split from sklearn.preprocessing import MultiLabelBinarizer # 处理单个JSON文件的函数 def process_json(json_file): images = [] labels = [] with open(json_file, 'r') as f: data = json.load(f) # 如果数据是字典,转为包含该字典的列表 if isinstance(data, dict): data = [data] for item in data: image_path = item['image_path'] label = item['labels'] # 将图像路径和标签加入列表 images.append(image_path) labels.append(label) return images, labels # 加载并预处理图像的函数 from PIL import Image def load_and_preprocess_images(image_paths, target_size=(112, 112)): images = [] for image_path in image_paths: # 加载图像 image = Image.open(image_path) # 调整图像尺寸 image = image.resize(target_size) # 转为numpy数组 image_array = np.array(image) # 打印图像形状 print("Image shape:", image_array.shape) # 检查图像是否符合预期形状 if image_array.shape == (112, 112, 3): images.append(image_array) return np.array(images) # 图像目录 image_dir = '/Users/norah/Desktop/Privacylables_preprossing/images112/train2017' # JSON文件目录 json_dir = '/Users/norah/Desktop/Privacylables_preprossing/annotations/train2017' # 存储图像路径和标签的列表 all_image_paths = [] all_labels = [] # 处理每张图像及其对应的JSON文件 for image_file in os.listdir(image_dir): if image_file.endswith('.jpg'): # 获取图像路径 image_path = os.path.join(image_dir, image_file) # 获取对应的JSON文件 json_file = os.path.join(json_dir, image_file.replace('.jpg', '.json')) # 处理JSON文件 images, labels = process_json(json_file) # 扩展列表 all_image_paths.extend([image_path] * len(images)) all_labels.extend(labels) # 获取所有唯一标签 unique_labels = np.unique([label for sublist in all_labels for label in sublist]) # 初始化存储编码后标签的空数组 all_labels_encoded = np.zeros((len(all_labels), len(unique_labels)), dtype=np.int32) # 对标签进行独热编码 for i, label_list in enumerate(all_labels): for label in label_list: label_index = np.where(unique_labels == label)[0][0] all_labels_encoded[i, label_index] = 1 # 拆分训练集和测试集 X_train_paths, X_test_paths, y_train, y_test = train_test_split(all_image_paths, all_labels_encoded, test_size=0.2, random_state=42) # 加载并预处理训练集和测试集的图像 X_train = load_and_preprocess_images(X_train_paths) X_test = load_and_preprocess_images(X_test_paths) # 打印数据形状 print("Shape of X_train:", X_train.shape) print("Shape of X_test:", X_test.shape) print("Shape of y_train:", y_train.shape) print("Shape of y_test:", y_test.shape) import tensorflow as tf from tensorflow.keras import layers, models # 定义CNN模型结构 model = models.Sequential([ layers.Conv2D(32, (3, 3), activation='relu', input_shape=(112, 112, 3)), layers.MaxPooling2D((2, 2)), layers.Conv2D(64, (3, 3), activation='relu'), layers.MaxPooling2D((2, 2)), layers.Conv2D(64, (3, 3), activation='relu'), layers.Flatten(), layers.Dense(128, activation='relu'), layers.Dense(68, activation='sigmoid') # 多标签分类使用sigmoid激活 ]) # 编译模型 model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy']) # 训练模型 history = model.fit(np.array(X_train), np.array(y_train), epochs=10, batch_size=32, validation_split=0.2) # 评估模型 loss, accuracy = model.evaluate(np.array(X_test), np.array(y_test)) print("Test Loss:", loss) print("Test Accuracy:", accuracy)
问题成因
- 图像预处理过滤了有效样本:
load_and_preprocess_images函数仅保留形状为(112,112,3)的图像,若部分图像是灰度图(形状为(112,112)或(112,112,1)),调整尺寸后仍不满足三通道要求,会被直接丢弃,导致X_train样本数少于y_train。 - 数据拆分顺序错误:先拆分图像路径与标签,再对图像做过滤处理,过滤后的图像样本数无法与已拆分的标签样本数对应,最终出现X和Y样本数不一致的情况。
修复方案
1. 统一图像为三通道
修改图像预处理函数,强制将所有图像转为RGB格式,确保形状符合要求,避免样本被过滤:
def load_and_preprocess_images(image_paths, target_size=(112, 112)): images = [] for image_path in image_paths: # 加载图像并强制转为RGB三通道 image = Image.open(image_path).convert('RGB') image = image.resize(target_size) image_array = np.array(image) images.append(image_array) return np.array(images)
2. 调整数据处理顺序(保留过滤逻辑时使用)
如果确实需要过滤不符合要求的图像,应先完成所有图像的预处理和过滤,再拆分训练集和测试集,确保X与Y的样本数始终匹配:
# 先加载所有图像 all_images = load_and_preprocess_images(all_image_paths) # 筛选出有效样本的索引 valid_indices = [] for idx, img in enumerate(all_images): if img.shape == (112, 112, 3): valid_indices.append(idx) # 过滤图像和标签 all_images_filtered = all_images[valid_indices] all_labels_encoded_filtered = all_labels_encoded[valid_indices] # 再拆分数据集 X_train, X_test, y_train, y_test = train_test_split(all_images_filtered, all_labels_encoded_filtered, test_size=0.2, random_state=42)
3. 校验数据对应关系
检查每个JSON文件中的image_path是否与实际图像路径一致,避免因重复添加或遗漏图像路径导致样本数不匹配。
内容的提问来源于stack exchange,提问作者Areej Almousa
相关产品推荐
相关产品推荐

