You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

训练隐私预测模型遇ValueError:数据样本数不匹配问题排查

隐私预测模型训练时的数据维度不匹配问题分析

问题描述

我正尝试基于JSON文件中的标签训练一个隐私预测模型,但训练过程中遇到如下错误:

ValueError: Data cardinality is ambiguous. Make sure all arrays contain the same number of samples. 'x' sizes: 1988, 'y' sizes: 1992

相关代码

import os
import json
import numpy as np
from PIL import Image
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import MultiLabelBinarizer

# 处理单个JSON文件的函数
def process_json(json_file):
    images = []
    labels = []
    with open(json_file, 'r') as f:
        data = json.load(f)
    # 如果数据是字典,转为包含该字典的列表
    if isinstance(data, dict):
        data = [data]
    for item in data:
        image_path = item['image_path']
        label = item['labels']
        # 将图像路径和标签加入列表
        images.append(image_path)
        labels.append(label)
    return images, labels

# 加载并预处理图像的函数
from PIL import Image

def load_and_preprocess_images(image_paths, target_size=(112, 112)):
    images = []
    for image_path in image_paths:
        # 加载图像
        image = Image.open(image_path)
        # 调整图像尺寸
        image = image.resize(target_size)
        # 转为numpy数组
        image_array = np.array(image)
        # 打印图像形状
        print("Image shape:", image_array.shape)
        # 检查图像是否符合预期形状
        if image_array.shape == (112, 112, 3):
            images.append(image_array)
    return np.array(images)

# 图像目录
image_dir = '/Users/norah/Desktop/Privacylables_preprossing/images112/train2017'
# JSON文件目录
json_dir = '/Users/norah/Desktop/Privacylables_preprossing/annotations/train2017'

# 存储图像路径和标签的列表
all_image_paths = []
all_labels = []

# 处理每张图像及其对应的JSON文件
for image_file in os.listdir(image_dir):
    if image_file.endswith('.jpg'):
        # 获取图像路径
        image_path = os.path.join(image_dir, image_file)
        # 获取对应的JSON文件
        json_file = os.path.join(json_dir, image_file.replace('.jpg', '.json'))
        # 处理JSON文件
        images, labels = process_json(json_file)
        # 扩展列表
        all_image_paths.extend([image_path] * len(images))
        all_labels.extend(labels)

# 获取所有唯一标签
unique_labels = np.unique([label for sublist in all_labels for label in sublist])

# 初始化存储编码后标签的空数组
all_labels_encoded = np.zeros((len(all_labels), len(unique_labels)), dtype=np.int32)

# 对标签进行独热编码
for i, label_list in enumerate(all_labels):
    for label in label_list:
        label_index = np.where(unique_labels == label)[0][0]
        all_labels_encoded[i, label_index] = 1

# 拆分训练集和测试集
X_train_paths, X_test_paths, y_train, y_test = train_test_split(all_image_paths, all_labels_encoded, test_size=0.2, random_state=42)

# 加载并预处理训练集和测试集的图像
X_train = load_and_preprocess_images(X_train_paths)
X_test = load_and_preprocess_images(X_test_paths)

# 打印数据形状
print("Shape of X_train:", X_train.shape)
print("Shape of X_test:", X_test.shape)
print("Shape of y_train:", y_train.shape)
print("Shape of y_test:", y_test.shape)

import tensorflow as tf
from tensorflow.keras import layers, models

# 定义CNN模型结构
model = models.Sequential([
    layers.Conv2D(32, (3, 3), activation='relu', input_shape=(112, 112, 3)),
    layers.MaxPooling2D((2, 2)),
    layers.Conv2D(64, (3, 3), activation='relu'),
    layers.MaxPooling2D((2, 2)),
    layers.Conv2D(64, (3, 3), activation='relu'),
    layers.Flatten(),
    layers.Dense(128, activation='relu'),
    layers.Dense(68, activation='sigmoid')  # 多标签分类使用sigmoid激活
])

# 编译模型
model.compile(optimizer='adam',
              loss='binary_crossentropy',
              metrics=['accuracy'])

# 训练模型
history = model.fit(np.array(X_train), np.array(y_train), epochs=10, batch_size=32, validation_split=0.2)

# 评估模型
loss, accuracy = model.evaluate(np.array(X_test), np.array(y_test))
print("Test Loss:", loss)
print("Test Accuracy:", accuracy)

问题成因

  • 图像预处理过滤了有效样本:load_and_preprocess_images函数仅保留形状为(112,112,3)的图像,若部分图像是灰度图(形状为(112,112)或(112,112,1)),调整尺寸后仍不满足三通道要求,会被直接丢弃,导致X_train样本数少于y_train。
  • 数据拆分顺序错误:先拆分图像路径与标签,再对图像做过滤处理,过滤后的图像样本数无法与已拆分的标签样本数对应,最终出现X和Y样本数不一致的情况。

修复方案

1. 统一图像为三通道

修改图像预处理函数,强制将所有图像转为RGB格式,确保形状符合要求,避免样本被过滤:

def load_and_preprocess_images(image_paths, target_size=(112, 112)):
    images = []
    for image_path in image_paths:
        # 加载图像并强制转为RGB三通道
        image = Image.open(image_path).convert('RGB')
        image = image.resize(target_size)
        image_array = np.array(image)
        images.append(image_array)
    return np.array(images)

2. 调整数据处理顺序(保留过滤逻辑时使用)

如果确实需要过滤不符合要求的图像,应先完成所有图像的预处理和过滤,再拆分训练集和测试集,确保X与Y的样本数始终匹配:

# 先加载所有图像
all_images = load_and_preprocess_images(all_image_paths)
# 筛选出有效样本的索引
valid_indices = []
for idx, img in enumerate(all_images):
    if img.shape == (112, 112, 3):
        valid_indices.append(idx)
# 过滤图像和标签
all_images_filtered = all_images[valid_indices]
all_labels_encoded_filtered = all_labels_encoded[valid_indices]
# 再拆分数据集
X_train, X_test, y_train, y_test = train_test_split(all_images_filtered, all_labels_encoded_filtered, test_size=0.2, random_state=42)

3. 校验数据对应关系

检查每个JSON文件中的image_path是否与实际图像路径一致,避免因重复添加或遗漏图像路径导致样本数不匹配。

内容的提问来源于stack exchange,提问作者Areej Almousa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 01:31:00