You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何我的自闭症分类CNN交叉验证模型出现100%验证准确率?

问题分析与解决方案

这种情况完全不正常,你的10折交叉验证结果存在明显的代码逻辑错误,并非模型真实性能的体现。

核心问题:模型实例复用导致的权重污染

你的代码中,your_cnn_model = model这一行直接复用了循环外定义的同一个模型实例。这意味着:

  • 第一折训练完成后,模型已经拥有了针对第一折数据训练后的权重;
  • 从第二折开始,训练并非从零初始化模型,而是在之前训练好的权重基础上继续更新;
  • 随着fold迭代,模型逐渐记住了整个数据集的模式,最终在后续测试集上出现100%准确率的虚假结果,这本质是一种隐性的数据泄露。

其他可能的辅助问题

除了模型复用,还需要排查以下点:

  • 数据预处理泄露:如果归一化、标准化等预处理步骤使用了整个数据集的统计量(而非当前fold的训练集),会导致测试集信息提前流入训练过程;
  • 标签划分错误:确认labels没有与数据索引错位,避免出现测试集标签与训练集完全重叠的极端情况。

修正方案

  1. 在每个fold内重新初始化模型
    将模型的定义与编译逻辑放入循环内部,或者封装成函数,确保每一轮fold都使用全新的模型实例:

    from sklearn.model_selection import KFold
    import numpy as np
    from tensorflow.keras.models import Sequential
    from tensorflow.keras.layers import Conv2D, MaxPooling2D, Flatten, Dense
    
    def create_cnn_model(input_shape):
        # 定义你的CNN模型结构,替换为实际的网络层
        model = Sequential([
            Conv2D(32, (3,3), activation='relu', input_shape=input_shape),
            MaxPooling2D((2,2)),
            Flatten(),
            Dense(64, activation='relu'),
            Dense(1, activation='sigmoid')
        ])
        model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'])
        return model
    
    data = conn_matrices
    labels = y
    
    kf = KFold(n_splits=10, shuffle=True, random_state=42)
    fold_accuracies = []
    
    for train_index, test_index in kf.split(data):
        X_train, X_test = data[train_index], data[test_index]
        y_train, y_test = labels[train_index], labels[test_index]
        
        # 每折创建全新的模型实例
        your_cnn_model = create_cnn_model(data.shape[1:])
        
        # 可选:添加早停防止单fold内过拟合
        from tensorflow.keras.callbacks import EarlyStopping
        early_stop = EarlyStopping(monitor='val_accuracy', patience=3, restore_best_weights=True)
        
        your_cnn_model.fit(X_train, y_train, epochs=25,
                          batch_size=32, validation_data=(X_test, y_test), 
                          verbose=1, callbacks=[early_stop])
    
        accuracy = your_cnn_model.evaluate(X_test, y_test, verbose=0)[1]
        fold_accuracies.append(accuracy)
    
    for i, accuracy in enumerate(fold_accuracies):
        print(f"Fold {i+1} Accuracy: {accuracy:.4f}")
    
    mean_accuracy = np.mean(fold_accuracies)
    std_deviation = np.std(fold_accuracies)
    print(f"Mean Accuracy: {mean_accuracy:.4f}")
    print(f"Standard Deviation: {std_deviation:.4f}")
    
  2. 规范数据预处理流程
    对每个fold的训练集单独做预处理,避免跨fold的数据泄露:

    # 仅基于当前fold的训练集计算统计量
    mean = X_train.mean(axis=0)
    std = X_train.std(axis=0)
    X_train = (X_train - mean) / std
    X_test = (X_test - mean) / std  # 用训练集的统计量标准化测试集
    

预期结果

修正后,各fold的准确率会回到你预期的70%-77%区间,与文献水平和之前的单分割测试结果一致,交叉验证的均值也会更具参考价值。

内容的提问来源于stack exchange,提问作者Galib Rabat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 14:03:25