KNeighborsClassifier结合10折交叉验证始终预测相同值如何解决?
你在结合KNN与10折交叉验证对全量数据集测试时,出现所有折叠的预测结果全部为类别3的问题,真实标签包含1、2、3三类。
问题代码
import numpy as np import pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv) import os import scipy.io from sklearn.neighbors import KNeighborsClassifier from sklearn import metrics from sklearn.model_selection import train_test_split from sklearn.preprocessing import StandardScaler from torch.utils.data import Dataset, DataLoader from sklearn import preprocessing import torch import numpy as np from sklearn.model_selection import KFold from sklearn.neighbors import KNeighborsClassifier from sklearn import metrics def load_mat_data(path): mat = scipy.io.loadmat(DATA_PATH) x,y = mat['data'], mat['class'] x = x.astype('float32') # stadardize values standardizer = preprocessing.StandardScaler() x = standardizer.fit_transform(x) return x, standardizer, y def numpyToTensor(x): x_train = torch.from_numpy(x) return x_train class DataBuilder(Dataset): def __init__(self, path): self.x, self.standardizer, self.y = load_mat_data(DATA_PATH) self.x = numpyToTensor(self.x) self.len=self.x.shape[0] self.y = numpyToTensor(self.y) def __getitem__(self,index): return (self.x[index], self.y[index]) def __len__(self): return self.len datasets = ['/home/katerina/Desktop/datasets/GSE75110.mat'] for DATA_PATH in datasets: print(DATA_PATH) data_set=DataBuilder(DATA_PATH) pred_rpknn = [0] * len(data_set.y) kf = KFold(n_splits=10, shuffle = True, random_state=7) for train_index, test_index in kf.split(data_set.x): #Create KNN Classifier knn = KNeighborsClassifier(n_neighbors=5) #print("TRAIN:", train_index, "TEST:", test_index) x_train, x_test = data_set.x[train_index], data_set.x[test_index] y_train, y_test = data_set.y[train_index], data_set.y[test_index] #Train the model using the training sets y1_train = y_train.ravel() knn.fit(x_train, y1_train) #Predict the response for test dataset y_pred = knn.predict(x_test) #print(y_pred) # Model Accuracy, how often is the classifier correct? print("Accuracy:",metrics.accuracy_score(y_test, y_pred)) c = 0 for idx in test_index: pred_rpknn[idx] = y_pred[c] c +=1 print("Accuracy:",metrics.accuracy_score(data_set.y, pred_rpknn)) print(pred_rpknn, data_set.y.reshape(1,-1))
输出结果
/home/katerina/Desktop/datasets/GSE75110.mat Accuracy: 0.2857142857142857 Accuracy: 0.38095238095238093 Accuracy: 0.14285714285714285 Accuracy: 0.4 Accuracy: 0.3 Accuracy: 0.25 Accuracy: 0.3 Accuracy: 0.6 Accuracy: 0.25 Accuracy: 0.45 Accuracy: 0.33497536945812806 [3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3]
真实标签
tensor([[1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1,1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 2, 2, 2,2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2,2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2,2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2,2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 3, 3, 3, 3, 3, 3, 3, 3, 3,3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3]])
解决方案
- 修正代码逻辑错误:
load_mat_data函数定义时传入了形参path,但实际调用scipy.io.loadmat时用的是全局变量DATA_PATH,建议修改为mat = scipy.io.loadmat(path),避免后续修改路径时出现逻辑错误。 - 修正数据预处理逻辑:你当前是对全量数据集做标准化,会引入数据泄露,应该在每个交叉验证折叠内,仅用当前训练集拟合StandardScaler,再分别对训练集和测试集做变换,不要提前处理全量数据。
- 调整K值:目前使用的k=5不一定适配你的数据集,建议从小到大尝试不同k值(如1、3、7、9、11),避免k值过大导致预测偏向多数类。
- 更换距离度量:高维特征下欧氏距离的区分度会大幅下降,可以尝试更换为曼哈顿距离、余弦距离,示例:
KNeighborsClassifier(n_neighbors=5, metric='cosine')。 - 优化权重策略:可以把KNN的权重设置为按距离加权,让距离更近的邻居对预测结果影响更大,降低多数类的干扰,示例:
KNeighborsClassifier(n_neighbors=5, weights='distance')。 - 过滤冗余特征:计算每个特征和标签的相关性,去掉和标签完全无关的噪声特征,提升距离计算的有效性。
- 处理类别不平衡:如果确认类别分布存在较大偏差,可以对少数类做过采样、对多数类做欠采样,平衡不同类别的样本量。
内容的提问来源于stack exchange,提问作者KateB

