KNeighborsClassifier predict方法抛出“Expected 2D array, got 1D array instead”异常的问题排查
你正在开发图像相似度算法,用cv2.calcHist提取图像特征后将其以numpy.float64列表的形式存入JSON文件,之后加载数据用sklearn的KNeighborsClassifier做分类,但在调用predict方法时遇到了维度不匹配的错误——明明X_train和X_test的形状看起来是对的((36,4096)和(9,4096)),但fit能成功运行,predict却报错说收到了1D数组。你尝试了各种reshape操作,但要么导致特征数不匹配,要么样本数不一致,始终没解决问题。
你的加载和训练代码如下:
import numpy as np import json from sklearn.neighbors import KNeighborsClassifier from sklearn.model_selection import train_test_split from sklearn.metrics.pairwise import cosine_similarity with open('data.json') as f: jsonData = json.load(f) X = [] y = [] for image in jsonData['images']: embeddingData = image['histogram'] X.append(embeddingData) y.append(image['classification']) X = np.array(X) y = np.array(y) #split dataset into train and test data X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=1, stratify=y) print('Shape of X_train:') print(X_train.shape) print('Shape of X_test:') print(X_test.shape) print('Shape of y_train:') print(y_train.shape) # Create KNN classifier knn = KNeighborsClassifier(n_neighbors = 1, metric=cosine_similarity) # Fit the classifier to the data knn.fit(X_train, y_train) #show predictions on the test data y_pred = knn.predict(X_test)
运行后在y_pred = knn.predict(X_test)这一行抛出的错误:
ValueError: Expected 2D array, got 1D array instead: array=[1.13707140e-01 9.81128156e-01 2.89475545e-02 ... 0.00000000e+00 5.02811105e-04 1.15502894e-01]. Reshape your data either using array.reshape(-1, 1) if your data has a single feature or array.reshape(1, -1) if it contains a single sample.
你尝试reshape后又遇到的错误:
- 执行
y_pred = knn.predict(X_test.reshape(-1, 1))时:
ValueError: X has 1 features, but KNeighborsClassifier is expecting 4096 features as input.
- 执行
knn.fit(X_train.reshape(-1, 1), y_train)时:
ValueError: Found input variables with inconsistent numbers of samples: [147456, 36]
问题根源分析
你遇到的核心问题是错误地将cosine_similarity函数直接传给了KNeighborsClassifier的metric参数。
sklearn的KNeighborsClassifier对metric参数的要求是:要么是字符串(比如'cosine'),要么是符合特定签名的自定义距离函数——自定义函数需要接收两个1D数组(单个样本),返回一个标量距离值;但sklearn.metrics.pairwise.cosine_similarity的设计是接收两个2D数组(样本集合),返回一个相似度矩阵,完全不符合KNN对metric函数的要求。
当你用cosine_similarity作为metric时,KNN内部在计算距离时会错误地处理输入维度,导致在predict阶段把整个测试集当成了单个1D样本,从而抛出维度不匹配的错误。而fit阶段没报错,是因为fit时的逻辑和predict时的样本处理逻辑不同,刚好没触发维度检查的问题。
解决方案
只需要把metric参数改成字符串'cosine'即可,sklearn会自动使用对应的余弦距离计算逻辑:
修改KNN初始化的代码:
# Create KNN classifier knn = KNeighborsClassifier(n_neighbors = 1, metric='cosine')
然后直接运行原代码,不需要对X_train或X_test做额外的reshape操作——因为你的X_train和X_test的形状已经是正确的2D数组(样本数×特征数)。
验证逻辑
修改后,KNN会正确识别输入的2D数组:
X_train是(36,4096),代表36个样本,每个样本4096个特征X_test是(9,4096),代表9个测试样本,特征数和训练集一致
此时predict方法会正常处理每个测试样本,不会再抛出维度错误。
备注:内容来源于stack exchange,提问作者papaya

