如何解决sklearn KNeighborsClassifier的FutureWarning问题?
问题场景
运行以下kNN分类代码时,执行y_hat = neigh.predict(X_test)触发了FutureWarning,已知屏蔽警告的方法,但希望从根源解决:
import pandas as pd from sklearn import preprocessing from sklearn.model_selection import train_test_split from sklearn.neighbors import KNeighborsClassifier df = pd.read_csv('https://cf-courses-data.s3.us.cloud-object-storage.appdomain.cloud/IBMDeveloperSkillsNetwork-ML0101EN-SkillsNetwork/labs/Module%203/data/teleCust1000t.csv') X = df[['region', 'tenure','age', 'marital', 'address', 'income', 'ed', 'employ','retire', 'gender', 'reside']].values #.astype(float) y = df['custcat'].values X = preprocessing.StandardScaler().fit(X).transform(X.astype(float)) X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.2, random_state=42) neigh = KNeighborsClassifier(n_neighbors = 4).fit(X_train, y_train) y_hat = neigh.predict(X_test)
警告内容:
FutureWarning: Unlike other reduction functions (e.g.
skew,kurtosis), the default behavior ofmodetypically preserves the axis it acts along. In SciPy 1.11.0, this behavior will change: the default value ofkeepdimswill become False, theaxisover which the statistic is taken will be eliminated, and the value None will no longer be accepted. Setkeepdimsto True or False to avoid this warning. mode, _ = stats.mode(_y[neigh_ind, k], axis=1)
用户找到的屏蔽警告代码:
# 导入警告过滤器 from warnings import simplefilter # 忽略所有未来警告 simplefilter(action='ignore', category=FutureWarning)
问题根源
这个警告的本质是:scikit-learn旧版本的KNeighborsClassifier内部调用scipy.stats.mode时,没有显式指定keepdims参数。而SciPy 1.11.0及后续版本会修改mode函数的默认keepdims行为,因此触发了兼容性警告。
根源修复方案
方案1:升级scikit-learn到最新稳定版
scikit-learn的后续版本(如1.2.0及以上)已经修复了这个问题,在内部调用stats.mode时显式设置了keepdims参数。直接升级库即可消除警告:
pip install --upgrade scikit-learn
方案2:自定义KNN类修改内部调用(不升级库时)
如果暂时无法升级scikit-learn,可以自定义一个继承KNeighborsClassifier的类,重写涉及stats.mode的方法,显式指定keepdims参数。示例代码如下:
import pandas as pd import numpy as np from sklearn import preprocessing from sklearn.model_selection import train_test_split from sklearn.neighbors import KNeighborsClassifier from scipy import stats class FixedKNeighborsClassifier(KNeighborsClassifier): def _kneighbors_predict(self, X, return_distance=False): neigh_dist, neigh_ind = self.kneighbors(X) classes_ = self.classes_ _y = self._y if not self.outputs_2d_: _y = self._y.reshape((-1, 1)) classes_ = [self.classes_] n_samples = X.shape[0] n_outputs = len(classes_) predictions = np.zeros((n_samples, n_outputs), dtype=classes_[0].dtype) for k in range(n_outputs): # 显式设置keepdims=False,匹配未来SciPy的默认行为 mode, _ = stats.mode(_y[neigh_ind, k], axis=1, keepdims=False) predictions[:, k] = mode if not self.outputs_2d_: predictions = predictions.ravel() if return_distance: return predictions, neigh_dist else: return predictions # 使用自定义类替换原类 df = pd.read_csv('https://cf-courses-data.s3.us.cloud-object-storage.appdomain.cloud/IBMDeveloperSkillsNetwork-ML0101EN-SkillsNetwork/labs/Module%203/data/teleCust1000t.csv') X = df[['region', 'tenure','age', 'marital', 'address', 'income', 'ed', 'employ','retire', 'gender', 'reside']].values #.astype(float) y = df['custcat'].values X = preprocessing.StandardScaler().fit(X).transform(X.astype(float)) X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.2, random_state=42) neigh = FixedKNeighborsClassifier(n_neighbors = 4).fit(X_train, y_train) y_hat = neigh.predict(X_test)
注:这里设置keepdims=False是为了匹配SciPy 1.11.0后的默认行为,确保代码在未来版本也能正常运行。
内容的提问来源于stack exchange,提问作者user_n_8093

