You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决sklearn KNeighborsClassifier的FutureWarning问题?

解决KNeighborsClassifier预测时的FutureWarning问题

问题场景

运行以下kNN分类代码时,执行y_hat = neigh.predict(X_test)触发了FutureWarning,已知屏蔽警告的方法,但希望从根源解决:

import pandas as pd
from sklearn import preprocessing
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KNeighborsClassifier

df = pd.read_csv('https://cf-courses-data.s3.us.cloud-object-storage.appdomain.cloud/IBMDeveloperSkillsNetwork-ML0101EN-SkillsNetwork/labs/Module%203/data/teleCust1000t.csv')
X = df[['region', 'tenure','age', 'marital', 'address', 'income', 'ed', 'employ','retire', 'gender', 'reside']].values  #.astype(float)
y = df['custcat'].values
X = preprocessing.StandardScaler().fit(X).transform(X.astype(float))
X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.2, random_state=42)
neigh = KNeighborsClassifier(n_neighbors = 4).fit(X_train, y_train)
y_hat = neigh.predict(X_test)

警告内容:

FutureWarning: Unlike other reduction functions (e.g. skew, kurtosis), the default behavior of mode typically preserves the axis it acts along. In SciPy 1.11.0, this behavior will change: the default value of keepdims will become False, the axis over which the statistic is taken will be eliminated, and the value None will no longer be accepted. Set keepdims to True or False to avoid this warning. mode, _ = stats.mode(_y[neigh_ind, k], axis=1)

用户找到的屏蔽警告代码:

# 导入警告过滤器
from warnings import simplefilter
# 忽略所有未来警告
simplefilter(action='ignore', category=FutureWarning)

问题根源

这个警告的本质是:scikit-learn旧版本的KNeighborsClassifier内部调用scipy.stats.mode时,没有显式指定keepdims参数。而SciPy 1.11.0及后续版本会修改mode函数的默认keepdims行为,因此触发了兼容性警告。

根源修复方案

方案1:升级scikit-learn到最新稳定版

scikit-learn的后续版本(如1.2.0及以上)已经修复了这个问题,在内部调用stats.mode时显式设置了keepdims参数。直接升级库即可消除警告:

pip install --upgrade scikit-learn

方案2:自定义KNN类修改内部调用(不升级库时)

如果暂时无法升级scikit-learn,可以自定义一个继承KNeighborsClassifier的类,重写涉及stats.mode的方法,显式指定keepdims参数。示例代码如下:

import pandas as pd
import numpy as np
from sklearn import preprocessing
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KNeighborsClassifier
from scipy import stats

class FixedKNeighborsClassifier(KNeighborsClassifier):
    def _kneighbors_predict(self, X, return_distance=False):
        neigh_dist, neigh_ind = self.kneighbors(X)
        classes_ = self.classes_
        _y = self._y
        if not self.outputs_2d_:
            _y = self._y.reshape((-1, 1))
            classes_ = [self.classes_]

        n_samples = X.shape[0]
        n_outputs = len(classes_)
        predictions = np.zeros((n_samples, n_outputs), dtype=classes_[0].dtype)

        for k in range(n_outputs):
            # 显式设置keepdims=False,匹配未来SciPy的默认行为
            mode, _ = stats.mode(_y[neigh_ind, k], axis=1, keepdims=False)
            predictions[:, k] = mode

        if not self.outputs_2d_:
            predictions = predictions.ravel()

        if return_distance:
            return predictions, neigh_dist
        else:
            return predictions

# 使用自定义类替换原类
df = pd.read_csv('https://cf-courses-data.s3.us.cloud-object-storage.appdomain.cloud/IBMDeveloperSkillsNetwork-ML0101EN-SkillsNetwork/labs/Module%203/data/teleCust1000t.csv')
X = df[['region', 'tenure','age', 'marital', 'address', 'income', 'ed', 'employ','retire', 'gender', 'reside']].values  #.astype(float)
y = df['custcat'].values
X = preprocessing.StandardScaler().fit(X).transform(X.astype(float))
X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.2, random_state=42)
neigh = FixedKNeighborsClassifier(n_neighbors = 4).fit(X_train, y_train)
y_hat = neigh.predict(X_test)

注:这里设置keepdims=False是为了匹配SciPy 1.11.0后的默认行为,确保代码在未来版本也能正常运行。

内容的提问来源于stack exchange,提问作者user_n_8093

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 03:16:04