You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

KNeighborsClassifier运行报错:输入含NaN值的KNN建模问题求助

问题:KNeighborsClassifier运行时因NaN值报错

运行基于KNeighborsClassifier的KNN建模代码时触发ValueError,报错位置在y_pred = neigh.predict(X_test_2)行,提示输入数据包含NaN,而KNN模型不原生支持缺失值。使用的是UCI在线新闻流行度数据集,代码来自对应的Kaggle探索性分析项目。

问题代码片段

# defining the model
from sklearn.neighbors import KNeighborsClassifier

k_range = np.arange(1,100)
accuracy = []

for n in k_range:
    neigh = KNeighborsClassifier(n_neighbors=n, n_jobs=-1)
    neigh.fit(X_train_2, y_train_2)

    # predict the result
    y_pred = neigh.predict(X_test_2)

    #print ("Random Forest Classifer Result")
    #print ("Performance - " + str(100*accuracy_score(y_pred, y_test_2)) + "%")
    accuracy.append(100*accuracy_score(y_pred, y_test_2))

报错信息

ValueError                                Traceback (most recent call last)
 in <cell line: 7>()
10
11     # predict the result
---> 12     y_pred = neigh.predict(X_test_2)
13
14     #print ("Random Forest Classifer Result")

4 frames
/usr/local/lib/python3.10/dist-packages/sklearn/utils/validation.py in _assert_all_finite(X, allow_nan, msg_dtype, estimator_name, input_name)
159                 "#estimators-that-handle-nan-values"
160             )
---> 161         raise ValueError(msg_err)
162
163

ValueError: Input X contains NaN.
KNeighborsClassifier does not accept missing values encoded as NaN natively. For supervised learning, you might want to consider sklearn.ensemble.HistGradientBoostingClassifier and Regressor which accept missing values encoded as NaNs natively. Alternatively, it is possible to preprocess the data, for instance by using an imputer transformer in a pipeline or drop samples with missing values. See <a href="https://scikit-learn.org/stable/modules/impute.html" rel="nofollow noreferrer">https://scikit-learn.org/stable/modules/impute.html</a> You can find a list of all estimators that handle NaN values at the following page: <a href="https://scikit-learn.org/stable/modules/impute.html#estimators-that-handle-nan-values" rel="nofollow noreferrer">https://scikit-learn.org/stable/modules/impute.html#estimators-that-handle-nan-values</a>
解决方案

要让KNeighborsClassifier正常运行,必须先处理数据中的NaN值,以下是几种可行方案:

方案1:删除含NaN的样本或特征

  • 删除样本:直接去掉训练集和测试集中包含NaN的行,注意同步处理特征集和标签集:
    # 处理训练集
    X_train_clean = X_train_2.dropna(axis=0)
    y_train_clean = y_train_2[X_train_clean.index]
    # 处理测试集
    X_test_clean = X_test_2.dropna(axis=0)
    y_test_clean = y_test_2[X_test_clean.index]
    
    之后用X_train_clean和y_train_clean训练模型,用X_test_clean执行预测。
  • 删除特征:如果某个特征的NaN占比极高,可直接删除该特征:
    # 删除含NaN的特征(列)
    X_train_clean = X_train_2.dropna(axis=1)
    X_test_clean = X_test_2[X_train_clean.columns]  # 测试集需与训练集保持特征一致
    

方案2:填充缺失值(基础版)

用统计值填充NaN,比如均值、中位数、众数,推荐结合Pipeline使用,避免数据泄露:

from sklearn.pipeline import Pipeline
from sklearn.impute import SimpleImputer
from sklearn.neighbors import KNeighborsClassifier

k_range = np.arange(1,100)
accuracy = []

for n in k_range:
    # 构建包含填充器和KNN的管道
    pipe = Pipeline([
        ('imputer', SimpleImputer(strategy='mean')),  # 可选median/most_frequent
        ('knn', KNeighborsClassifier(n_neighbors=n, n_jobs=-1))
    ])
    pipe.fit(X_train_2, y_train_2)
    y_pred = pipe.predict(X_test_2)
    accuracy.append(100*accuracy_score(y_pred, y_test_2))
  • 策略选择:数值特征常用mean或median(中位数更抗异常值),分类特征用most_frequent。

方案3:进阶缺失值填充

如果数据量充足,可尝试KNNImputer(用K近邻的均值填充NaN),同样结合Pipeline:

from sklearn.impute import KNNImputer
from sklearn.pipeline import Pipeline
from sklearn.neighbors import KNeighborsClassifier

k_range = np.arange(1,100)
accuracy = []

for n in k_range:
    pipe = Pipeline([
        ('imputer', KNNImputer(n_neighbors=5)),  # 可根据数据调整K值
        ('knn', KNeighborsClassifier(n_neighbors=n, n_jobs=-1))
    ])
    pipe.fit(X_train_2, y_train_2)
    y_pred = pipe.predict(X_test_2)
    accuracy.append(100*accuracy_score(y_pred, y_test_2))

内容的提问来源于stack exchange,提问作者skm

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 21:28:15