You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

已处理NaN、无穷值及超大值仍触发ValueError报错排查

报错根因

你遇到的ValueError和数据中的空值、无穷值无关,核心是模型类型和评估指标完全不匹配:

  • 你选用的DecisionTreeRegressor是回归模型,预测输出是连续浮点值。而accuracy_score是分类任务专用指标,它的计算逻辑是逐元素对比预测值和真实值是否完全相等,仅支持离散类别标签输入。
  • 连续浮点值受计算精度影响,预测值和真实值几乎不可能完全精确相等,直接传入就会触发值类型不匹配的报错。
  • 另外你代码中用reindex匹配切分后的数据集索引属于冗余操作,直接用.iloc按位置取数更稳妥,能避免索引乱序引发的潜在问题。
问题复现的原始代码与流程

首先完成了异常值与空值清洗,统一数据类型:

df.replace([np.inf, -np.inf], np.nan, inplace=True)
df = df.dropna()

数据预处理后样例

之后提取特征、标签构建10折交叉验证流程,触发报错的原代码如下:

X = df.iloc[:,:-1].astype('float32')
y = df.iloc[:,-1].astype('float32')

from sklearn.model_selection import KFold
from sklearn.tree import DecisionTreeRegressor 
from sklearn.metrics import accuracy_score
import matplotlib.pyplot as plt
k=10
kf = KFold(n_splits=k, random_state=1, shuffle=True)
reg = DecisionTreeRegressor(random_state=1)
acc_score = []

for train_index, test_index in kf.split(X):
    X_train, X_test = X.reindex(index = train_index), X.reindex(index = test_index)
    y_train, y_test = y.reindex(index = train_index), y.reindex(index = test_index)
    reg.fit(X_train, y_train)
    y_pred = reg.predict(X_test)
    acc_score.append(accuracy_score(y_pred, y_test))

报错信息截图

修正方案

根据实际任务类型二选一即可:

方案1:任务为回归预测(输出连续数值)

放弃使用accuracy_score,替换为回归任务专用评估指标,比如R²决定系数、平均绝对误差(MAE)、均方根误差(RMSE),修正后代码:

from sklearn.model_selection import KFold
from sklearn.tree import DecisionTreeRegressor 
from sklearn.metrics import r2_score, mean_absolute_error
import numpy as np

X = df.iloc[:,:-1].astype('float32')
y = df.iloc[:,-1].astype('float32')

k=10
kf = KFold(n_splits=k, random_state=1, shuffle=True)
reg = DecisionTreeRegressor(random_state=1)
r2_list = []
mae_list = []

for train_index, test_index in kf.split(X):
    # 直接按位置索引取数,移除冗余的reindex操作
    X_train, X_test = X.iloc[train_index], X.iloc[test_index]
    y_train, y_test = y.iloc[train_index], y.iloc[test_index]
    reg.fit(X_train, y_train)
    y_pred = reg.predict(X_test)
    r2_list.append(r2_score(y_test, y_pred))
    mae_list.append(mean_absolute_error(y_test, y_pred))

print(f"10折交叉验证平均R²得分:{np.mean(r2_list):.4f}")
print(f"10折交叉验证平均MAE得分:{np.mean(mae_list):.4f}")

方案2:任务为分类预测(输出离散类别)

放弃使用回归器DecisionTreeRegressor,替换为分类器DecisionTreeClassifier,此时可以正常使用accuracy_score计算准确率,修正后代码:

from sklearn.model_selection import KFold
from sklearn.tree import DecisionTreeClassifier
from sklearn.metrics import accuracy_score
import numpy as np

X = df.iloc[:,:-1].astype('float32')
# 分类任务标签需为离散整数类型
y = df.iloc[:,-1].astype('int32')

k=10
kf = KFold(n_splits=k, random_state=1, shuffle=True)
clf = DecisionTreeClassifier(random_state=1)
acc_list = []

for train_index, test_index in kf.split(X):
    X_train, X_test = X.iloc[train_index], X.iloc[test_index]
    y_train, y_test = y.iloc[train_index], y.iloc[test_index]
    clf.fit(X_train, y_train)
    y_pred = clf.predict(X_test)
    acc_list.append(accuracy_score(y_test, y_pred))

print(f"10折交叉验证平均分类准确率:{np.mean(acc_list):.4f}")

内容的提问来源于stack exchange,提问作者adithya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 00:54:32