已处理NaN、无穷值及超大值仍触发ValueError报错排查
报错根因
你遇到的ValueError和数据中的空值、无穷值无关,核心是模型类型和评估指标完全不匹配:
- 你选用的
DecisionTreeRegressor是回归模型,预测输出是连续浮点值。而accuracy_score是分类任务专用指标,它的计算逻辑是逐元素对比预测值和真实值是否完全相等,仅支持离散类别标签输入。 - 连续浮点值受计算精度影响,预测值和真实值几乎不可能完全精确相等,直接传入就会触发值类型不匹配的报错。
- 另外你代码中用
reindex匹配切分后的数据集索引属于冗余操作,直接用.iloc按位置取数更稳妥,能避免索引乱序引发的潜在问题。
问题复现的原始代码与流程
首先完成了异常值与空值清洗,统一数据类型:
df.replace([np.inf, -np.inf], np.nan, inplace=True) df = df.dropna()

之后提取特征、标签构建10折交叉验证流程,触发报错的原代码如下:
X = df.iloc[:,:-1].astype('float32') y = df.iloc[:,-1].astype('float32') from sklearn.model_selection import KFold from sklearn.tree import DecisionTreeRegressor from sklearn.metrics import accuracy_score import matplotlib.pyplot as plt k=10 kf = KFold(n_splits=k, random_state=1, shuffle=True) reg = DecisionTreeRegressor(random_state=1) acc_score = [] for train_index, test_index in kf.split(X): X_train, X_test = X.reindex(index = train_index), X.reindex(index = test_index) y_train, y_test = y.reindex(index = train_index), y.reindex(index = test_index) reg.fit(X_train, y_train) y_pred = reg.predict(X_test) acc_score.append(accuracy_score(y_pred, y_test))

修正方案
根据实际任务类型二选一即可:
方案1:任务为回归预测(输出连续数值)
放弃使用accuracy_score,替换为回归任务专用评估指标,比如R²决定系数、平均绝对误差(MAE)、均方根误差(RMSE),修正后代码:
from sklearn.model_selection import KFold from sklearn.tree import DecisionTreeRegressor from sklearn.metrics import r2_score, mean_absolute_error import numpy as np X = df.iloc[:,:-1].astype('float32') y = df.iloc[:,-1].astype('float32') k=10 kf = KFold(n_splits=k, random_state=1, shuffle=True) reg = DecisionTreeRegressor(random_state=1) r2_list = [] mae_list = [] for train_index, test_index in kf.split(X): # 直接按位置索引取数,移除冗余的reindex操作 X_train, X_test = X.iloc[train_index], X.iloc[test_index] y_train, y_test = y.iloc[train_index], y.iloc[test_index] reg.fit(X_train, y_train) y_pred = reg.predict(X_test) r2_list.append(r2_score(y_test, y_pred)) mae_list.append(mean_absolute_error(y_test, y_pred)) print(f"10折交叉验证平均R²得分:{np.mean(r2_list):.4f}") print(f"10折交叉验证平均MAE得分:{np.mean(mae_list):.4f}")
方案2:任务为分类预测(输出离散类别)
放弃使用回归器DecisionTreeRegressor,替换为分类器DecisionTreeClassifier,此时可以正常使用accuracy_score计算准确率,修正后代码:
from sklearn.model_selection import KFold from sklearn.tree import DecisionTreeClassifier from sklearn.metrics import accuracy_score import numpy as np X = df.iloc[:,:-1].astype('float32') # 分类任务标签需为离散整数类型 y = df.iloc[:,-1].astype('int32') k=10 kf = KFold(n_splits=k, random_state=1, shuffle=True) clf = DecisionTreeClassifier(random_state=1) acc_list = [] for train_index, test_index in kf.split(X): X_train, X_test = X.iloc[train_index], X.iloc[test_index] y_train, y_test = y.iloc[train_index], y.iloc[test_index] clf.fit(X_train, y_train) y_pred = clf.predict(X_test) acc_list.append(accuracy_score(y_test, y_pred)) print(f"10折交叉验证平均分类准确率:{np.mean(acc_list):.4f}")
内容的提问来源于stack exchange,提问作者adithya
相关产品推荐
相关产品推荐

