You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无法计算决策树回归器性能:Repl正常运行但Katacoda报错求助

排查Katacoda中决策树回归模型性能计算失败的问题

嘿,我仔细看了你的代码和问题情况——在Repl上运行一切正常,但Katacoda环境里完全没法计算决策树回归器的模型性能,对吧?咱们来拆解一下问题根源,再给出修复方案。

核心问题分析

1. 波士顿数据集的兼容性问题

Katacoda的环境大概率用的是scikit-learn 1.0及以上版本,而经典的波士顿房价数据集(load_boston)在这个版本之后被官方移除了,原因是数据包含种族歧视相关的特征,存在伦理争议。这是最可能导致代码在Katacoda中直接报错的原因。

2. 测试集交叉验证的不合理性

你在测试集上用了cross_val_score(dt_reg, X_test,Y_test, cv=10):

  • 波士顿数据集总共只有506个样本,按默认拆分比例,测试集大概只有127个样本
  • 10折交叉验证意味着每折仅包含12-13个样本,样本量过小会导致交叉验证无法正常执行,甚至抛出错误

3. RandomizedSearchCV参数设置冗余

你的参数搜索空间只有max_depth=2-5四个值,但n_iter=90会让模型重复搜索90次,完全是资源浪费,甚至可能触发环境的性能限制。

修复后的可运行代码

# 替换为sklearn推荐的加州住房数据集(替代波士顿数据集)
from sklearn.datasets import fetch_california_housing
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeRegressor
from sklearn.model_selection import RandomizedSearchCV
from sklearn.model_selection import cross_val_score
import numpy as np

# 设置随机种子保证可复现性
np.random.seed(100)

# 加载替代数据集
housing = fetch_california_housing()
X, y = housing.data, housing.target

# 拆分训练集和测试集
X_train, X_test, Y_train, Y_test = train_test_split(X, y, random_state=30)
print(X_train.shape)
print(X_test.shape)

# 初始化并训练基础决策树回归器
dt_reg = DecisionTreeRegressor(random_state=1)
dt_reg.fit(X_train, Y_train)

# 训练集用交叉验证评估,测试集直接用模型预测评估(更合理)
print('训练集交叉验证R²分数(10折):', cross_val_score(dt_reg, X_train,Y_train, cv=10))
print('测试集R²分数:', dt_reg.score(X_test, Y_test))

# 预测测试集前两个样本
predicted = dt_reg.predict(X_test[:2])
print('前两个测试样本预测值:', predicted)

# 优化决策树的max_depth参数
max_depth = range(2, 6)
dt_reg_base = DecisionTreeRegressor()
random_grid = {'max_depth': max_depth}

# 调整n_iter为参数空间的大小(4),避免冗余计算
dt_random = RandomizedSearchCV(estimator=dt_reg_base, 
                               param_distributions=random_grid, 
                               n_iter=4,  # 仅搜索4个不同的max_depth值
                               cv=3, 
                               verbose=2, 
                               random_state=42, 
                               n_jobs=-1)
dt_random.fit(X_train, Y_train)

# 定义性能评估函数
def evaluate(model, test_features, test_labels):
    predictions = model.predict(test_features)
    errors = abs(predictions - test_labels)
    mape = 100 * np.mean(errors / test_labels)
    accuracy = 100 - mape
    print('\n模型性能')
    print('平均误差: {:0.4f}'.format(np.mean(errors)))
    print('准确率(100-MAPE): {:0.2f}%'.format(accuracy))
    return accuracy

# 评估最优模型
best_random = dt_random.best_estimator_
random_accuracy = evaluate(best_random, X_test, Y_test)
print('最优模型准确率:', random_accuracy)
print('最优max_depth参数:', dt_random.best_params_['max_depth'])

关键改动说明

  1. 数据集替换:用fetch_california_housing替代已被移除的load_boston,解决版本兼容性问题
  2. 测试集评估方式调整:测试集不再用交叉验证,直接调用model.score()计算R²值,或者用自定义函数计算MAE、MAPE,符合机器学习评估的最佳实践
  3. RandomizedSearchCV参数优化:将n_iter设置为参数空间的大小(4),避免无效重复计算,提升运行效率
  4. 输出信息优化:调整打印内容的可读性,避免歧义

内容的提问来源于stack exchange,提问作者Sam_2207

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 19:27:37