You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为RandomForestRegressor模型计算绘制房价预测置信区间

随机森林回归预测置信区间计算与可视化落地方案

无需安装额外第三方扩展库,直接基于scikit-learn原生接口即可实现,不存在版本兼容问题,可直接接入你现有代码流程。

核心原理:随机森林由多棵独立训练的决策树组成,每棵子树会输出独立的预测结果,统计所有子树对同一样本的预测值分布,取对应置信水平的分位数即可得到置信区间上下界。


可直接复用的代码

将你原有代码中pred = reg.predict(xtest)及之后的部分替换为以下代码即可:

# 计算所有子树对测试集的独立预测结果,输出维度为(子树总数, 测试集样本数)
all_subtree_predictions = np.array([tree.predict(xtest.to_numpy()) for tree in reg.estimators_])

# 配置置信水平,此处以常用的95%置信区间为例,可按需调整
confidence = 0.95
# 取对应分位数作为区间上下界
ci_lower = np.percentile(all_subtree_predictions, (1 - confidence)/2 * 100, axis=0)
ci_upper = np.percentile(all_subtree_predictions, (1 + confidence)/2 * 100, axis=0)
# 模型点预测结果(即所有子树预测值的均值,和原reg.predict输出一致)
y_pred = reg.predict(xtest)

# 验证区间可靠性,输出真实值落在置信区间内的比例
actual_coverage = np.mean((ytest >= ci_lower) & (ytest <= ci_upper))
print(f"训练集R2: {r2_score(ytrain, reg.predict(xtrain)):.3f}")
print(f"测试集R2: {r2_score(ytest, y_pred):.3f}")
print(f"测试集MSE: {metrics.mean_squared_error(ytest, y_pred):.3f}")
print(f"{confidence*100}%置信区间实际覆盖率: {actual_coverage:.3f}")

# 可视化部分
plt.figure(figsize=(12, 6))
# 按预测值排序样本,避免填充区间时线条杂乱
sort_index = np.argsort(y_pred)
sorted_y_true = ytest.to_numpy()[sort_index]
sorted_y_pred = y_pred[sort_index]
sorted_ci_lower = ci_lower[sort_index]
sorted_ci_upper = ci_upper[sort_index]
x_axis = range(len(sorted_y_pred))

# 绘制真实值散点
plt.scatter(x_axis, sorted_y_true, s=12, c="gray", alpha=0.4, label="真实房价")
# 绘制预测值折线
plt.plot(x_axis, sorted_y_pred, c="#2563eb", lw=2, label="模型预测值")
# 填充置信区间
plt.fill_between(x_axis, sorted_ci_lower, sorted_ci_upper, color="#2563eb", alpha=0.2, label=f"{confidence*100}%置信区间")

plt.xlabel("测试集样本(按预测值从小到大排序)")
plt.ylabel("房屋价格")
plt.legend()
plt.title("房价预测结果与置信区间展示")
plt.tight_layout()
plt.show()

注意事项

  • 单棵决策树的predict方法不兼容pandas DataFrame格式输入,因此传入特征时需要用.to_numpy()转为numpy数组,避免特征名匹配警告
  • 如果你需要更严谨的分位数置信区间,可直接使用scikit-learn 1.0以上版本内置的HistGradientBoostingRegressor,通过设置loss="quantile"分别训练对应分位数的上下界模型,不需要依赖长期未更新维护的第三方扩展库
  • 置信水平参数可按需调整,例如需要90%置信区间时将confidence设为0.9即可
  • 该方法计算效率远高于第三方扩展库的重抽样方案,在当前房价数据集上运行无性能压力

内容的提问来源于stack exchange,提问作者user17224304

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 01:31:19