You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言tree包predict函数subset参数作用及使用疑问

问题解析:tree包predict函数subset参数的误用

为什么两段代码结果不同?

tree包的predict.tree()函数里的subset参数不是用来筛选newdata的行的,它的作用是指定使用树模型中的哪些节点对应的预测规则(比如针对剪枝后的树节点),和传入的newdata数据集的行筛选完全无关。

  • 第一段代码直接传入newdata = Credit[-train,],相当于只给模型喂测试集的200行数据,自然返回200个预测结果。
  • 第二段代码传入newdata = Credit(全量400行数据),同时加了subset=-train,但这个参数不会对newdata做行筛选,模型依然会对全量400行做预测,所以返回400个结果。你误以为subset能筛选newdata,这是对参数功能的误解。

正确获取测试集预测结果的方式

如果想获取测试集的预测结果,有两种靠谱的方式:

  1. 像第一段代码那样,直接传入测试集作为newdata:
tree.pred <- predict(object = tree.Credit, newdata = Credit[-train,])
length(tree.pred) # 200
  1. 先对全量数据预测,再提取测试集对应的结果:
tree.pred_full <- predict(object = tree.Credit, newdata = Credit)
tree.pred <- tree.pred_full[-train]
length(tree.pred) # 200

tree包的predict函数没有提供通过subset参数筛选newdata的功能,别把它和其他函数(比如lm的predict)的subset用法混淆了。

内容的提问来源于stack exchange,提问作者Antonio Piemontese

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 10:53:16