You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中绘制多个分类器的提升曲线(Lift Curves)

嘿,作为R语言新手碰到这种可视化问题太正常了,我来帮你搞定在同一张图里合并多个分类器提升曲线的事儿!

首先先确认几个前提:你的Response应该是二分类变量对吧?提升曲线主要用于二分类模型的效果对比,另外建议用测试集来生成提升曲线,别用训练集,不然结果会偏乐观哦。

步骤1:安装并加载必要的包

我推荐用lift包(专门做提升曲线的)加上ggplot2来绘图,先搞定包的安装:

install.packages(c("lift", "ggplot2"))
library(lift)
library(ggplot2)

步骤2:拟合模型并生成预测值

先把你的GLM模型补全(记得加family=binomial,因为是二分类问题):

# 补全你的GLM模型
fullmod <- glm(Response ~ page_views_90d+win_visits+osx_visits+mc_1+mc_2+mc_3+mc_4+mc_5+mc_6+store_page+orders+orderlines+bookings+purchase, 
               data=train, 
               family=binomial)

# 假设你有第二个"本质相同"的模型(比如不同参数/数据集训练的)
# 这里我用另一个示例模型,你替换成自己的即可
other_mod <- glm(Response ~ page_views_90d+win_visits+store_page+orders+purchase, 
                 data=train, 
                 family=binomial)

# 用测试集生成预测概率(没有测试集的话也可以用训练集,但不推荐)
pred_full <- predict(fullmod, newdata = test, type = "response")
pred_other <- predict(other_mod, newdata = test, type = "response")

步骤3:生成提升曲线数据并合并绘图

接下来我们分别生成两个模型的提升曲线数据,然后合并到同一张图里:

# 生成第一个模型的提升数据
lift_full <- lift(Response ~ pred_full, data = test, class = "1") # class参数填你的正类标签,比如"1"或者"yes"
# 生成第二个模型的提升数据
lift_other <- lift(Response ~ pred_other, data = test, class = "1")

# 给每个数据集加上模型标识,方便区分
lift_data_full <- cbind(lift_full$data, Model = "Full GLM Model")
lift_data_other <- cbind(lift_other$data, Model = "Simplified GLM Model")

# 合并两个数据集
combined_lift_data <- rbind(lift_data_full, lift_data_other)

# 绘制合并后的提升曲线
ggplot(combined_lift_data, aes(x = Cume.Pct.of.Population, y = Cume.Pct.of.Response, color = Model)) +
  geom_line(size = 1) + # 画曲线
  geom_abline(intercept = 0, slope = 1, linetype = "dashed", color = "gray") + # 加基准线(随机模型)
  labs(title = "Lift Curve Comparison of Two Classifiers",
       x = "Cumulative Percentage of Population",
       y = "Cumulative Percentage of Positive Responses",
       color = "Classifier") +
  theme_minimal()

关于你提到的"两个分类器本质相同但图表不同"的小提醒

如果两个模型本质相同但图表不一样,大概率是因为:

  • 一个用了训练集、一个用了测试集(训练集的提升曲线会更漂亮)
  • 生成曲线时的分组数或者参数设置不同
  • 模型拟合时的随机种子不一样(如果涉及到抽样的话)

你可以检查下这些点,确保两个模型的评估条件一致哦~

内容的提问来源于stack exchange,提问作者paddy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 11:17:41