You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

构建二分类数据学习曲线时遇ID未找到错误求助

解决逻辑回归曲线绘制报错的方案

问题根源

你在绘制时,geom_point调用了prediction_data中不存在的ID列——因为prediction_data是用于生成预测概率的数据集,通常只包含Measurement序列和对应的预测概率值,没有原始数据的ID字段。

修正步骤

1. 正确生成prediction_data

生成预测数据集时,仅构建覆盖原始Measurement范围的序列即可,无需包含ID:

# 基于原始数据的Measurement范围生成连续序列
prediction_data <- data.frame(
  Measurement = seq(min(df$Measurement), max(df$Measurement), length.out = 100)
)
# 预测成功概率并加入数据集
prediction_data$pred_prob <- predict(glm_model, newdata = prediction_data, type = "response")

2. 拆分图层的数据来源

绘制时,原始散点(带ID分组)用原始数据集,逻辑回归曲线用prediction_data,避免跨数据集引用不存在的字段:

library(ggplot2)

ggplot() +
  # 原始数据的Success散点(使用原始数据集df,保留ID分组)
  geom_point(data = df, aes(x = Measurement, y = Success, color = factor(ID))) +
  # 逻辑回归曲线(使用prediction_data,仅调用存在的字段)
  geom_line(data = prediction_data, aes(x = Measurement, y = pred_prob), color = "red", linewidth = 1) +
  labs(x = "Measurement", y = "Success Probability", color = "ID") +
  theme_bw()

3. 可选:叠加原始Value折线图

如果需要同时保留初始的Value随Measurement的分组折线,可叠加图层:

ggplot() +
  # 原始Value分组折线
  geom_line(data = df, aes(x = Measurement, y = Value, color = factor(ID))) +
  # 原始Success散点
  geom_point(data = df, aes(x = Measurement, y = Success, shape = factor(Success))) +
  # 逻辑回归曲线
  geom_line(data = prediction_data, aes(x = Measurement, y = pred_prob), color = "darkblue", linewidth = 1.2) +
  labs(x = "Measurement", y = "Value / Success Probability", color = "ID", shape = "Success") +
  theme_bw()

关键提醒

  • 每个geom_*图层可单独指定data参数,确保不同图层使用对应的数据,避免字段缺失报错。
  • predict函数设置type = "response"会输出0-1区间的概率值,适配逻辑回归曲线的y轴需求。

内容的提问来源于stack exchange,提问作者Conda_suman07

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 11:22:07