You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

ggplot2二维核密度估计中level参数含义及概率相关性技术咨询

关于ggplot2中stat_density_2d的after_stat(level)参数解读

问题描述

我用ggplot2生成了多组二维核密度估计图,代码能正常运行,但不清楚stat_density_2d里after_stat(level)的level具体含义,它和概率有关吗?该怎么解读这个参数?

代码示例

m <- ggplot(EMOD29, aes(x=GPP, y=ER, shape=Model)) +
geom_point(aes(shape=Model, color=Model), size=3.5)+
scale_shape_manual(values=c(3,16)) +
scale_color_manual(values=c("Black","Black"))+
theme(legend.title = element_blank(),legend.text = element_text(size=15, "Times New Roman"), legend.position = "right", aspect.ratio = 1)+
xlim(0, 25) +
ylim(-10,5)+
guides(color = guide_legend(override.aes = list(size = 1)))

KDE.plot.m<- m + stat_density_2d(geom = "polygon", contour = TRUE,
              aes(fill = after_stat(level)), colour="transparent",
              bins = 30)+
geom_point() +
scale_fill_distiller(palette = "Blues", direction = 1) +
theme(text=element_text(family="Times New Roman", size=18),
    panel.grid = element_blank(),
    panel.background = element_blank(),
    axis.line = element_line(),
    legend.position = "right")

KDE.plot.m

参数解读

  • 核心含义:after_stat(level)中的level是二维核密度估计的等高线阈值,每条填充的多边形轮廓,对应着「样本点落在该轮廓内的累积概率等于某一固定值」,本质是概率密度积分后得到的累积概率对应的密度临界值。
  • 和概率的关联:它直接和累积概率挂钩。你代码里用了bins=30,ggplot会把总概率(100%)平均分成30个区间,每个level就对应一个区间的密度临界值。比如相邻两个level之间的区域,对应约3.33%(1/30)的样本概率占比。
  • 实际解读方式:
    • 你用了scale_fill_distiller(palette = "Blues", direction = 1),所以颜色越深的区域对应越高的level,代表样本点密度越高,随机抽取一个样本落在该区域的概率也越大。
    • 如果想要更直观的概率对应关系,可以不用bins,改用breaks手动指定level值,比如breaks=c(0.1, 0.5, 0.9),这样生成的轮廓就分别对应10%、50%、90%的累积概率区域,一眼就能看懂不同区域的概率占比。

内容的提问来源于stack exchange,提问作者Nicole Gotkowski

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 06:25:21