如何用ggplot2生成以计数而非比例为Y轴的逆ECDF图
解决方案
方法1:直接调整stat_ecdf的Y轴
通过将比例乘以总样本数,把Y轴从比例转换为计数。以USArrests为例:
library(tidyverse) USArrests %>% ggplot(aes(x = Assault)) + stat_ecdf(aes(y = (1 - ..y..) * nrow(USArrests)), geom = "step") + labs(x = "袭击事件数", y = "至少发生该数量袭击事件的州数", title = "州袭击事件累计计数分布") + theme_minimal()
此方法无需额外数据预处理,直接在stat_ecdf中计算计数,X=200对应的Y值即为19,符合需求。
方法2:手动计算累计计数(更直观可控)
先统计每个个体的治疗次数,再计算每个次数对应的累计患者数:
# 处理USArrests模拟数据 arrest_summary <- USArrests %>% mutate(state = rownames(.)) %>% # 州名作为"患者标识" select(state, n_treatments = Assault) %>% arrange(n_treatments) %>% mutate(cum_patients = nrow(.) - row_number() + 1) # 计算至少该次数的州数 # 绘图 ggplot(arrest_summary, aes(x = n_treatments, y = cum_patients)) + geom_step() + labs(x = "袭击事件数", y = "至少发生该数量袭击事件的州数", title = "州袭击事件累计计数分布") + theme_minimal()
迁移到你的treatment数据集
将逻辑应用到真实治疗数据:
# 统计每位患者的治疗次数 patient_summary <- treatment %>% group_by(patKey) %>% summarise(n_treatments = n()) %>% arrange(n_treatments) %>% mutate(cum_patients = nrow(.) - row_number() + 1) # 生成目标图 ggplot(patient_summary, aes(x = n_treatments, y = cum_patients)) + geom_step() + labs(x = "治疗次数", y = "至少接受该次数治疗的患者数", title = "患者治疗次数累计计数分布") + theme_minimal()
此时查看X=10对应的Y值,就是满足至少10次治疗的患者总数。
内容的提问来源于stack exchange,提问作者Mark
相关产品推荐
相关产品推荐

