在R中基于另一列的因子层级变化创建新列
基于因子层级分组计算百分位变化值
我需要基于数据框中Percentile列的因子层级信息,按ID分组创建新列PercentileChange,用来标记相邻月份间百分位的升降及层级变化数。
原始数据框
df <- data.frame(ID = c("1","1","2","2"), Month = c("01", "02", "01", "02"), Percentile = c("P50", "P95", "P97", "P85")) df
输出:
ID Month Percentile 1 1 01 P50 2 1 02 P95 3 2 01 P97 4 2 02 P85
需求说明
创建名为PercentileChange的新列,规则如下:
- 按单个
ID分组 - 对比同一
ID下相邻月份的Percentile因子层级 - 首行标记为
NA,后续行标记升降(+表示上升,-表示下降,0表示无变化)及变化的层级数
期望输出:
df2
ID Month Percentile PercentileChange 1 1 01 P50 NA 2 1 02 P95 +4 3 2 01 P97 NA 4 2 02 P85 -3
因子层级定义
Percentile列的因子层级顺序固定为:
df$Percentile <- factor(df$Percentile,levels=c("P01","P1","P3","P5","P10","P15","P25","P50","P75","P85","P90","P95","P97","P99","P999"))
扩展数据框示例
为覆盖更多场景(多月份、无变化、跨非连续月份),补充扩展数据框:
df <- data.frame(ID = c("1","1","1","1","2","2","3","3","3"), Month = c("01", "02", "03", "04", "01", "02", "02", "03", "05"), Percentile = c("P50", "P95", "P97", "P85", "P01","P01", "P5","P5","P3")) df
输出:
ID Month Percentile 1 1 01 P50 2 1 02 P95 3 1 03 P97 4 1 04 P85 5 2 01 P01 6 2 02 P01 7 3 02 P5 8 3 03 P5 9 3 05 P3
对应的期望输出:
ID Month Percentile PercentileChange 1 1 01 P50 NA 2 1 02 P95 +4 3 1 03 P97 +1 4 1 04 P85 +3 5 2 01 P01 NA 6 2 02 P01 0 7 3 02 P5 NA 8 3 03 P5 0 9 3 05 P3 -1
解决方案
使用dplyr包实现分组计算,步骤如下:
- 先将
Percentile转换为指定层级的因子 - 提取因子的数值索引(对应层级位置)
- 按
ID分组,计算当前行与上一行的索引差值 - 将差值格式化为符合要求的字符串格式
代码实现:
library(dplyr) # 定义固定的百分位因子层级 percentile_levels <- c("P01","P1","P3","P5","P10","P15","P25","P50","P75","P85","P90","P95","P97","P99","P999") df_result <- df %>% # 转换为指定层级的因子 mutate(Percentile = factor(Percentile, levels = percentile_levels)) %>% # 按ID分组计算 group_by(ID) %>% mutate( # 提取因子对应的层级索引 pct_rank = as.integer(Percentile), # 计算与上一行的层级差值 change = pct_rank - lag(pct_rank), # 格式化为要求的字符串 PercentileChange = case_when( is.na(change) ~ NA_character_, change > 0 ~ paste0("+", change), change < 0 ~ as.character(change), TRUE ~ "0" ) ) %>% # 移除中间计算列 select(-pct_rank, -change) %>% ungroup() print(df_result)
内容的提问来源于stack exchange,提问作者user19779614
相关产品推荐
相关产品推荐

