R语言含NA值时的group_by优化:计算基准值与变化率
简洁实现分组基准值提取与百分比变化计算
先给出原始数据:
df <- data.frame( id = c(1, 1, 1, 1, 2, 2, 2, 2, 3, 3), result = c(NA, 33, 13, 44, 23, 44, 52, 11, NA, NA), flag = c("", "", "Y", "", "Y", "", "", "", "", ""), col_1 = c("a", "b", "c", "d" , "e" , "f" , "g" , "h", "i", "j") )
方案一:使用dplyr(tidyverse)
利用group_by()分组后,直接提取每组中flag == "Y"对应的result作为基准值,再计算百分比变化(假设每个id最多对应一个flag="Y"的行):
library(dplyr) df_processed <- df %>% group_by(id) %>% mutate( base_value = result[flag == "Y"][1], percentage_change = ifelse(!is.na(result) & !is.na(base_value), (result - base_value)/base_value * 100, NA) ) %>% ungroup() print(df_processed)
说明:
result[flag == "Y"][1]:若组内无flag="Y"的行(如id=3),会返回NA;若有多行也只会取第一个,适配多数场景。- 百分比计算时增加NA判断,避免无效值引发的计算错误。
方案二:使用data.table(大数据场景优先)
如果处理大规模数据集,data.table的语法更高效简洁:
library(data.table) setDT(df) df_processed <- df[, `:=`( base_value = result[flag == "Y"][1], percentage_change = fcase( !is.na(result) & !is.na(result[flag == "Y"][1]), (result - result[flag == "Y"][1])/result[flag == "Y"][1] * 100, default = NA_real_ ) ), by = id] print(df_processed)
处理后结果示例
# A tibble: 10 × 6 id result flag col_1 base_value percentage_change <dbl> <dbl> <chr> <chr> <dbl> <dbl> 1 1 NA "" a 13 NA 2 1 33 "" b 13 153. 3 1 13 "Y" c 13 0 4 1 44 "" d 13 238. 5 2 23 "Y" e 23 0 6 2 44 "" f 23 91.3 7 2 52 "" g 23 126. 8 2 11 "" h 23 -52.2 9 3 NA "" i NA NA 10 3 NA "" j NA NA
内容的提问来源于stack exchange,提问作者Joe the Second
相关产品推荐
相关产品推荐

