R语言:用aggregate()替代嵌套for()循环配合cut()实现分组分箱标记
分组分位数分箱的实现方法
问题描述
我希望将带有多个属性标记的数据划分至分位数区间,且分位数仅对具备对应属性的数据生效。目前通过嵌套for()循环已实现双属性分组的分箱功能,代码如下(已修正原代码的语法错误):
df <- data.frame( type1 = rep(c('A','B'), each=50), type2 = rep(c('D','E','F','G'), each=25), f1 = rnorm(100, mean=0.0, sd=1.0), f2 = rnorm(100, mean=0.0, sd=1.0) ) summary(df) id_breaks <- c(-2.0,-1.0,0,1.0,2.0) q_df <- NULL for (t1 in unique(df$type1)) { for (t2 in unique(df$type2)) { df_sub <- df[df$type1 == t1 & df$type2 == t2,] this_q_df <- within(df_sub, f1_quart <- cut(f1, id_breaks, include.lowest=FALSE, labels=FALSE)) print(paste("t1", t1, "t2", t2)) print(head(this_q_df)) # 修正原代码:将rbind结果赋值回q_df q_df <- rbind(q_df, this_q_df) } }
请问是否可以通过
aggregate(df$f1~t1+t2,...)的方式实现该功能?
解决方案
可以用aggregate结合自定义函数实现
aggregate支持按多分组变量处理数据,但要保留原数据的所有行,结合ave函数会比直接用aggregate后合并更高效:
- 先定义分箱的自定义函数:
bin_func <- function(x, breaks) { cut(x, breaks, include.lowest=FALSE, labels=FALSE) }
- 用
ave按type1和type2分组,对每个f1值应用分箱函数,直接在原数据中生成分箱列:
df$f1_quart <- ave(df$f1, df$type1, df$type2, FUN=function(x) bin_func(x, id_breaks))
如果一定要用aggregate,可以先生成分组的分箱结果,再合并回原数据,步骤稍繁琐:
# 生成每组的分箱结果(每组对应一个向量) bin_groups <- aggregate(f1 ~ type1 + type2, data=df, FUN=function(x) bin_func(x, id_breaks)) # 展开分组结果并与原数据合并 df <- merge(df, bin_groups, by=c("type1", "type2")) names(df)[names(df) == "f1.y"] <- "f1_quart" df <- df[, c("type1", "type2", "f1.x", "f2", "f1_quart")] names(df)[names(df) == "f1.x"] <- "f1"
更简洁的tidyverse实现(可选)
如果使用dplyr包,代码会更直观易读:
library(dplyr) df <- df %>% group_by(type1, type2) %>% mutate(f1_quart = cut(f1, id_breaks, include.lowest=FALSE, labels=FALSE)) %>% ungroup()
内容的提问来源于stack exchange,提问作者Bill Pearson
相关产品推荐
相关产品推荐

