R语言实现多列多条件(含部分字符串匹配)的行计数与结果聚合
方案1:tidyverse 实现(语法简洁易读)
依赖dplyr和stringr包,适合日常分析场景:
# 加载依赖 library(dplyr) library(stringr) # 你的模拟数据(加种子方便结果复现) set.seed(123) x1<- sample(c("t1xy", "t2xy", "m1xy", "m2xy","t1yx", "t2yx", "m1yx", "m2yx"), 20, replace = T) x2<- sample(1:4, 20, replace = T) x3<- sample(0:1, 20, replace = T) df_x <- data.frame(x1,x2,x3) # 核心统计逻辑 stat_res <- df_x %>% # 提取x1中匹配的前缀t1/t2/m1/m2 mutate(prefix = str_extract(x1, "^(t1|t2|m1|m2)")) %>% # 按前缀、x2、x3分组统计行数 count(prefix, x2, x3, name = "cnt") # 生成你需要的字符串格式输出 stat_res %>% mutate(output = sprintf("(%s, x2=%d, x3=%d) = %d", prefix, x2, x3, cnt)) %>% pull(output) %>% print(quote = FALSE) # 如果需要按前缀拆分存储为独立对象,可直接拆分到列表: stat_by_group <- split(stat_res, stat_res$prefix) # 调用示例:stat_by_group[["t1"]] 即可取出t1对应的所有组合统计
方案2:基础R实现(无需第三方依赖)
适合不能安装加载包的场景:
# 模拟数据 set.seed(123) x1<- sample(c("t1xy", "t2xy", "m1xy", "m2xy","t1yx", "t2yx", "m1yx", "m2yx"), 20, replace = T) x2<- sample(1:4, 20, replace = T) x3<- sample(0:1, 20, replace = T) df_x <- data.frame(x1,x2,x3) # 提取前缀 df_x$prefix <- substr(df_x$x1, 1, 2) # 分组统计 stat_res <- aggregate(.~prefix+x2+x3, transform(df_x, n=1), length)[,c("prefix","x2","x3","n")] colnames(stat_res)[4] <- "cnt" # 生成字符串输出 stat_res$output <- sprintf("(%s, x2=%d, x3=%d) = %d", stat_res$prefix, stat_res$x2, stat_res$x3, stat_res$cnt) print(stat_res$output, quote = FALSE) # 按前缀拆分 stat_by_group <- split(stat_res, stat_res$prefix)
两种方案均会输出所有前缀和x2、x3的组合统计,用你给出的示例数据测试时,t2, x2=4, x3=0对应的统计值为3,符合预期。
内容的提问来源于stack exchange,提问作者Taiku
相关产品推荐
相关产品推荐

