You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言按分组统计数据框目标向量元素出现次数的实现方法

R语言实现方案

全程使用R基础函数实现,无需安装第三方包,支持两种输出格式。


实现步骤

  • 按Subject分组,筛选出属于class.interest的Class值,去重后统计匹配个数、拼接匹配到的类名字符串
  • 生成0到length(class.interest)的完整整数档位序列,保证所有档位都在结果中展示
  • 根据需要选择聚合输出(同档位合并Subject)或逐行输出(同档位每个Subject单独一行),填充无匹配档位的数值为0、文本为NA

完整代码

1. 初始化输入数据

class.interest <- c("a", "b", "c", "d", "e")
df <- data.frame(
  "Subject" = c(rep("A",3), rep("B",3), rep("C",5), "D", "E"),
  "Class" = c("a", "b", "f", "b", "b", "e", "a", "b", "c", "e", "f", "c", "f")
)

2. 分组统计每个Subject的匹配结果

sub_stat <- tapply(df$Class, df$Subject, \(x) {
  match_cls <- unique(x[x %in% class.interest])
  data.frame(
    match_num = length(match_cls),
    count.class = ifelse(length(match_cls) == 0, NA_character_, paste(match_cls, collapse = ","))
  )
}, simplify = FALSE)

sub_df <- do.call(rbind, lapply(names(sub_stat), \(sub) {
  cbind(Subject = sub, sub_stat[[sub]])
}))

3. 生成全档位基础表

full_occur <- data.frame(
  `# of occurrences` = 0:length(class.interest)
)

两种输出格式实现

格式1:同档位Subject合并为一行(初始期望格式)

# 按匹配档位聚合
occur_stat <- tapply(1:nrow(sub_df), sub_df$match_num, \(idx) {
  cur_data <- sub_df[idx, ]
  data.frame(
    count = nrow(cur_data),
    count.from.subject = paste(cur_data$Subject, collapse = ","),
    count.class = paste(cur_data$count.class, collapse = ";")
  )
}, simplify = FALSE)

occur_df <- do.call(rbind, lapply(names(occur_stat), \(num) {
  cbind(`# of occurrences` = as.integer(num), occur_stat[[num]])
}))

# 合并全档位,填充空值
final_res <- merge(full_occur, occur_df, by = "# of occurrences", all.x = TRUE)
final_res$count[is.na(final_res$count)] <- 0
final_res[final_res$count == 0, c("count.from.subject", "count.class")] <- NA_character_
final_res <- final_res[order(final_res$`# of occurrences`), ]
rownames(final_res) <- NULL

运行输出结果:

# of occurrences count count.from.subject count.class
1                0     1                  E        <NA>
2                1     1                  D           c
3                2     2                A,B     a,b;b,e
4                3     0               <NA>        <NA>
5                4     1                  C     a,b,c,e
6                5     0               <NA>        <NA>

格式2:同档位每个Subject单独一行(补充需求格式)

final_res2 <- merge(full_occur, sub_df, by.x = "# of occurrences", by.y = "match_num", all.x = TRUE)
final_res2$count <- ave(final_res2$Subject, final_res2$`# of occurrences`, FUN = \(x) length(na.omit(x)))
final_res2 <- final_res2[, c("# of occurrences", "count", "Subject", "count.class")]
colnames(final_res2)[3] <- "count.from.subject"
final_res2 <- final_res2[order(final_res2$`# of occurrences`), ]
rownames(final_res2) <- NULL

运行输出结果:

# of occurrences count count.from.subject count.class
1                0     1                  E        <NA>
2                1     1                  D           c
3                2     2                  A         a,b
4                2     2                  B         b,e
5                3     0               <NA>        <NA>
6                4     1                  C     a,b,c,e
7                5     0               <NA>        <NA>

内容的提问来源于stack exchange,提问作者Jen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 11:09:16