R语言实现:统计组合数据框二进制组合在基准数据框中的匹配行数及RatingBinary分组计数
我来帮你搞定这个需求,下面是两种R语言的实现方案(tidyverse风格和基础R风格),都能完美完成你的统计要求:
第一步:先定义你的数据框
首先把你提供的两个数据框准确还原出来:
# 组合数据框 combinations_df <- structure(list( `Q1` = c(0, 0, 0, 1, 0, 0, 0, 0, 0, 0), `Q2` = c(0, 0, 0, 0, 0, 0, 0, 0, 0, 0), `Q3` = c(0, 1, 0, 0, 0, 1, 1, 0, 0, 0), `Q4` = c(0, 0, 0, 0, 0, 0, 0, 0, 0, 0), `Q5` = c(0, 0, 1, 0, 0, 1, 0, 1, 1, 0), `Q6` = c(1, 1, 0, 1, 1, 0, 0, 1, 1, 1), `Q7` = c(0, 0, 0, 0, 0, 0, 0, 0, 0, 0), `Q8` = c(0, 0, 0, 0, 0, 0, 0, 0, 0, 1), `Q9` = c(1, 0, 1, 0, 0, 1, 1, 0, 1, 0), `Q10` = c(0, 0, 0, 0, 0, 0, 0, 0, 0, 0), `Q11` = c(0, 0, 0, 0, 1, 0, 0, 0, 0, 0), `Q12` = c(1, 1, 1, 1, 1, 0, 1, 1, 0, 1) ), row.names = c(NA, -10L), class = "data.frame") # 基准数据框 benchmark_df <- structure(list( Q1 = c(0, 0, 0, 0, 0, 1, 0, 0, 0, 1), Q2 = c(0, 1, 1, 0, 0, 0, 0, 0, 0, 0), Q3 = c(1, 0, 0, 1, 0, 0, 0, 0, 0, 0), Q4 = c(0, 0, 0, 0, 0, 0, 0, 0, 0, 0), Q5 = c(1, 0, 1, 0, 0, 0, 1, 0, 0, 1), Q6 = c(1, 1, 1, 0, 1, 0, 0, 1, 0, 1), Q7 = c(0, 0, 1, 1, 1, 0, 0, 0, 0, 0), Q8 = c(1, 0, 1, 0, 0, 1, 0, 0, 0, 0), Q9 = c(1, 0, 0, 0, 0, 0, 0, 1, 1, 0), Q10 = c(0, 0, 1, 0, 0, 1, 0, 0, 0, 0), Q11 = c(0, 0, 1, 0, 0, 1, 0, 0, 0, 0), Q12 = c(1, 0, 0, 0, 1, 0, 1, 0, 0, 0), RatingBinary = c(1, 1, 0, 1, 0, 1, 0, 1, 1, 1) ), row.names = c(NA, 10L), class = "data.frame")
方案一:用tidyverse包(更直观易读)
如果你习惯用tidyverse生态的工具,这个方案更清晰:
library(tidyverse) # 定义处理单一行组合的函数 process_combination <- function(row) { # 提取当前行中值为1的列名 selected_cols <- names(row)[row == 1] # 处理全0的特殊情况:此时所有基准行都符合条件 if (length(selected_cols) == 0) { matched_rows <- benchmark_df } else { # 筛选基准数据中,所有选中列都为1的行 matched_rows <- benchmark_df %>% filter(across(all_of(selected_cols), ~ .x == 1)) } # 统计需要的数值 total_matches <- nrow(matched_rows) count_0 <- sum(matched_rows$RatingBinary == 0) count_1 <- sum(matched_rows$RatingBinary == 1) # 返回结果作为一个小数据框 tibble(total_matches, count_0, count_1) } # 逐行处理组合数据框,合并结果 result <- combinations_df %>% rowwise() %>% do(process_combination(.)) %>% ungroup() # 查看最终统计结果 print(result)
方案二:基础R实现(无需额外包)
如果你不想加载第三方包,用基础R也能实现:
# 基础R版本的处理函数 process_combination_base <- function(row) { selected_cols <- names(row)[row == 1] if (length(selected_cols) == 0) { matched_rows <- benchmark_df } else { # 构建筛选条件:检查每行的选中列是否全为1 filter_condition <- apply(benchmark_df[, selected_cols], 1, function(x) all(x == 1)) matched_rows <- benchmark_df[filter_condition, ] } total_matches <- nrow(matched_rows) count_0 <- sum(matched_rows$RatingBinary == 0, na.rm = TRUE) count_1 <- sum(matched_rows$RatingBinary == 1, na.rm = TRUE) data.frame(total_matches, count_0, count_1) } # 应用函数并合并结果 result_base <- do.call(rbind, apply(combinations_df, 1, process_combination_base)) # 查看结果 print(result_base)
代码逻辑说明
不管哪种方案,核心逻辑都是一致的:
- 对组合数据框的每一行,先提取出值为1的列,这就是我们要匹配的二进制特征组合。
- 在基准数据框中筛选出所有这些列的值都为1的行。
- 对筛选出来的匹配行,统计总行数,以及其中
RatingBinary为0和1的数量。 - 把每一行的统计结果合并成一个最终的结果数据框,方便你查看和后续分析。
比如组合数据框的第一行,选中的列是Q6、Q9、Q12,基准数据中只有第1行满足这三列全为1,所以统计结果是total_matches=1, count_0=0, count_1=1,和实际情况完全一致。
内容的提问来源于stack exchange,提问作者Nevedha Ayyanar
相关产品推荐
相关产品推荐

