You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言实现:统计组合数据框二进制组合在基准数据框中的匹配行数及RatingBinary分组计数

我来帮你搞定这个需求,下面是两种R语言的实现方案(tidyverse风格和基础R风格),都能完美完成你的统计要求:

第一步:先定义你的数据框

首先把你提供的两个数据框准确还原出来:

# 组合数据框
combinations_df <- structure(list(
  `Q1` = c(0, 0, 0, 1, 0, 0, 0, 0, 0, 0), 
  `Q2` = c(0, 0, 0, 0, 0, 0, 0, 0, 0, 0), 
  `Q3` = c(0, 1, 0, 0, 0, 1, 1, 0, 0, 0), 
  `Q4` = c(0, 0, 0, 0, 0, 0, 0, 0, 0, 0), 
  `Q5` = c(0, 0, 1, 0, 0, 1, 0, 1, 1, 0), 
  `Q6` = c(1, 1, 0, 1, 1, 0, 0, 1, 1, 1), 
  `Q7` = c(0, 0, 0, 0, 0, 0, 0, 0, 0, 0), 
  `Q8` = c(0, 0, 0, 0, 0, 0, 0, 0, 0, 1), 
  `Q9` = c(1, 0, 1, 0, 0, 1, 1, 0, 1, 0), 
  `Q10` = c(0, 0, 0, 0, 0, 0, 0, 0, 0, 0), 
  `Q11` = c(0, 0, 0, 0, 1, 0, 0, 0, 0, 0), 
  `Q12` = c(1, 1, 1, 1, 1, 0, 1, 1, 0, 1)
), row.names = c(NA, -10L), class = "data.frame")

# 基准数据框
benchmark_df <- structure(list(
  Q1 = c(0, 0, 0, 0, 0, 1, 0, 0, 0, 1), 
  Q2 = c(0, 1, 1, 0, 0, 0, 0, 0, 0, 0), 
  Q3 = c(1, 0, 0, 1, 0, 0, 0, 0, 0, 0), 
  Q4 = c(0, 0, 0, 0, 0, 0, 0, 0, 0, 0), 
  Q5 = c(1, 0, 1, 0, 0, 0, 1, 0, 0, 1), 
  Q6 = c(1, 1, 1, 0, 1, 0, 0, 1, 0, 1), 
  Q7 = c(0, 0, 1, 1, 1, 0, 0, 0, 0, 0), 
  Q8 = c(1, 0, 1, 0, 0, 1, 0, 0, 0, 0), 
  Q9 = c(1, 0, 0, 0, 0, 0, 0, 1, 1, 0), 
  Q10 = c(0, 0, 1, 0, 0, 1, 0, 0, 0, 0), 
  Q11 = c(0, 0, 1, 0, 0, 1, 0, 0, 0, 0), 
  Q12 = c(1, 0, 0, 0, 1, 0, 1, 0, 0, 0), 
  RatingBinary = c(1, 1, 0, 1, 0, 1, 0, 1, 1, 1)
), row.names = c(NA, 10L), class = "data.frame")

方案一:用tidyverse包(更直观易读)

如果你习惯用tidyverse生态的工具,这个方案更清晰:

library(tidyverse)

# 定义处理单一行组合的函数
process_combination <- function(row) {
  # 提取当前行中值为1的列名
  selected_cols <- names(row)[row == 1]
  
  # 处理全0的特殊情况:此时所有基准行都符合条件
  if (length(selected_cols) == 0) {
    matched_rows <- benchmark_df
  } else {
    # 筛选基准数据中,所有选中列都为1的行
    matched_rows <- benchmark_df %>%
      filter(across(all_of(selected_cols), ~ .x == 1))
  }
  
  # 统计需要的数值
  total_matches <- nrow(matched_rows)
  count_0 <- sum(matched_rows$RatingBinary == 0)
  count_1 <- sum(matched_rows$RatingBinary == 1)
  
  # 返回结果作为一个小数据框
  tibble(total_matches, count_0, count_1)
}

# 逐行处理组合数据框,合并结果
result <- combinations_df %>%
  rowwise() %>%
  do(process_combination(.)) %>%
  ungroup()

# 查看最终统计结果
print(result)

方案二:基础R实现(无需额外包)

如果你不想加载第三方包,用基础R也能实现:

# 基础R版本的处理函数
process_combination_base <- function(row) {
  selected_cols <- names(row)[row == 1]
  
  if (length(selected_cols) == 0) {
    matched_rows <- benchmark_df
  } else {
    # 构建筛选条件:检查每行的选中列是否全为1
    filter_condition <- apply(benchmark_df[, selected_cols], 1, function(x) all(x == 1))
    matched_rows <- benchmark_df[filter_condition, ]
  }
  
  total_matches <- nrow(matched_rows)
  count_0 <- sum(matched_rows$RatingBinary == 0, na.rm = TRUE)
  count_1 <- sum(matched_rows$RatingBinary == 1, na.rm = TRUE)
  
  data.frame(total_matches, count_0, count_1)
}

# 应用函数并合并结果
result_base <- do.call(rbind, apply(combinations_df, 1, process_combination_base))

# 查看结果
print(result_base)

代码逻辑说明

不管哪种方案,核心逻辑都是一致的:

  1. 对组合数据框的每一行,先提取出值为1的列,这就是我们要匹配的二进制特征组合。
  2. 在基准数据框中筛选出所有这些列的值都为1的行。
  3. 对筛选出来的匹配行,统计总行数,以及其中RatingBinary为0和1的数量。
  4. 把每一行的统计结果合并成一个最终的结果数据框,方便你查看和后续分析。

比如组合数据框的第一行,选中的列是Q6、Q9、Q12,基准数据中只有第1行满足这三列全为1,所以统计结果是total_matches=1, count_0=0, count_1=1,和实际情况完全一致。

内容的提问来源于stack exchange,提问作者Nevedha Ayyanar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 03:57:44