R语言统计两个dataframe对应列行匹配的出现次数
问题根源
- 你代码中使用
any()仅会判断是否存在匹配,返回逻辑值,转成数值后只能得到1或0,自然无法统计匹配次数 - 在
apply的匿名函数内修改全局design$check属于错误用法,apply逐行迭代时的赋值逻辑完全不符合计数需求 - 按
x[1]、x[2]这种位置取列容错率极低,若design列顺序发生变化会直接导致匹配错误
解决方案
方案1:修改原有base R代码
直接将any()替换为sum()即可统计匹配行数,同时调整为按列名取值提升稳定性:
# 指定需要匹配的列名 match_cols <- c("undIssue", "feelConf", "setup", "undContex", "undChang") design$check <- apply(design[match_cols], 1, function(x) { sum( x[["undIssue"]] == actconjoint$undIssue & x[["feelConf"]] == actconjoint$feelConf & x[["setup"]] == actconjoint$setup & x[["undContex"]] == actconjoint$undContex & x[["undChang"]] == actconjoint$undChang ) })
方案2:用dplyr实现(更适合大数据量,运行效率更高)
先统计actconjoint中每个列组合的出现次数,再通过左连接赋值给design:
library(dplyr) match_cols <- c("undIssue", "feelConf", "setup", "undContex", "undChang") # 统计actconjoint各组合的出现次数 act_cnt <- actconjoint %>% count(across(all_of(match_cols)), name = "check") # 连接到design,无匹配的行默认填充0 design <- design %>% left_join(act_cnt, by = match_cols) %>% mutate(check = coalesce(check, 0))
内容的提问来源于stack exchange,提问作者Fabio Marcos Santos
相关产品推荐
相关产品推荐

