如何按组检查其他组作者是否存在于本组Citations列并统计次数
解决方案(R实现)
步骤1:准备示例数据集
先模拟一个匹配需求的数据集,方便后续演示:
library(dplyr) library(stringr) # 构造示例数据 df <- tibble( Group = c("A", "A", "B", "B", "C"), Citations = c( "引用了Smith, Johnson", "引用了Williams, Smith", "引用了Brown, Johnson", "引用了Davis, Williams", "引用了Miller, Smith" ), Authors = c("Smith", "Johnson", "Williams", "Brown", "Miller") )
步骤2:提取各分组的作者集合
先把每个Group对应的所有Authors提取出来,形成分组-作者的映射表:
group_authors <- df %>% distinct(Group, Authors) %>% group_split(Group) %>% set_names(sapply(., function(x) unique(x$Group)))
步骤3:对每个组统计外部作者的引用情况
遍历每个分组,排除自身组的作者,逐个检查这些作者在当前组所有Citations中的出现次数:
result_list <- lapply(names(group_authors), function(current_group) { # 合并当前组的所有Citations文本 current_citations <- df %>% filter(Group == current_group) %>% pull(Citations) %>% str_c(collapse = " ") # 获取其他组的所有作者 other_authors <- bind_rows(group_authors[names(group_authors) != current_group]) %>% pull(Authors) # 统计每个外部作者的出现次数 occurrences <- str_count(current_citations, fixed(other_authors)) # 整理成结果数据框 tibble( Group = current_group, Author = other_authors, Occurrences = occurrences, Present = as.integer(occurrences > 0) ) }) # 合并所有结果 final_df <- bind_rows(result_list)
关键说明
- 用
fixed()包裹作者名,避免正则特殊字符干扰匹配(比如名字里的空格、点号) str_count()直接统计每个作者在当前组Citations中的总出现次数,替代你之前尝试的colSumas.integer(occurrences > 0)快速生成存在标记(1表示存在,0表示不存在)
运行后final_df就是你需要的结构:每个Group对应其他组的所有作者,附带出现次数和存在状态。
内容的提问来源于stack exchange,提问作者Clara HL
相关产品推荐
相关产品推荐

