You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按组检查其他组作者是否存在于本组Citations列并统计次数

解决方案(R实现)

步骤1:准备示例数据集

先模拟一个匹配需求的数据集,方便后续演示:

library(dplyr)
library(stringr)

# 构造示例数据
df <- tibble(
  Group = c("A", "A", "B", "B", "C"),
  Citations = c(
    "引用了Smith, Johnson",
    "引用了Williams, Smith",
    "引用了Brown, Johnson",
    "引用了Davis, Williams",
    "引用了Miller, Smith"
  ),
  Authors = c("Smith", "Johnson", "Williams", "Brown", "Miller")
)

步骤2:提取各分组的作者集合

先把每个Group对应的所有Authors提取出来,形成分组-作者的映射表:

group_authors <- df %>%
  distinct(Group, Authors) %>%
  group_split(Group) %>%
  set_names(sapply(., function(x) unique(x$Group)))

步骤3:对每个组统计外部作者的引用情况

遍历每个分组,排除自身组的作者,逐个检查这些作者在当前组所有Citations中的出现次数:

result_list <- lapply(names(group_authors), function(current_group) {
  # 合并当前组的所有Citations文本
  current_citations <- df %>%
    filter(Group == current_group) %>%
    pull(Citations) %>%
    str_c(collapse = " ")
  
  # 获取其他组的所有作者
  other_authors <- bind_rows(group_authors[names(group_authors) != current_group]) %>%
    pull(Authors)
  
  # 统计每个外部作者的出现次数
  occurrences <- str_count(current_citations, fixed(other_authors))
  
  # 整理成结果数据框
  tibble(
    Group = current_group,
    Author = other_authors,
    Occurrences = occurrences,
    Present = as.integer(occurrences > 0)
  )
})

# 合并所有结果
final_df <- bind_rows(result_list)

关键说明

  • 用fixed()包裹作者名,避免正则特殊字符干扰匹配(比如名字里的空格、点号)
  • str_count()直接统计每个作者在当前组Citations中的总出现次数,替代你之前尝试的colSum
  • as.integer(occurrences > 0)快速生成存在标记(1表示存在,0表示不存在)

运行后final_df就是你需要的结构:每个Group对应其他组的所有作者,附带出现次数和存在状态。

内容的提问来源于stack exchange,提问作者Clara HL

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 02:41:26