You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于年份多条件使用dplyr::summarise生成年度论文计数列求助

问题:为关键词生成各年度论文计数列

我希望基于每年的论文数量为关键词创建一组列,需要在dplyr中通过summarise过滤多条件。现有代码如下:

words_list <- data %>% 
  select(Keywords, year) %>% 
  unnest_tokens(word, Keywords) %>% 
  filter(between(year,1990,2017)) %>% 
  group_by(word) %>% 
  summarise(papers_count = n()) %>% 
  arrange(desc(papers_count))

这段代码生成了word和papers_count两列,我需要创建更多类似papers_count的列(如papers_count1990、papers_count1991等),对应1990至2017年的每一年。期望代码形式如下:

words_list <- data %>% 
  select(Keywords, year) %>% 
  unnest_tokens(word, Keywords) %>% 
  filter(between(year,1990,2017)) %>% 
  group_by(word) %>% 
  summarise(tot_papers_count = n(), 
            papers_count_1991 = n()year="1991", ...) %>% 
  arrange(desc(papers_count))

恳请各位提供解决方案建议。


解决方案

这里有两种实用的方法帮你实现需求,不用手动编写每一年的重复代码,效率更高:

方法1:count + pivot_wider(推荐)

这是tidyverse里最简洁直观的方案,先按关键词和年份统计数量,再把行结构转成列结构,自动生成各年份的计数列:

library(tidyverse)

words_list <- data %>% 
  select(Keywords, year) %>% 
  unnest_tokens(word, Keywords) %>% 
  filter(between(year, 1990, 2017)) %>% 
  # 按关键词+年份统计每年的论文数
  count(word, year, name = "papers_count") %>% 
  # 将年份转换为列名,缺失年份填充0(可选)
  pivot_wider(
    names_from = year,
    values_from = papers_count,
    names_prefix = "papers_count_",
    values_fill = 0  # 若某关键词在某年份无论文,用0填充
  ) %>% 
  # 计算总论文数
  mutate(tot_papers_count = rowSums(across(starts_with("papers_count_")))) %>% 
  # 按总数量降序排列
  arrange(desc(tot_papers_count))

方法2:summarise + across动态生成列

如果你更习惯在分组汇总环节直接生成所有年份列,可以用across配合年份向量动态创建:

library(tidyverse)

# 先定义目标年份范围
target_years <- 1990:2017

words_list <- data %>% 
  select(Keywords, year) %>% 
  unnest_tokens(word, Keywords) %>% 
  filter(between(year, 1990, 2017)) %>% 
  group_by(word) %>% 
  summarise(
    tot_papers_count = n(),
    # 为每个年份生成对应的计数列
    across(all_of(target_years), ~sum(year == .x), .names = "papers_count_{.col}")
  ) %>% 
  ungroup() %>% 
  arrange(desc(tot_papers_count))

小提示:

  • 方法1的pivot_wider是tidyverse的标准操作,代码可读性强,后续调整年份范围只需要修改filter条件即可,维护成本低。
  • 方法2的across可以直接在summarise里完成所有列的生成,适合喜欢在汇总步骤一次性搞定所有计算的场景。
  • 两种方法都能避免手动编写几十行重复的年份判断代码,大大降低出错概率。

内容的提问来源于stack exchange,提问作者Amleto

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:35:12