You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按年份统计语料库总词数?R语言quanteda求助

解决方案

tokens_by()确实在quanteda的新版本中已被弃用,现在可以通过分组操作结合quanteda核心函数实现按年份统计总词数,以下是两种实用方法:

方法一:基于分词结果分组统计

先对语料分词,再按Year分组统计每组总词数:

library(quanteda)

# 分词(可按需添加清洗规则,比如移除标点、数字)
SPARA_tokens <- tokens(SPARA_corp, remove_punct = TRUE, remove_numbers = TRUE)

# 按年份分组并统计总词数
word_count_by_year <- tokens_group(SPARA_tokens, groups = Year) %>%
  lengths() %>%
  as.data.frame() %>%
  rename(total_words = ".") %>%
  mutate(Year = rownames(.)) %>%
  relocate(Year, .before = total_words)

方法二:通过文档-特征矩阵(dfm)聚合统计

生成dfm后按年份聚合,再计算总词数,这种方法更灵活,方便后续扩展分析:

# 生成dfm(可添加移除停用词等清洗参数)
SPARA_dfm <- dfm(SPARA_corp, remove_punct = TRUE, remove_numbers = TRUE, remove_stopwords = TRUE)

# 按年份聚合并统计总词数
word_count_by_year <- dfm_group(SPARA_dfm, groups = Year) %>%
  rowSums() %>%
  as.data.frame() %>%
  rename(total_words = ".") %>%
  mutate(Year = rownames(.)) %>%
  relocate(Year, .before = total_words)

补充提示

  • 如果需要更精细的文本清洗(比如自定义停用词、移除特定字符),可以在tokens()或dfm()中添加对应参数,比如stopwords = "english"或自定义停用词向量。
  • 两种方法最终都会得到包含Year(年份)和total_words(当年总词数)的数据框,直接用于后续分析或可视化即可。

内容的提问来源于stack exchange,提问作者Peter

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 15:50:15