You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R语言中按列统计指定词汇的出现频率

R语言实现词汇提取与指定词频统计

1. 加载必要工具包

处理这类文本统计任务,tidyverse用于数据操作,tidytext专门负责文本拆分与词汇处理:

# 首次使用需安装包
install.packages(c("tidyverse", "tidytext"))
library(tidyverse)
library(tidytext)

2. 提取全量词汇列表

假设你的原始数据框名为raw_data,包含2000、2001、2002、2003这几列文本数据。先把数据转成长格式统一处理,再拆分出所有词汇:

# 仅保留目标年份列并转为长格式
long_format_data <- raw_data %>%
  select(2000, 2001, 2002, 2003) %>%
  pivot_longer(cols = everything(), names_to = "year", values_to = "text")

# 拆分文本为单个词汇,生成纯词汇列表
all_words <- long_format_data %>%
  unnest_tokens(word, text, token = "words", strip_punct = TRUE) %>%
  pull(word)

3. 统计指定词汇的出现频率

针对ich、möchte、doner、und、ayran这几个目标词,按年份统计出现次数与频率:

target_words <- c("ich", "möchte", "doner", "und", "ayran")

frequency_result <- long_format_data %>%
  unnest_tokens(word, text, token = "words", strip_punct = TRUE) %>%
  filter(word %in% target_words) %>%
  count(year, word, name = "occurrence") %>%
  group_by(year) %>%
  mutate(
    total_word_count = nrow(unnest_tokens(long_format_data %>% filter(year == cur_group()$year), word, text)),
    frequency = occurrence / total_word_count
  ) %>%
  ungroup()

# 输出统计结果
print(frequency_result)

补充说明

  • 若文本包含德语变音等特殊字符,确保数据编码为UTF-8避免乱码;
  • 代码中total_word_count统计的是对应年份的总词汇出现次数(含重复),如果需要统计去重后的词汇总数,可替换为n_distinct(word);
  • strip_punct = TRUE参数用于自动去除文本中的标点符号,若不需要可删除该参数。

内容的提问来源于stack exchange,提问作者Meryem Çiftçi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 14:40:34