在R语言中按列统计指定词汇的出现频率
R语言实现词汇提取与指定词频统计
1. 加载必要工具包
处理这类文本统计任务,tidyverse用于数据操作,tidytext专门负责文本拆分与词汇处理:
# 首次使用需安装包 install.packages(c("tidyverse", "tidytext")) library(tidyverse) library(tidytext)
2. 提取全量词汇列表
假设你的原始数据框名为raw_data,包含2000、2001、2002、2003这几列文本数据。先把数据转成长格式统一处理,再拆分出所有词汇:
# 仅保留目标年份列并转为长格式 long_format_data <- raw_data %>% select(2000, 2001, 2002, 2003) %>% pivot_longer(cols = everything(), names_to = "year", values_to = "text") # 拆分文本为单个词汇,生成纯词汇列表 all_words <- long_format_data %>% unnest_tokens(word, text, token = "words", strip_punct = TRUE) %>% pull(word)
3. 统计指定词汇的出现频率
针对ich、möchte、doner、und、ayran这几个目标词,按年份统计出现次数与频率:
target_words <- c("ich", "möchte", "doner", "und", "ayran") frequency_result <- long_format_data %>% unnest_tokens(word, text, token = "words", strip_punct = TRUE) %>% filter(word %in% target_words) %>% count(year, word, name = "occurrence") %>% group_by(year) %>% mutate( total_word_count = nrow(unnest_tokens(long_format_data %>% filter(year == cur_group()$year), word, text)), frequency = occurrence / total_word_count ) %>% ungroup() # 输出统计结果 print(frequency_result)
补充说明
- 若文本包含德语变音等特殊字符,确保数据编码为
UTF-8避免乱码; - 代码中
total_word_count统计的是对应年份的总词汇出现次数(含重复),如果需要统计去重后的词汇总数,可替换为n_distinct(word); strip_punct = TRUE参数用于自动去除文本中的标点符号,若不需要可删除该参数。
内容的提问来源于stack exchange,提问作者Meryem Çiftçi
相关产品推荐
相关产品推荐

