R语言如何使用dplyr统计字符串中出现至少2次的单词
tidyverse 实现字符串重复单词识别
测试输入
input_str <- "the home home on the range the friend"
目标:识别字符串中出现次数≥2的单词,输出按单词在原字符串首次出现的顺序排列,最终返回c("the", "home"),优先使用stringr、dplyr实现。
方法1:stringr + dplyr 实现(无额外依赖)
library(tidyverse) result <- tibble(word = str_split(input_str, "\\s+")[[1]]) %>% count(word, name = "freq") %>% filter(freq >= 2) %>% mutate(first_pos = str_locate(input_str, paste0("\\b", word, "\\b"))[,1]) %>% arrange(first_pos) %>% pull(word)
代码逻辑说明
str_split按空白符拆分原字符串,得到独立的单词向量count统计每个单词的出现次数filter保留出现次数≥2的单词- 新增
first_pos列记录每个单词在原字符串中第一次出现的位置,按该列排序可保证输出顺序和要求一致,避免默认按字母排序导致home排在the前的问题 pull提取最终的单词列表
运行输出:
> print(result) [1] "the" "home"
方法2:tidytext 实现
library(tidyverse) library(tidytext) result <- tibble(text = input_str) %>% unnest_tokens(word, text) %>% count(word, name = "freq") %>% filter(freq >= 2) %>% mutate(first_pos = str_locate(input_str, paste0("\\b", word, "\\b"))[,1]) %>% arrange(first_pos) %>% pull(word)
运行结果和方法1完全一致。
内容的提问来源于stack exchange,提问作者John Thomas
相关产品推荐
相关产品推荐

