You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言如何使用dplyr统计字符串中出现至少2次的单词

tidyverse 实现字符串重复单词识别

测试输入

input_str <- "the home home on the range the friend"

目标:识别字符串中出现次数≥2的单词,输出按单词在原字符串首次出现的顺序排列,最终返回c("the", "home"),优先使用stringr、dplyr实现。

方法1:stringr + dplyr 实现(无额外依赖)

library(tidyverse)

result <- tibble(word = str_split(input_str, "\\s+")[[1]]) %>%
  count(word, name = "freq") %>%
  filter(freq >= 2) %>%
  mutate(first_pos = str_locate(input_str, paste0("\\b", word, "\\b"))[,1]) %>%
  arrange(first_pos) %>%
  pull(word)

代码逻辑说明

  • str_split按空白符拆分原字符串,得到独立的单词向量
  • count统计每个单词的出现次数
  • filter保留出现次数≥2的单词
  • 新增first_pos列记录每个单词在原字符串中第一次出现的位置,按该列排序可保证输出顺序和要求一致,避免默认按字母排序导致home排在the前的问题
  • pull提取最终的单词列表

运行输出:

> print(result)
[1] "the"  "home"

方法2:tidytext 实现

library(tidyverse)
library(tidytext)

result <- tibble(text = input_str) %>%
  unnest_tokens(word, text) %>%
  count(word, name = "freq") %>%
  filter(freq >= 2) %>%
  mutate(first_pos = str_locate(input_str, paste0("\\b", word, "\\b"))[,1]) %>%
  arrange(first_pos) %>%
  pull(word)

运行结果和方法1完全一致。


内容的提问来源于stack exchange,提问作者John Thomas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 21:12:19