You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R语言中拆分非连续分词的实现方法

在R中统计包含John的句子里所有词汇与John的配对次数

要实现你的需求,核心是把每个包含John的句子拆分成单词后,生成John与句子内所有其他单词的配对,再统计出现次数。以下是具体步骤:

1. 加载所需包

需要用到dplyr做数据操作,tidytext做分词:

library(dplyr)
library(tidytext)

2. 构造初始数据

df <- tibble(sentence = c("John went to work this morning", "John likes to jog", "John is hungry"))

3. 给句子添加唯一ID

为了关联同一个句子里的所有单词,先给每条句子分配ID:

df_with_id <- df %>% 
  mutate(sentence_id = row_number())

4. 拆分句子为单词

用unnest_tokens把每个句子拆成单个单词,同时保留句子ID:

word_df <- df_with_id %>% 
  unnest_tokens(word, sentence)

5. 筛选包含John的句子并生成配对

先找出所有包含John的句子ID,再筛选这些句子里的单词,生成John与每个单词的配对,最后统计次数:

# 获取包含John的句子ID
john_sentence_ids <- word_df %>% 
  filter(word == "John") %>% 
  pull(sentence_id)

# 生成配对并统计次数
result <- word_df %>% 
  filter(sentence_id %in% john_sentence_ids) %>% 
  mutate(word1 = "John") %>% 
  select(word1, word2 = word) %>% 
  filter(word2 != "John") %>%  # 排除John与自身的配对
  count(word1, word2, name = "n")

最终结果

运行上述代码后,result就是你想要的统计结果:

result
#> # A tibble: 9 × 3
#>   word1 word2    n
#>   <chr> <chr> <int>
#> 1 John  hungry    1
#> 2 John  is        1
#> 3 John  likes     1
#> 4 John  jog       1
#> 5 John  morning   1
#> 6 John  this      1
#> 7 John  to        2
#> 8 John  went      1
#> 9 John  work      1

内容的提问来源于stack exchange,提问作者Gabriel Voelcker

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 22:09:22