You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R文本分析:统计两关键词列表词汇在指定距离内的共现次数

跨关键词列表指定距离共现统计解决方案(R语言)

核心思路

先对文本分词,定位两个关键词列表中所有词汇的位置,再计算不同列表词汇间的位置差,统计符合指定距离阈值的共现次数。

实现代码

1. 安装并加载依赖包

install.packages(c("tokenizers", "dplyr"))
library(tokenizers)
library(dplyr)

2. 基础版函数(逻辑直观,适合小型文本)

count_cross_cooccur <- function(text, words1, words2, max_dist) {
  # 分词并标准化为小写(需区分大小写可移除tolower)
  tokens <- tokenize_words(text, lowercase = TRUE)[[1]]
  # 获取两个列表词汇在分词结果中的位置
  pos1 <- which(tokens %in% tolower(words1))
  pos2 <- which(tokens %in% tolower(words2))
  
  cooccur_num <- 0
  # 遍历所有位置对,判断距离是否符合要求
  for (p1 in pos1) {
    for (p2 in pos2) {
      if (abs(p1 - p2) <= max_dist) {
        cooccur_num <- cooccur_num + 1
      }
    }
  }
  return(cooccur_num)
}

3. 高效版函数(向量操作,适配大型文本)

count_cross_cooccur_fast <- function(text, words1, words2, max_dist) {
  tokens <- tokenize_words(text, lowercase = TRUE)[[1]]
  pos1 <- which(tokens %in% tolower(words1))
  pos2 <- which(tokens %in% tolower(words2))
  
  # 批量计算所有位置对的距离,统计符合条件的数量
  dist_matrix <- outer(pos1, pos2, function(x, y) abs(x - y))
  sum(dist_matrix <= max_dist)
}

测试示例

text <- c("The house is blue. The car is very big and red.")
words1 <- c("car", "house") 
words2 <- c("blue", "red") 

# 设定最大距离为3,返回结果1,符合预期
count_cross_cooccur_fast(text, words1, words2, max_dist = 3)

自定义调整说明

  • 大小写区分:若需保留大小写,移除代码中所有tolower相关操作即可。
  • 距离定义:代码中采用"词汇位置差绝对值"作为距离(如house在位置2,blue在位置4,位置差为2)。若你的距离定义为"两个词之间最多包含N个词",将判断条件改为abs(p1 - p2) - 1 <= max_dist。
  • 去重需求:若同一词汇对(如house-blue)多次出现仅需统计一次,可使用以下去重版函数:
count_cross_cooccur_unique <- function(text, words1, words2, max_dist) {
  tokens <- tokenize_words(text, lowercase = TRUE)[[1]]
  token_df <- tibble(token = tokens, pos = seq_along(tokens))
  
  w1_df <- filter(token_df, token %in% tolower(words1))
  w2_df <- filter(token_df, token %in% tolower(words2))
  
  # 生成词汇对并筛选符合距离要求的结果,最后去重计数
  cooccur_pairs <- expand.grid(w1 = w1_df$token, w2 = w2_df$token) %>%
    mutate(dist = abs(w1_df$pos - w2_df$pos)) %>%
    filter(dist <= max_dist) %>%
    distinct(w1, w2)
  
  nrow(cooccur_pairs)
}

内容的提问来源于stack exchange,提问作者dmort

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 09:30:51