R文本分析:统计两关键词列表词汇在指定距离内的共现次数
跨关键词列表指定距离共现统计解决方案(R语言)
核心思路
先对文本分词,定位两个关键词列表中所有词汇的位置,再计算不同列表词汇间的位置差,统计符合指定距离阈值的共现次数。
实现代码
1. 安装并加载依赖包
install.packages(c("tokenizers", "dplyr")) library(tokenizers) library(dplyr)
2. 基础版函数(逻辑直观,适合小型文本)
count_cross_cooccur <- function(text, words1, words2, max_dist) { # 分词并标准化为小写(需区分大小写可移除tolower) tokens <- tokenize_words(text, lowercase = TRUE)[[1]] # 获取两个列表词汇在分词结果中的位置 pos1 <- which(tokens %in% tolower(words1)) pos2 <- which(tokens %in% tolower(words2)) cooccur_num <- 0 # 遍历所有位置对,判断距离是否符合要求 for (p1 in pos1) { for (p2 in pos2) { if (abs(p1 - p2) <= max_dist) { cooccur_num <- cooccur_num + 1 } } } return(cooccur_num) }
3. 高效版函数(向量操作,适配大型文本)
count_cross_cooccur_fast <- function(text, words1, words2, max_dist) { tokens <- tokenize_words(text, lowercase = TRUE)[[1]] pos1 <- which(tokens %in% tolower(words1)) pos2 <- which(tokens %in% tolower(words2)) # 批量计算所有位置对的距离,统计符合条件的数量 dist_matrix <- outer(pos1, pos2, function(x, y) abs(x - y)) sum(dist_matrix <= max_dist) }
测试示例
text <- c("The house is blue. The car is very big and red.") words1 <- c("car", "house") words2 <- c("blue", "red") # 设定最大距离为3,返回结果1,符合预期 count_cross_cooccur_fast(text, words1, words2, max_dist = 3)
自定义调整说明
- 大小写区分:若需保留大小写,移除代码中所有
tolower相关操作即可。 - 距离定义:代码中采用"词汇位置差绝对值"作为距离(如
house在位置2,blue在位置4,位置差为2)。若你的距离定义为"两个词之间最多包含N个词",将判断条件改为abs(p1 - p2) - 1 <= max_dist。 - 去重需求:若同一词汇对(如
house-blue)多次出现仅需统计一次,可使用以下去重版函数:
count_cross_cooccur_unique <- function(text, words1, words2, max_dist) { tokens <- tokenize_words(text, lowercase = TRUE)[[1]] token_df <- tibble(token = tokens, pos = seq_along(tokens)) w1_df <- filter(token_df, token %in% tolower(words1)) w2_df <- filter(token_df, token %in% tolower(words2)) # 生成词汇对并筛选符合距离要求的结果,最后去重计数 cooccur_pairs <- expand.grid(w1 = w1_df$token, w2 = w2_df$token) %>% mutate(dist = abs(w1_df$pos - w2_df$pos)) %>% filter(dist <= max_dist) %>% distinct(w1, w2) nrow(cooccur_pairs) }
内容的提问来源于stack exchange,提问作者dmort
相关产品推荐
相关产品推荐

