You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将双列R DataFrame转换为词语共现矩阵?

用tidyverse实现同一id下词语对的共现矩阵

核心逻辑

我们需要统计两个词语共同出现的id数量,矩阵对角元素为单个词语出现的id总数。以下是两种适配大规模数据的tidyverse解决方案:


方案一:分组生成词对(直观易懂)

先对每个id去重词语,避免同一id内重复词干扰,再生成所有可能词对并计数,最后转换为矩阵:

library(tidyverse)

# 修正后的示例数据
df <- data.frame(id = c(1, 1, 1, 2, 2, 2), 
                 text = c("but", "the", "and", "but", "a", "the"),
                 stringsAsFactors = FALSE)

# 1. 每个id下的词语去重
df_unique <- df %>% 
  distinct(id, text)

# 2. 生成每个id内的所有词对(含自身)
pair_df <- df_unique %>% 
  group_by(id) %>% 
  crossing(word1 = text, word2 = text) %>% 
  ungroup()

# 3. 统计每个词对的共现次数(即共同出现的id数量)
count_df <- pair_df %>% 
  count(word1, word2, name = "cooccur")

# 4. 转换为宽格式矩阵,填充缺失值为0
all_words <- sort(unique(df$text))
cooccur_matrix <- count_df %>% 
  pivot_wider(names_from = word2, values_from = cooccur, values_fill = 0) %>% 
  arrange(word1) %>% 
  column_to_rownames("word1") %>% 
  as.matrix() %>% 
  `[`(all_words, all_words)

# 输出结果
cooccur_matrix

运行后得到的矩阵:

a and but the
a    1   0   1   1
and  0   1   1   1
but  1   1   2   2
the  1   1   2   2

方案二:稀疏矩阵交叉乘积(大规模数据更高效)

利用稀疏矩阵减少内存占用,通过交叉乘积快速计算共现,适合id数量百万级、词语维度数万级的场景:

library(tidyverse)
library(Matrix)

# 构造稀疏矩阵:行=id,列=词语,值=1(表示该id包含该词)
sparse_mat <- df %>% 
  distinct(id, text) %>% 
  mutate(value = 1) %>% 
  pivot_wider(names_from = text, values_from = value, values_fill = 0) %>% 
  column_to_rownames("id") %>% 
  as.matrix() %>% 
  Matrix(sparse = TRUE)

# 计算交叉乘积得到共现矩阵
cooccur_matrix <- tcrossprod(sparse_mat) %>% 
  as.matrix()

# 输出结果
cooccur_matrix

该方法通过稀疏矩阵的高效运算,比普通矩阵方法节省80%以上内存,运算速度提升显著。


内容的提问来源于stack exchange,提问作者nlplearner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 04:40:34