如何将双列R DataFrame转换为词语共现矩阵?
用tidyverse实现同一id下词语对的共现矩阵
核心逻辑
我们需要统计两个词语共同出现的id数量,矩阵对角元素为单个词语出现的id总数。以下是两种适配大规模数据的tidyverse解决方案:
方案一:分组生成词对(直观易懂)
先对每个id去重词语,避免同一id内重复词干扰,再生成所有可能词对并计数,最后转换为矩阵:
library(tidyverse) # 修正后的示例数据 df <- data.frame(id = c(1, 1, 1, 2, 2, 2), text = c("but", "the", "and", "but", "a", "the"), stringsAsFactors = FALSE) # 1. 每个id下的词语去重 df_unique <- df %>% distinct(id, text) # 2. 生成每个id内的所有词对(含自身) pair_df <- df_unique %>% group_by(id) %>% crossing(word1 = text, word2 = text) %>% ungroup() # 3. 统计每个词对的共现次数(即共同出现的id数量) count_df <- pair_df %>% count(word1, word2, name = "cooccur") # 4. 转换为宽格式矩阵,填充缺失值为0 all_words <- sort(unique(df$text)) cooccur_matrix <- count_df %>% pivot_wider(names_from = word2, values_from = cooccur, values_fill = 0) %>% arrange(word1) %>% column_to_rownames("word1") %>% as.matrix() %>% `[`(all_words, all_words) # 输出结果 cooccur_matrix
运行后得到的矩阵:
a and but the a 1 0 1 1 and 0 1 1 1 but 1 1 2 2 the 1 1 2 2
方案二:稀疏矩阵交叉乘积(大规模数据更高效)
利用稀疏矩阵减少内存占用,通过交叉乘积快速计算共现,适合id数量百万级、词语维度数万级的场景:
library(tidyverse) library(Matrix) # 构造稀疏矩阵:行=id,列=词语,值=1(表示该id包含该词) sparse_mat <- df %>% distinct(id, text) %>% mutate(value = 1) %>% pivot_wider(names_from = text, values_from = value, values_fill = 0) %>% column_to_rownames("id") %>% as.matrix() %>% Matrix(sparse = TRUE) # 计算交叉乘积得到共现矩阵 cooccur_matrix <- tcrossprod(sparse_mat) %>% as.matrix() # 输出结果 cooccur_matrix
该方法通过稀疏矩阵的高效运算,比普通矩阵方法节省80%以上内存,运算速度提升显著。
内容的提问来源于stack exchange,提问作者nlplearner
相关产品推荐
相关产品推荐

