R语言统计studyID分组下post_id值共现次数并生成下三角矩阵
R语言实现共现统计及下三角矩阵输出
实现逻辑
我们需要统计每个studyID分组下,任意两个不同post_id的共同出现次数,最终仅在4×4矩阵的下三角填充统计值,其余位置设为NA,核心步骤如下:
- 按
studyID分组,提取每组内的唯一post_id值 - 对每组的
post_id生成所有两两从小到大的不重复组合 - 统计所有组合的出现频次
- 构造目标矩阵,对应填充统计结果
完整实现代码
依赖tidyverse的实现方式
# 读取数据集 data <- read.csv("https://raw.githubusercontent.com/ilzl/i/master/pr.csv")[c(1,7)] # 加载依赖包 library(tidyverse) # 统计共现频次 count_res <- data %>% # 先去重,避免同一studyID下重复的post_id干扰统计 distinct(studyID, post_id) %>% group_by(studyID) %>% # 过滤掉只有单个post_id的分组,无共现可能 filter(n_distinct(post_id) >= 2) %>% summarise( # 生成当前分组所有两两post_id组合 pair = list(combn(sort(unique(post_id)), 2, simplify = FALSE)) ) %>% unnest(pair) %>% # 提取组合对应的矩阵行、列索引(大值为行,小值为列,对应下三角位置) mutate( mat_row = map_int(pair, ~.x[2]), mat_col = map_int(pair, ~.x[1]) ) %>% count(mat_row, mat_col) # 构造结果矩阵 result_mat <- matrix(NA, nrow = 4, ncol = 4, dimnames = list(1:4, 1:4)) # 填充统计值到对应下三角位置 for (i in seq(nrow(count_res))) { result_mat[count_res$mat_row[i], count_res$mat_col[i]] <- count_res$n[i] } # 查看结果 print(result_mat)
基础R实现方式(无需额外安装包)
# 读取数据集 data <- read.csv("https://raw.githubusercontent.com/ilzl/i/master/pr.csv")[c(1,7)] # 按studyID拆分post_id post_by_id <- split(data$post_id, data$studyID) # 统计所有共现组合的频次 pair_count <- table(unlist(lapply(post_by_id, function(x) { unique_post <- sort(unique(x)) # 跳过只有单个post_id的分组 if (length(unique_post) < 2) return(NULL) # 生成两两组合并拼接为标识字符串 apply(combn(unique_post, 2), 2, paste, collapse = "_") }))) # 构造并填充结果矩阵 result_mat <- matrix(NA, nrow = 4, ncol = 4, dimnames = list(1:4, 1:4)) for (pair_name in names(pair_count)) { idx <- as.integer(strsplit(pair_name, "_")[[1]]) result_mat[idx[2], idx[1]] <- as.integer(pair_count[pair_name]) } # 查看结果 print(result_mat)
输出结果验证
运行以上代码后得到的矩阵和预期完全一致:
- [2,1](post_id 1与2共现):31次
- [3,1](post_id 1与3共现):3次
- [4,1](post_id 1与4共现):1次
- [3,2](post_id 2与3共现):3次
- [4,2](post_id 2与4共现):1次
- [4,3](post_id 3与4共现):1次
- 其余位置均为NA
内容的提问来源于stack exchange,提问作者rnorouzian
相关产品推荐
相关产品推荐

