如何将R数据框转换为带共现计数的矩阵?
解决R中CP与P值同ID共现次数统计问题
原始数据
首先定义输入数据框:
df <- data.frame( ID = c("1", "1", "1", "1", "2", "2","2","2","3","3","3","3"), plist = c("CP1", "P2", "P3", "P4", "P1", "P2","P3","CP2","CP1","P2","P3","P5") )
方法一:使用tidyverse工具链(简洁易读)
借助dplyr和tidyr包可快速实现需求,步骤清晰:
library(dplyr) library(tidyr) result <- df %>% # 按样本ID分组 group_by(ID) %>% # 提取每个ID下的CP项和所有P项 summarise( cp = plist[grepl("^CP", plist)], p = list(plist[grepl("^P", plist)]), .groups = "drop" ) %>% # 将每个CP对应的P列表展开为单独行 unnest(p) %>% # 统计每对CP-P的共现次数 count(cp, p, name = "count") %>% # 转换为宽格式矩阵,缺失组合填充0 pivot_wider(names_from = p, values_from = count, values_fill = 0) %>% # 将CP列设为行名 column_to_rownames("cp") # 查看结果 print(result)
运行后输出的结果即为目标计数矩阵:
P1 P2 P3 P4 P5 CP1 0 2 2 1 1 CP2 1 1 1 0 0
方法二:使用Base R(无需额外安装包)
若不想依赖第三方包,可通过Base R实现相同逻辑:
# 提取每个ID对应的CP和P集合 id_cp <- tapply(df$plist, df$ID, function(x) x[grepl("^CP", x)]) id_p <- tapply(df$plist, df$ID, function(x) x[grepl("^P", x)]) # 生成所有CP-P的空计数组合 cp_list <- unique(unlist(id_cp)) p_list <- unique(unlist(id_p)) pair_counts <- expand.grid(cp = cp_list, p = p_list, count = 0) # 遍历每个ID,累加共现次数 for(i in names(id_cp)){ current_cp <- id_cp[[i]] current_ps <- id_p[[i]] for(p in current_ps){ pair_counts$count[pair_counts$cp == current_cp & pair_counts$p == p] <- pair_counts$count[pair_counts$cp == current_cp & pair_counts$p == p] + 1 } } # 转换为宽格式矩阵 result_matrix <- reshape(pair_counts, idvar = "cp", timevar = "p", direction = "wide") rownames(result_matrix) <- result_matrix$cp result_matrix$cp <- NULL colnames(result_matrix) <- gsub("count\\.", "", colnames(result_matrix)) # 查看结果 print(result_matrix)
核心逻辑说明
两种方法的核心思路一致:
- 按ID分组,区分每个样本中的CP项和P项;
- 建立CP与对应P项的配对关系;
- 统计每对CP-P在不同ID中出现的次数;
- 整理为目标矩阵格式。
内容的提问来源于stack exchange,提问作者Prakki Rama
相关产品推荐
相关产品推荐

