WGCNA Signed网络分析出现负kME值的原因咨询
提问:WGCNA Signed网络中出现负kME值的困惑
我正在开展WGCNA分析,遇到了诸多难题。我在bwnet分析中采用了Signed网络(TOMType = "signed"),原本认为这意味着仅考虑正相关,负相关会被设为0。但当我绘制kME与Gene Significance的关系图时,发现大量基因与模块特征基因表现相反(存在负kME值),我无法理解该现象的原因,希望能得到帮助,目前我对所有数据都产生了质疑。
网络构建代码
#Network construction bwnet <- blockwiseModules(norm_counts, maxBlockSize = 18456, TOMType = "signed", power = soft_power, mergeCutHeight = 0.25, numericLabels = FALSE, randomSeed = 1234, verbose = 3) cor <- temp_cor
kME与Trait相关性绘图代码
#Plotting kME vs Correlation to trait # load needed libraries---- library(ggplot2) library(WGCNA) library(dplyr) library(tidyr) library(patchwork) # or use facets in ggplot #load all files needed and prepare data---- load("results_V4/norm_counts.RData") load("results_V4/bwnet.RData") data <- read.csv("count_data_df_JT.csv", sep= ";", row.names = "genes") meta <- read.csv("sample_info.csv", sep =";", row.names = "sampletype") growth_condition <- meta$growth # Convert growth condition to numeric (1 for 'c', 0 for 'sg') growth_condition_numeric <- ifelse(growth_condition == "c", 1, 0) # Transpose norm_counts if needed genes(columns) , samples (rows) norm_counts_t <- t(norm_counts) # Prepare eigengenes module_eigengenes <- bwnet$MEs module_colors <- bwnet$colors unique_modules <- setdiff(unique(module_colors), "grey") # Recalculate module membership (kME) module.membership.measure <- cor(norm_counts_t, module_eigengenes, use = "p") # Get gene significance (GS) for growth GS <- apply(norm_counts, 1, function(x) cor(x, growth_condition_numeric)) # Prepare plotting dataframe plot_df <- data.frame() for (mod in unique_modules) { # Get genes belonging to the current module module_genes <- names(module_colors[module_colors == mod]) # Get the module membership (kME) for the genes in the current module module_kME <- module.membership.measure[module_genes, paste0("ME", mod)] # Get Gene Significance (GS) for the current module GS_module <- GS[module_genes] # Create a dataframe to store the module's kME and GS df <- data.frame( Module = mod, kME = module_kME, GS = GS_module ) # Calculate Pearson correlation and p-value for the TS-MM corr <- cor(df$kME, df$GS, use = "complete.obs") pval <- corPvalueStudent(corr, nrow(df)) # Calculate p-value for correlation # Add correlation label and number of genes in the module df$corr_label <- paste0("r = ", round(corr, 2), ", p = ", format.pval(pval, digits = 2)) df$n_genes <- length(module_genes) # Append to the plot dataframe plot_df <- rbind(plot_df, df) } # Step 5: Plotting using ggplot2 pdf("module_membership_vs_trait_significance.pdf", width = 10, height = 8) ggplot(plot_df, aes(x = kME, y = GS, color = Module)) + geom_point(alpha = 0.7) + # Scatter plot of kME vs GS geom_smooth(method = "lm", color = "black") + # Regression line for TS-MM facet_wrap(~Module, scales = "free", ncol = 4) + # Create a facet for each module labs(x = "Module Membership (kME)", y = "Gene Significance for Growth Condition", title = "Module Membership vs Trait Significance (TS-MM)") + theme_minimal() + geom_text( data = plot_df %>% distinct(Module, corr_label, n_genes), aes(x = -Inf, y = Inf, label = paste0("n=", n_genes, "\n", corr_label)), hjust = -0.1, vjust = 1.2, inherit.aes = FALSE ) dev.off()
解答
核心误解澄清
你对TOMType = "signed"的功能理解有误:
signed模式不会将负相关置为0:它通过公式(1 + cor)/2计算邻接矩阵,负相关会被转换为0~0.5的低值,而非直接丢弃。该模式的作用是让模块划分优先聚合正共表达的基因,但不排除模块内存在表达模式相反的基因。- 若需要完全忽略负相关、仅基于正相关构建网络,应使用
TOMType = "signed hybrid",该模式会直接将负相关设为0。
负kME值的合理性
模块特征基因(ME)是模块内所有基因表达的第一主成分,它捕捉的是模块内最显著的整体表达趋势。即使模块由正共表达基因聚合而成,少数基因(比如模块内的负调控因子、表达量波动与整体趋势相反的基因)仍可能与ME呈现负相关,从而出现负kME值——这是正常的生物学现象,不代表分析流程出错。
代码优化与验证建议
- 规范kME计算:你手动计算kME的方式是正确的,但可以使用WGCNA内置的
signedKME函数,结果更贴合WGCNA的分析逻辑:
# 替换手动计算kME的代码 module.membership.measure <- signedKME(norm_counts_t, module_eigengenes, outputColumnName = "kME")
- 验证模块表达模式:绘制模块内基因的表达热图,确认模块内大部分基因的表达模式是否一致,少数负kME基因是否属于合理例外:
# 替换为你的目标模块颜色 target_module <- "blue" module_genes <- names(module_colors[module_colors == target_module]) module_expr <- norm_counts[module_genes, ] heatmap(t(module_expr), scale = "row", col = colorRampPalette(c("blue", "white", "red"))(100), main = paste("Expression Heatmap of", target_module, "Module"))
内容的提问来源于stack exchange,提问作者Laura
相关产品推荐
相关产品推荐

