函数内limma::removeBatchEffect报错,外部运行正常问题排查
问题描述
我编写了一个函数,通过布尔型参数use_batch控制是否执行批次效应校正,函数内的校正代码块如下:
if (use_batch){ scores = as.data.frame(scores) cancer.type = metadata[['Cancer_Type']] scores = limma::removeBatchEffect(t(scores) ,batch = cancer.type) scores = t(scores) }
运行该函数时出现报错:
Error in rowMeans(y$exprs, na.rm = TRUE) : 'x' must be numeric
但直接用真实数据my_scores和my_meta在函数外执行完全相同的逻辑却正常运行:
> if (TRUE){ scoresx = as.data.frame(my_scores) cancer.type = my_meta[['Cancer_Type']] scoresx = limma::removeBatchEffect(t(scoresx) ,batch = cancer.type) scoresx = t(scoresx) }
注意:当use_batch=FALSE时函数运行完全正常,回溯显示错误出在scores = limma::removeBatchEffect(t(scores) ,batch = cancer.type)这一行。
以下是数据子集:
> dput(my_scores[1:6,1:6]) structure(c(-0.0741098696855045, -0.094401270881699, 0.0410284948786532, -0.163302950330185, -0.0942478217207681, -0.167314411991775, -0.178835447722468, -0.253897294559596, -0.0372301980787381, -0.230579110769457, -0.224125346052727, -0.196933050675633, -0.0384421660291032, -0.0275306107582565, 0.186447606591857, -0.124972070102036, -0.15348122673842, -0.106812144494277, -0.0924083910597563, -0.172356328661097, -0.0172673823614314, 0.0280649471541352, -0.128925304635747, -0.0875076743713435, -0.0680848706469295, -0.173427291586957, -0.0106773958944477, -0.0015805672257001, -0.0751114943036091, -0.0737177243152751, -0.0391833488213571, -0.0275279418713283, 0.0156454755097513, 0.0285160860867748, -0.0633367938488132, 0.0252778805872529), dim = c(6L, 6L), dimnames = list(c("Pt1", "Pt10", "Pt101", "Pt103", "Pt106", "Pt11"), c("CD4-T-cells", "CD8-T-cells", "T-helpers", "NK-cells", "Monocytes", "Neutrophils" ))) > dput(my_meta[1:6,2:7]) structure(list(samples = c("Pt1", "Pt10", "Pt101", "Pt103", "Pt106", "Pt11"), Response = c("NoResponse", "NoResponse", "Response", "NoResponse", "NoResponse", "NoResponse"), RECIST = c("PD", "SD", "PR", "PD", "PD", "PD"), Gender = c("male", "female", "female", "female", "male", "female"), treatment = c("anti-PD1", "anti-PD1", "anti-PD1", "anti-PD1", "anti-PD1", "anti-PD1"), Cancer_Type = c("Melanoma", "Melanoma", "Melanoma", "Melanoma", "Melanoma", "Melanoma")), row.names = c("Pt1", "Pt10", "Pt101", "Pt103", "Pt106", "Pt11"), class = "data.frame")
问题原因分析
- 变量作用域不匹配:函数内部的
metadata可能未正确绑定到输入的元数据参数,而是引用了全局环境中其他非预期对象,导致cancer.type的长度/结构与scores转置后的样本数不匹配;或者函数内的scores被其他代码修改,转置后不再是纯数值结构。而全局环境的my_scores和my_meta是明确匹配的,因此运行正常。 - 数据类型识别偏差:
limma::removeBatchEffect会优先将data.frame识别为ExpressionSet对象(这也是报错中出现y$exprs的原因),而非普通数值矩阵。函数内转置后的scores可能因结构问题被误判为非数值矩阵,导致函数尝试访问ExpressionSet专属的exprs属性,但实际对象无此属性,进而引发rowMeans处理非数值对象的错误。而全局环境中t(scoresx)是纯数值矩阵,函数会按普通矩阵处理,因此无问题。 - 隐式类型转换问题:函数内的
scores转换为data.frame后,可能混入了非数值元素(比如字符型列),转置后整个对象变成字符型,不满足removeBatchEffect对数值输入的要求。而全局环境的my_scores是纯数值矩阵,转置后仍保持数值型,因此能正常运行。
解决建议
- 在函数内明确将
scores转换为数值矩阵:替换scores = as.data.frame(scores)为scores = as.matrix(scores),确保转置后是纯数值矩阵再传入removeBatchEffect。 - 检查函数定义,确保
metadata作为参数传入,避免依赖全局变量,保证批次变量与样本一一对应。 - 验证
cancer.type的长度与t(scores)的行数(即原scores的列数)完全一致,避免维度不匹配问题。
内容的提问来源于stack exchange,提问作者Kev
相关产品推荐
相关产品推荐

