在dplyr分组流程中用临时变量计算组内经验FDR遇错求助
问题:分组计算经验FDR时出现"object 'Type' not found"错误
我需要按Group变量分组,利用组内Type为Random的p值分布,计算每组各类型的经验FDR(错误发现率)。已经编写了calculate_empirical_fdr函数并准备了示例数据集,但在dplyr分组流程中尝试用临时变量存储Random组的p值时,触发了Error in filter(Type == "Random") : object 'Type' not found错误,求解决办法。
相关函数与数据集代码
经验FDR计算函数
library(dplyr) library(data.table) calculate_empirical_fdr = function(control_pVal, test_pVal) { m_control = length(control_pVal) m_test = length(test_pVal) unlist(lapply(test_pVal, function(significance_threshold) { m_control = length(control_pVal) m_test = length(test_pVal) FP_expected = length(control_pVal[control_pVal<=significance_threshold])*m_test/m_control S = length(test_pVal[test_pVal<=significance_threshold]) return(FP_expected/S) })) }
示例数据集
set.seed(42) library(dplyr); library(data.table) dataset_test = data.table(Type = c(rep("Random", 500), rep("test1", 500), rep("test2", 500)), Group = sample(c("group1", "group2", "group3"), 1500, replace = T), Pvalue = c(runif(n = 500), rbeta(n = 500, shape1 = 1, shape2 = 4), rbeta(n = 500, shape1 = 1, shape2 = 6)) )
报错的dplyr代码
dataset_test %>% group_by(Group) %>% {filter(Type=="Random") %>% select(Pvalue) ->> control_set } %>% group_by(Type, add = T) %>% mutate(FDR_empirical = calculate_empirical_fdr(control_pVal = control_set, test_pVal = Pvalue)) %>% data.table()
错误信息:
Error in filter(Type == "Random") : object 'Type' not found
错误原因
大括号{}包裹的代码块未显式引用管道传递的数据集,导致filter无法找到Type列;同时,用->>赋值全局变量的方式不适合分组场景——每个组的Random型p值是独立的,全局变量会被后续组覆盖,无法实现分组计算的需求。
解决方法
方法1:用dplyr分组嵌套实现
先按Group分组,提取每组的Random型p值作为嵌套列,再对组内不同Type计算FDR:
dataset_test %>% group_by(Group) %>% # 提取当前组的Random p值存入嵌套列 mutate(control_p = list(Pvalue[Type == "Random"])) %>% # 按Group+Type分组计算FDR group_by(Type, .add = TRUE) %>% mutate(FDR_empirical = calculate_empirical_fdr(control_pVal = unlist(control_p), test_pVal = Pvalue)) %>% # 清理临时列并转换为data.table select(-control_p) %>% ungroup() %>% as.data.table()
方法2:用data.table原生分组(更高效)
利用data.table的分组语法直接实现,逻辑更简洁:
# 先按Group提取各组的Random p值,生成对应字典 control_list = dataset_test[Type == "Random", .(control_p = list(Pvalue)), by = Group] control_dict = setNames(control_list$control_p, control_list$Group) # 按Group+Type分组计算FDR dataset_test[, FDR_empirical := calculate_empirical_fdr(control_pVal = unlist(control_dict[.BY$Group]), test_pVal = Pvalue), by = .(Group, Type)]
内容的提问来源于stack exchange,提问作者lizaveta
相关产品推荐
相关产品推荐

