You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在dplyr分组流程中用临时变量计算组内经验FDR遇错求助

问题:分组计算经验FDR时出现"object 'Type' not found"错误

我需要按Group变量分组,利用组内Type为Random的p值分布,计算每组各类型的经验FDR(错误发现率)。已经编写了calculate_empirical_fdr函数并准备了示例数据集,但在dplyr分组流程中尝试用临时变量存储Random组的p值时,触发了Error in filter(Type == "Random") : object 'Type' not found错误,求解决办法。


相关函数与数据集代码

经验FDR计算函数

library(dplyr)
library(data.table)
calculate_empirical_fdr = function(control_pVal, test_pVal) {
  m_control = length(control_pVal)
  m_test = length(test_pVal)
  unlist(lapply(test_pVal, function(significance_threshold) {
    m_control = length(control_pVal)
    m_test = length(test_pVal)

    FP_expected = length(control_pVal[control_pVal<=significance_threshold])*m_test/m_control
    S = length(test_pVal[test_pVal<=significance_threshold])
    return(FP_expected/S)
  }))
}

示例数据集

set.seed(42)
library(dplyr); library(data.table)
dataset_test = data.table(Type = c(rep("Random", 500),
                                   rep("test1", 500),
                                   rep("test2", 500)),
                          Group = sample(c("group1", "group2", "group3"), 1500, replace = T),
                          Pvalue = c(runif(n = 500),
                                     rbeta(n = 500, shape1 = 1, shape2 = 4),
                                     rbeta(n = 500, shape1 = 1, shape2 = 6))
                          )

报错的dplyr代码

dataset_test %>%
  group_by(Group) %>%
  {filter(Type=="Random") %>% select(Pvalue) ->> control_set } %>%
  group_by(Type, add = T) %>%
  mutate(FDR_empirical = calculate_empirical_fdr(control_pVal = control_set,
                              test_pVal = Pvalue)) %>%
  data.table()

错误信息:

Error in filter(Type == "Random") : object 'Type' not found


错误原因

大括号{}包裹的代码块未显式引用管道传递的数据集,导致filter无法找到Type列;同时,用->>赋值全局变量的方式不适合分组场景——每个组的Random型p值是独立的,全局变量会被后续组覆盖,无法实现分组计算的需求。


解决方法

方法1:用dplyr分组嵌套实现

先按Group分组,提取每组的Random型p值作为嵌套列,再对组内不同Type计算FDR:

dataset_test %>%
  group_by(Group) %>%
  # 提取当前组的Random p值存入嵌套列
  mutate(control_p = list(Pvalue[Type == "Random"])) %>%
  # 按Group+Type分组计算FDR
  group_by(Type, .add = TRUE) %>%
  mutate(FDR_empirical = calculate_empirical_fdr(control_pVal = unlist(control_p),
                                                 test_pVal = Pvalue)) %>%
  # 清理临时列并转换为data.table
  select(-control_p) %>%
  ungroup() %>%
  as.data.table()

方法2:用data.table原生分组(更高效)

利用data.table的分组语法直接实现,逻辑更简洁:

# 先按Group提取各组的Random p值,生成对应字典
control_list = dataset_test[Type == "Random", .(control_p = list(Pvalue)), by = Group]
control_dict = setNames(control_list$control_p, control_list$Group)

# 按Group+Type分组计算FDR
dataset_test[, FDR_empirical := calculate_empirical_fdr(control_pVal = unlist(control_dict[.BY$Group]),
                                                        test_pVal = Pvalue),
             by = .(Group, Type)]

内容的提问来源于stack exchange,提问作者lizaveta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 20:42:33