You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用dplyr转换数据,对每个甲基化位点做Fisher精确检验

用dplyr对甲基化数据集逐行执行Fisher精确检验

我们需要对包含肿瘤组(Tumor)和正常组(Normal)甲基化/非甲基化计数的数据集,按每个位点(每行)构建2x2列联表并执行Fisher精确检验,以下是基于dplyr的实现方案:

示例数据集

首先定义示例数据:

methyl_dat <- data.frame(loci = c("site1", "site2", "site3", "site4"), 
           Methy.tumor = c(50, 5, 60, 12), 
           UnMethy.tumor = c(60, 0, 65, 5), 
           Methy.Normal = c(13, 5, 22, 3),
           UnMethy.Normal = c(86, 0, 35, 3) )

每个位点的列联表结构以site1为例:

Normal
Tumor          Methyl  UnMethy
  Methy         50      13
  UnMethy       60      86

实现步骤

  1. 加载必要的包
library(dplyr)
  1. 逐行处理并执行Fisher检验
# 逐行处理每个位点,执行Fisher检验并提取结果
methyl_fisher_result <- methyl_dat %>%
  rowwise(loci) %>%
  mutate(
    # 构建2x2列联表矩阵
    cont_table = list(
      matrix(
        c(Methy.tumor, Methy.Normal,
          UnMethy.tumor, UnMethy.Normal),
        nrow = 2,
        dimnames = list(
          c("Tumor_Methyl", "Tumor_UnMethyl"),
          c("Normal_Methyl", "Normal_UnMethyl")
        )
      )
    ),
    # 执行Fisher精确检验
    fisher_res = list(fisher.test(cont_table)),
    # 提取关键结果:p值、优势比(OR)及置信区间
    p_value = fisher_res$p.value,
    odds_ratio = fisher_res$estimate,
    or_ci_low = fisher_res$conf.int[1],
    or_ci_high = fisher_res$conf.int[2]
  ) %>%
  # 保留有用列,取消分组
  select(loci, p_value, odds_ratio, or_ci_low, or_ci_high) %>%
  ungroup()
  1. 查看结果
print(methyl_fisher_result)

说明

  • rowwise(loci):指定按loci列分组,实现逐行独立处理
  • 列联表矩阵的维度对应检验需求,确保行是肿瘤组的甲基化/非甲基化计数,列是正常组的对应计数
  • 提取的结果包含每个位点的显著性p值、优势比及其95%置信区间,可直接用于后续分析

内容的提问来源于stack exchange,提问作者sahuno

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 02:45:23