You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何移除全相同值列以避免R语言fisher.test报错

问题描述

我有分属两类样本的基因存在/缺失(1/0)计数数据,需要对每个基因执行fisher.test分析,但遇到所有样本中全为1或0的基因时会报错。基因数量有数百个,没法手动剔除,想知道怎么在现有dplyr循环代码里加入移除或跳过这类基因的步骤。

样本数据

mydata <- data.frame(sampleID = c("A", "B", "C", "D", "E", "F", "G"),
                     category = c("high", "low", "high", "high", "low", "high", "low"),
                     Gene1 = c(1, 1, 0, 0, 0, 1, 1),
                     Gene2 = c(0, 1, 1, 1, 1, 1, 0),
                     Gene3 = c(0, 0, 0, 1, 1, 1, 1),
                     Gene4 = c(1, 1, 1, 1, 1, 1, 1))

原循环代码

library(dplyr)
library(tidyr)
library(broom)

mydata %>%
  select(-sampleID) %>%
  pivot_longer(cols = -category, names_to = "gene") %>%
  group_by(gene) %>%
  summarise(fisher_test = list(tidy(fisher.test(table(category, value))))) %>%
  unnest(fisher_test) %>%
  mutate(odds_ratio = exp(estimate)) %>% 
  select(-method, -alternative)

报错信息

Caused by error in `fisher.test()`:
! 'x' must have at least 2 rows and columns
Run `rlang::last_error()` to see where the error occurred.

解决方案

报错根源是全0或全1的基因无法生成符合要求的2×2列联表,导致fisher.test无法运行。以下两种方式可跳过这类无效基因:

方法1:宽格式提前筛选有效基因

先在宽表阶段过滤掉全0或全1的基因列,减少后续处理的数据量,效率更高:

library(dplyr)
library(tidyr)
library(broom)

# 筛选出同时存在0和1两种取值的基因
valid_genes <- mydata %>%
  select(starts_with("Gene")) %>% # 匹配所有基因列(可根据实际列名调整规则)
  select(where(~ n_distinct(.x) > 1)) %>% # 保留取值不止一种的列
  colnames()

# 基于有效基因执行分析
mydata %>%
  select(sampleID, category, all_of(valid_genes)) %>%
  select(-sampleID) %>%
  pivot_longer(cols = -category, names_to = "gene") %>%
  group_by(gene) %>%
  summarise(fisher_test = list(tidy(fisher.test(table(category, value))))) %>%
  unnest(fisher_test) %>%
  mutate(odds_ratio = exp(estimate)) %>% 
  select(-method, -alternative)

方法2:长格式分组后过滤无效基因

在分组后直接过滤掉取值单一的基因,逻辑更直观:

library(dplyr)
library(tidyr)
library(broom)

mydata %>%
  select(-sampleID) %>%
  pivot_longer(cols = -category, names_to = "gene") %>%
  group_by(gene) %>%
  filter(n_distinct(value) > 1) %>% # 只保留有两种取值的基因组
  summarise(fisher_test = list(tidy(fisher.test(table(category, value))))) %>%
  unnest(fisher_test) %>%
  mutate(odds_ratio = exp(estimate)) %>% 
  select(-method, -alternative)

两种方法核心都是通过n_distinct(value)判断基因是否同时存在0和1两种取值,以此跳过无法执行fisher.test的无效基因。

内容的提问来源于stack exchange,提问作者ABee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 06:37:05