You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R筛选非零值占比超过75%的数据框列(基因)

R数据框筛选:保留非零值占比>75%的列

实现思路

核心逻辑是计算每列中非零值的占比,筛选出占比超过75%的列,步骤如下:

  • 统计每列非零值的比例
  • 基于比例阈值(>0.75)筛选目标列

具体代码

首先加载数据(需指定表头和行名参数,确保数据结构正确):

x <- read.table(text = "    Gene1   Gene2   Gene3   Gene4
cell1   0   0   0   0
cell2   1   2   1.7 1.5
cell3   2   0   2.5 0
cell4   2.5 0   2.1 2.5
cell5   2   0   0.8 1.5
cell6   3   0   0.5 2.1
cell7   0   0   0   1.2
cell8   0   0   1.6 1.5
cell9   1.6 0   2.3 2.7
cell10  2.4 0   2.1 2.1", header = TRUE, row.names = 1)

基础R实现

# 计算每列非零值占比
non_zero_ratio <- colMeans(x != 0)
# 筛选占比>75%的列
filtered_x <- x[, non_zero_ratio > 0.75]
# 查看筛选结果
filtered_x

tidyverse(dplyr)实现

如果习惯使用tidyverse工具链,可采用更简洁的写法:

library(dplyr)
filtered_x <- x %>% 
  select(where(~ mean(.x != 0) > 0.75))

结果说明

根据给定数据的实际计算:

  • Gene1非零占比70%(不满足>75%)
  • Gene2非零占比10%(被剔除)
  • Gene3非零占比80%(保留)
  • Gene4非零占比90%(保留)
    最终结果保留Gene3和Gene4。若需求为≥75%,只需将代码中的> 0.75改为>= 0.75,即可保留Gene1、Gene3、Gene4。

内容的提问来源于stack exchange,提问作者sswzrdss

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 18:15:16