You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按指定区间统计R数据框各列数值的分组数量

在R中对多列基因表达数据按自定义区间分组统计计数

我有一个1246×60660的dataframe,数据示例如下:

gene1             gene2              gene3
sample1          1615.7529         41.932474           697.9728
sample2           663.2001          8.602831          1198.1398
sample3          2406.1532         12.622443          1033.4625
sample4           836.3808         60.144235           259.3720
sample5          1217.8192         22.775497           695.9924
sample6           865.0344         15.350298           683.5397
sample7           935.3658         20.380676           540.6242
sample8           667.3883         56.939874          1056.6981

我需要将每个基因列的样本值归入以下分组并统计每组的样本数:

  • none = 0
  • ultra low = 1-4
  • low = 5 - 100
  • medium = 101 - 1000
  • high = 1000及以上

最终要得到如下格式的结果:

gene1      gene2      gene3
none                  0          0          0 
ultra low             0          0          0
low                   0          8          0
medium                5          0          5
high                  3          0          3

我考虑过用count或aggregate函数,但这些函数的示例大多只针对单列统计,不知道怎么批量应用到每一列,请问该怎么实现?


解决方案

方法1:使用dplyr + tidyr(代码简洁易读)

先定义分组区间和标签,通过数据格式转换实现多列批量统计:

library(dplyr)
library(tidyr)

# 定义分组规则
groups <- c("none", "ultra low", "low", "medium", "high")
breaks <- c(-Inf, 0, 4, 100, 1000, Inf)

# 处理数据
result <- df %>%
  # 转成长格式,统一处理所有列
  pivot_longer(cols = everything(), names_to = "gene", values_to = "value") %>%
  # 按基因分组,给每个值分配对应分组标签
  group_by(gene) %>%
  mutate(group = cut(value, breaks = breaks, labels = groups, include.lowest = TRUE)) %>%
  # 统计每个基因下各分组的样本数
  count(group) %>%
  # 转回宽格式,缺失分组填充0
  pivot_wider(names_from = "gene", values_from = "n", values_fill = 0) %>%
  # 按预设分组顺序排序
  arrange(factor(group, levels = groups)) %>%
  # 将分组列设为行名
  column_to_rownames("group")

# 查看结果
print(result)

方法2:使用base R(无需额外包,适配大数据量)

针对6万列的大数据量,base R的apply系列函数效率更稳定:

# 定义分组规则
groups <- c("none", "ultra low", "low", "medium", "high")
breaks <- c(-Inf, 0, 4, 100, 1000, Inf)

# 对每一列批量执行分组计数
count_matrix <- sapply(df, function(col) {
  # 给列中每个值分配分组标签
  group_labels <- cut(col, breaks = breaks, labels = groups, include.lowest = TRUE)
  # 强制保留所有分组,缺失分组填充0
  table(factor(group_labels, levels = groups))
})

# 转换为目标格式的dataframe
result_df <- as.data.frame(count_matrix)

# 查看结果
print(result_df)

关键注意点

  • 先拿小样本测试逻辑,再处理全量6万列数据,避免内存溢出
  • cut函数的include.lowest = TRUE确保0被归入none组、1000被归入high组
  • 若数据存在NA值,可在cut前用na.omit(col)过滤,或在table中加入useNA = "always"统计缺失值

内容的提问来源于stack exchange,提问作者mfeldbauer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 10:24:46