如何按样本对R数据框物种丰度列求和并保留前21列?
按样本汇总物种相对丰度的dplyr解决方案
问题描述
我有一个大型R数据框,前21列为非生物变量(包含样本名),22-72列为物种相对丰度值。每个样本对应多行数据,且同一样本的非生物变量值完全一致。需要按样本对每个物种的相对丰度值求和,但使用dplyr的group_by和summarise时出现错误:明明ncol(df)返回72,却提示仅存在51列,无法索引超出范围的列。
原始数据示例
df <- data.frame( sample = c(1,1,1,1,1,1,1,2,2,2,2,2,2,2,3,3,3,3,3,3), var1 = c(3,3,3,3,3,3,3,7,7,7,7,7,7,7,2,2,2,2,2,2), var2 = c(4,4,4,4,4,4,4,42,42,42,42,42,42,42,2,2,2,2,2,2), species1 = c(0,0,0.05,0,0,0.02,0,0,0,0,0,0,0,0,0,0.001,0.02,0.03,0.001,0), species2 = c(0,0,0,0,0,0,0,0,0,0,0,0,0,0,0.001,0.002,0.03,0,0,0) )
期望结果
df_summed <- data.frame( sample = c(1, 2, 3), var1 = c(3, 7, 2), var2 = c(4, 42, 2), species1 = c(0.07, 0, 0.052), species2 = c(0, 0, 0.033) )
尝试的错误代码
df_summed <- df %>% group_by(across(1:21)) %>% summarise(across(22:ncol(df), sum), .groups = "drop")
错误信息
Caused by error in `across()`: ! Can't subset columns past the end. ℹ Locations 52, 53, 54, …, 71, and 72 don't exist. ℹ There are only 51 columns.
解决方案
错误原因
报错的核心是:在group_by分组后,summarise中的across函数会将列索引的基准默认设为非分组列(共72-21=51列),而你用22:ncol(df)引用的是原始数据框的列索引(22到72),这个范围远超当前可用的非分组列数量,导致索引越界。
可行方案
方案1:提前定义列范围,用all_of引用
提前把非生物变量列和物种列的索引存为变量,避免上下文导致的索引混乱:
# 定义列索引 non_bio_cols <- 1:21 species_cols <- 22:ncol(df) # 汇总计算 df_summed <- df %>% group_by(across(all_of(non_bio_cols))) %>% summarise(across(all_of(species_cols), sum), .groups = "drop")
方案2:排除分组列选择物种列
直接通过排除非生物变量列来选择物种列,无需关注具体索引:
df_summed <- df %>% group_by(across(1:21)) %>% summarise(across(-all_of(1:21), sum), .groups = "drop")
方案3:按列名模式选择物种列
如果物种列有统一命名规则(比如都以species开头),可以直接按名称匹配:
df_summed <- df %>% group_by(across(1:21)) %>% summarise(across(starts_with("species"), sum), .groups = "drop")
以上三种方案都能正确按样本汇总物种相对丰度,同时保留非生物变量的唯一值。
内容的提问来源于stack exchange,提问作者RobH
相关产品推荐
相关产品推荐

