dplyr多级分组求和报错问题求助
鸟类调查数据分组求和问题解决
问题描述
我在处理鸟类调查数据的分组求和时遇到问题:数据包含多个样地(site)、调查点(point),每年开展1-2次调查(visit)统计鸟类物种(species)数量(abundance),但因距离区间划分,同一物种在单次调查中会有多条记录。为估算物种丰富度,需要对单次调查的同一物种数量求和。
我用dplyr编写的代码:
bird.group <- group_by(bird.use, site, point, year, visit, species, sum.abund=sum(abundance))
运行后报错:
Error in
group_by():
! Problem adding computed columns.
Caused by error inmutate():
! Problem while computingsum.abund = sum(abundance).
Caused by error insum():
! invalid 'type' (character) of argument
Runrlang::last_error()to see where the error occurred.
附示例数据:
structure(list(site = c("BethWater", "BethWater", "BethWater", "BethWater", "BethWater", "BethWater", "BethWater", "BethWater", "BethWater", "BethWater", "BethWater", "BethWater", "BethWater", "BethWater", "BethWater", "BethWater", "BethWater", "BethWater", "BethWater", "BethWater", "BethWater", "BethWater", "BethWater", "BethWater", "BethWater"), point = c("1", "1", "1", "1", "1", "1", "1", "1", "1", "1", "1", "1", "1", "1", "1", "1", "1", "1", "2", "2", "2", "2", "2", "2", "2"), year = c(2020L, 2020L, 2020L, 2020L, 2020L, 2020L, 2020L, 2020L, 2020L, 2020L, 2020L, 2020L, 2020L, 2020L, 2020L, 2020L, 2020L, 2020L, 2020L, 2020L, 2020L, 2020L, 2020L, 2020L, 2020L), visit = c(1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L), species = c("BAWW", "CSWA", "CSWA", "COYE", "EATO", "EATO", "EATO", "FISP", "GRCA", "GRCA", "MODO", "OVEN", "REVI", "REVI", "PRAW", "AMRE", "ALFL", "YBCU", "BAWW", "BAWW", "CSWA", "COYE", "COYE", "EATO", "EATO"), abundance = c("2", "1", "1", "1", "1", "2", "1", "1", "1", "2", "1", "1", "1", "2", "1", "2", "1", "1", "1", "1", "1", "1", "1", "2", "2")), row.names = c(NA, 25L), class = "data.frame")
问题分析与修正
1. 核心错误原因
group_by()仅用于指定分组字段,不能直接在函数内执行求和计算,求和需配合summarise()或mutate()完成- 示例数据中
abundance列是字符类型(值带引号),sum()函数无法对字符类型进行运算,需先转换为数值型
2. 修正后的代码
library(dplyr) # 第一步:将字符型的abundance转为数值型 bird.use <- bird.use %>% mutate(abundance = as.numeric(abundance)) # 第二步:分组并计算每个组的总数量 bird.group <- bird.use %>% group_by(site, point, year, visit, species) %>% summarise(sum.abund = sum(abundance), .groups = "drop")
代码说明
mutate(abundance = as.numeric(abundance)):转换数据类型,确保求和函数能正常工作group_by(...):仅定义分组维度,不包含计算逻辑summarise(sum.abund = sum(abundance), .groups = "drop"):对每个分组执行求和,.groups = "drop"参数用于取消分组状态,返回普通数据框(若需保留分组可省略该参数)
内容的提问来源于stack exchange,提问作者Jeff Stratford
相关产品推荐
相关产品推荐

