使用vegan包specaccum计算多采样点物种累积时遇维度错误求助
问题:使用vegan包specaccum函数(collector方法)计算物种累积曲线时触发维度错误
背景
数据集包含120个调查点,每个调查点每年开展4次调查(共480次)。每次15分钟调查按每3分钟时段(共5个时段)记录物种存在/缺失状态,目标是用vegan包的specaccum函数,以collector方法计算每个「调查点+调查日」组合的物种累积(即每新增一个3分钟时段的物种数量变化)。
当前实现代码
# Loading packages library(vegan) # setwd first Toronto_data=read.csv(file.choose()) # Converting columns that won't be used in the species accumulation curve into factors cols_to_convert <- c("SurveyDay", "TimeBlock") Toronto_data[cols_to_convert] <- lapply(Toronto_data[cols_to_convert], as.factor) # Split the data frame by multiple factors: Point ID and Number of Days Survey Occurred (~120 points x 4 survey dates) list_of_samples <- split(Toronto_data, interaction(Toronto_data$PointID, Toronto_data$SurveyDay)) # Remove empty data frames from split_data list_of_samples <- list_of_samples[sapply(list_of_samples, nrow) > 0] # Create function to apply specaccum function to numeric columns. Want to use "collector" rather than "exact but it will not work. specaccum_function <- function(data) { # Ensure that only numeric columns are passed numeric_data <- data[sapply(data, is.numeric)] # Check if numeric_data has more than one row and one column if (nrow(numeric_data) > 1 && ncol(numeric_data) > 1) { return(specaccum(numeric_data, "collector")) } else { stop("Not enough data to calculate species accumulation.") } } # Apply specaccum to each subset results <- lapply(list_of_samples, function(list_of_samples) { specaccum_function(list_of_samples) })
错误信息
Error in rowSums(apply(x[ind, , drop = FALSE], 2, cumsum) > 0) : 'x' must be an array of at least two dimensions
数据结构示例
structure(list(PointID = c(1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L), SurveyDay = c(1L, 1L, 1L, 1L, 1L, 2L, 2L, 2L, 2L, 2L, 3L, 3L, 3L, 3L, 3L, 4L, 4L, 4L, 4L, 4L), TimeBlock = c(1L, 2L, 3L, 4L, 5L, 1L, 2L, 3L, 4L, 5L, 1L, 2L, 3L, 4L, 5L, 1L, 2L, 3L, 4L, 5L), SpeciesA = c(1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L), SpeciesB = c(1L, 1L, 1L, 1L, 0L, 0L, 0L, 0L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 0L, 1L, 1L), SpeciesC = c(0L, 0L, 0L, 0L, 0L, 0L, 1L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L), SpeciesD = c(1L, 0L, 0L, 0L, 0L, 0L, 1L, 1L, 1L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L), SpeciesE = c(0L, 0L, 0L, 0L, 0L, 0L, 1L, 1L, 0L, 1L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L), SpeciesF = c(0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L)), class = "data.frame", row.names = c(NA, -20L))
问题原因与修正方案
错误原因
原代码中numeric_data <- data[sapply(data, is.numeric)]会把PointID(整数型,属于numeric)也纳入计算,但specaccum要求输入的是纯物种存在/缺失矩阵(行=调查时段,列=物种)。每个「调查点+调查日」子集中的PointID是固定值,这一列会干扰collector方法的累积计算逻辑,导致内部处理时出现维度异常。
修正代码
修改specaccum_function,明确筛选物种列(这里假设物种列以Species开头,可根据实际列名调整):
specaccum_function <- function(data) { # 筛选物种列(仅保留列名以Species开头的列) species_data <- data[, grepl("^Species", colnames(data))] # 检查数据维度:至少2个时段(行)和2个物种(列) if (nrow(species_data) > 1 && ncol(species_data) > 1) { # 转换为矩阵格式,避免dataframe可能的兼容性问题 return(specaccum(as.matrix(species_data), method = "collector")) } else { # 用warning替代stop,避免单个无效子集中断整个批量计算 warning("Not enough data to calculate species accumulation for this subset.") return(NULL) } }
补充说明
- 如果物种列的命名规则不是
Species开头,可以手动指定列名,比如species_data <- data[, c("SpeciesA", "SpeciesB", "SpeciesC", ...)] - 转换为矩阵格式是为了确保
specaccum处理时的稳定性,避免dataframe的潜在问题
内容的提问来源于stack exchange,提问作者pprest00
相关产品推荐
相关产品推荐

