如何为数据框添加缺失物种的零TPH观测值并计算标准误?
为缺失树种添加TPH=0的观测值并计算统计量
我正在处理一份站点树种密度(每公顷树木数,TPH)的数据集。为计算所有站点的平均密度,需要给那些在某站点不存在但其他站点存在的树种添加TPH值为0的观测值。以下是模拟我所用数据框的代码:
# Define number of rows and levels for the grouping variable set.seed(120) n_rows <- 10 site_levels <- c("A", "B", "C", "D") # Create a map of sites and species that can be absent absent_species <- list( North = c("Quercus alba", "Betula papyrifera"), South = c("Pinus strobus", "Tsuga canadensis"), East = c("Acer rubrum"), West = c("Quercus alba", "Tsuga canadensis") ) # Define species pool and pre-fill empty site vectors species_pool <- c("Acer rubrum", "Quercus alba", "Pinus strobus", "Tsuga canadensis", "Betula papyrifera") site_species <- lapply(site_levels, function(site) character(0)) # Simulate Site column data <- data.frame(Site = sample(site_levels, size = n_rows, replace = TRUE)) # Loop through rows and assign unique species per site for (i in 1:n_rows) { site <- data$Site[i] absent_list <- absent_species[[site]] species_pool_filtered <- setdiff(species_pool, absent_list) # Check if all species have been used at this site if (length(site_species[[site]]) == length(species_pool_filtered)) { # No more species available, skip this row next } # Choose a random species from the filtered pool species <- sample(species_pool_filtered, size = 1, replace = FALSE) # Assign species and add it to the site's list data$Species[i] <- species site_species[[site]] <- c(site_species[[site]], species) } # Simulate tree densities with some variation by site data$TPH <- rnorm(n_rows, mean = c(500, 250, 100, 350)[match(data$Site, site_levels)], sd = c(100, 50, 25, 75)[match(data$Site, site_levels)]) # Print the simulated dataframe print(data)
简便实现方法
方法一:使用tidyverse的complete()函数
这是最简洁的方案,tidyr::complete()可以直接生成所有站点与树种的组合,并将缺失的TPH值填充为0:
library(tidyverse) # 生成全组合并填充缺失TPH为0 full_data <- data %>% complete(Site, Species = species_pool, fill = list(TPH = 0)) # 查看补全后的数据集 print(full_data)
补全后即可直接计算每个树种的平均密度和标准误:
# 按树种分组计算统计量 summary_stats <- full_data %>% group_by(Species) %>% summarise( 平均TPH = mean(TPH), 标准误 = sd(TPH)/sqrt(n()) ) print(summary_stats)
方法二:使用Base R实现
如果不想加载额外包,可以用expand.grid生成全组合后合并数据,再填充缺失值:
# 生成所有站点和树种的全组合 full_combinations <- expand.grid(Site = site_levels, Species = species_pool, stringsAsFactors = FALSE) # 合并原数据与全组合,缺失TPH填充为0 full_data_base <- merge(full_combinations, data, by = c("Site", "Species"), all.x = TRUE) full_data_base$TPH[is.na(full_data_base$TPH)] <- 0 # 查看补全后的数据集 print(full_data_base)
同样可以计算统计量:
# 按树种分组计算均值和标准误 summary_stats_base <- aggregate(TPH ~ Species, data = full_data_base, function(x) { c(均值 = mean(x), 标准误 = sd(x)/sqrt(length(x))) }) # 整理结果格式为数据框 summary_stats_base <- do.call(data.frame, summary_stats_base) print(summary_stats_base)
内容的提问来源于stack exchange,提问作者River
相关产品推荐
相关产品推荐

