You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为数据框添加缺失物种的零TPH观测值并计算标准误?

为缺失树种添加TPH=0的观测值并计算统计量

我正在处理一份站点树种密度(每公顷树木数,TPH)的数据集。为计算所有站点的平均密度,需要给那些在某站点不存在但其他站点存在的树种添加TPH值为0的观测值。以下是模拟我所用数据框的代码:

# Define number of rows and levels for the grouping variable
set.seed(120)
n_rows <- 10
site_levels <- c("A", "B", "C", "D")

# Create a map of sites and species that can be absent
absent_species <- list(
  North = c("Quercus alba", "Betula papyrifera"),
  South = c("Pinus strobus", "Tsuga canadensis"),
  East = c("Acer rubrum"),
  West = c("Quercus alba", "Tsuga canadensis")
)

# Define species pool and pre-fill empty site vectors
species_pool <- c("Acer rubrum", "Quercus alba", "Pinus strobus", "Tsuga canadensis", "Betula papyrifera")
site_species <- lapply(site_levels, function(site) character(0))

# Simulate Site column
data <- data.frame(Site = sample(site_levels, size = n_rows, replace = TRUE))

# Loop through rows and assign unique species per site
for (i in 1:n_rows) {
  site <- data$Site[i]
  absent_list <- absent_species[[site]]
  species_pool_filtered <- setdiff(species_pool, absent_list)
  
  # Check if all species have been used at this site
  if (length(site_species[[site]]) == length(species_pool_filtered)) {
    # No more species available, skip this row
    next
  }
  
  # Choose a random species from the filtered pool
  species <- sample(species_pool_filtered, size = 1, replace = FALSE)
  
  # Assign species and add it to the site's list
  data$Species[i] <- species
  site_species[[site]] <- c(site_species[[site]], species)
}

# Simulate tree densities with some variation by site
data$TPH <- rnorm(n_rows, 
                  mean = c(500, 250, 100, 350)[match(data$Site, site_levels)],
                  sd = c(100, 50, 25, 75)[match(data$Site, site_levels)])

# Print the simulated dataframe
print(data)

简便实现方法

方法一:使用tidyverse的complete()函数

这是最简洁的方案,tidyr::complete()可以直接生成所有站点与树种的组合,并将缺失的TPH值填充为0:

library(tidyverse)

# 生成全组合并填充缺失TPH为0
full_data <- data %>%
  complete(Site, Species = species_pool, fill = list(TPH = 0))

# 查看补全后的数据集
print(full_data)

补全后即可直接计算每个树种的平均密度和标准误:

# 按树种分组计算统计量
summary_stats <- full_data %>%
  group_by(Species) %>%
  summarise(
    平均TPH = mean(TPH),
    标准误 = sd(TPH)/sqrt(n())
  )

print(summary_stats)

方法二:使用Base R实现

如果不想加载额外包,可以用expand.grid生成全组合后合并数据,再填充缺失值:

# 生成所有站点和树种的全组合
full_combinations <- expand.grid(Site = site_levels, Species = species_pool, stringsAsFactors = FALSE)

# 合并原数据与全组合,缺失TPH填充为0
full_data_base <- merge(full_combinations, data, by = c("Site", "Species"), all.x = TRUE)
full_data_base$TPH[is.na(full_data_base$TPH)] <- 0

# 查看补全后的数据集
print(full_data_base)

同样可以计算统计量:

# 按树种分组计算均值和标准误
summary_stats_base <- aggregate(TPH ~ Species, data = full_data_base, function(x) {
  c(均值 = mean(x), 标准误 = sd(x)/sqrt(length(x)))
})

# 整理结果格式为数据框
summary_stats_base <- do.call(data.frame, summary_stats_base)
print(summary_stats_base)

内容的提问来源于stack exchange,提问作者River

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 00:32:51