如何在R中对相机数据集下采样以确定新检测率
R中相机数据集下采样与相机数量优化指南
1. 先整理数据
确保数据集包含camera_id(相机ID)、sub_region(子区域)、species_detection_rate(物种检测率)这几个核心字段,先清理缺失值:
library(dplyr) # 假设你的数据集名为camera_data camera_data <- camera_data %>% select(camera_id, sub_region, species_detection_rate) %>% na.omit()
2. 下采样的两种实用方法
分层随机下采样(推荐)
按子区域分层随机保留一定比例的相机,避免某区域相机被过度移除,保证区域结构的代表性:
# 示例:每个子区域保留70%的相机,可调整prop参数(0-1之间) downsampled_data <- camera_data %>% group_by(sub_region) %>% slice_sample(prop = 0.7) %>% ungroup()
基于检测率的定向下采样
如果想优先保留检测率高的相机,或者剔除检测率极低的,可按检测率排序后采样:
# 示例:每个子区域保留检测率前60%的相机 downsampled_data <- camera_data %>% group_by(sub_region) %>% arrange(desc(species_detection_rate)) %>% slice_head(prop = 0.6) %>% ungroup()
3. 评估下采样后的检测率变化
对比原始数据和下采样数据的子区域检测率,判断是否在可接受范围内:
# 计算原始数据各子区域的平均检测率和相机数量 original_summary <- camera_data %>% group_by(sub_region) %>% summarise(avg_detection_original = mean(species_detection_rate), total_cameras_original = n()) # 计算下采样后的数据指标 downsampled_summary <- downsampled_data %>% group_by(sub_region) %>% summarise(avg_detection_down = mean(species_detection_rate), total_cameras_down = n()) # 合并对比,计算检测率下降百分比 comparison <- left_join(original_summary, downsampled_summary, by = "sub_region") %>% mutate(detection_drop_percent = (avg_detection_original - avg_detection_down)/avg_detection_original * 100)
如果detection_drop_percent在你设定的阈值内(比如低于10%),说明对应比例的相机减少是可行的。
4. 批量测试不同采样比例
写个循环自动测试多个保留比例,快速找到最优解:
# 定义要测试的保留比例 test_props <- c(0.9, 0.8, 0.7, 0.6, 0.5) results <- list() for(p in test_props){ # 按当前比例下采样 temp_down <- camera_data %>% group_by(sub_region) %>% slice_sample(prop = p) %>% ungroup() # 统计下采样后的指标 temp_summary <- temp_down %>% group_by(sub_region) %>% summarise(avg_detection = mean(species_detection_rate), camera_count = n()) # 和原始数据合并对比 temp_compare <- left_join(temp_summary, original_summary, by = "sub_region") %>% mutate(drop_rate = (avg_detection_original - avg_detection)/avg_detection_original * 100, sample_proportion = p) results[[as.character(p)]] <- temp_compare } # 合并所有结果,查看不同比例下的平均检测率下降情况 all_results <- bind_rows(results) all_results %>% group_by(sample_proportion) %>% summarise(mean_drop_rate = mean(drop_rate, na.rm = TRUE))
内容的提问来源于stack exchange,提问作者Sierra McMurry
相关产品推荐
相关产品推荐

