You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中对相机数据集下采样以确定新检测率

R中相机数据集下采样与相机数量优化指南

1. 先整理数据

确保数据集包含camera_id(相机ID)、sub_region(子区域)、species_detection_rate(物种检测率)这几个核心字段,先清理缺失值:

library(dplyr)
# 假设你的数据集名为camera_data
camera_data <- camera_data %>%
  select(camera_id, sub_region, species_detection_rate) %>%
  na.omit()

2. 下采样的两种实用方法

分层随机下采样(推荐)

按子区域分层随机保留一定比例的相机,避免某区域相机被过度移除,保证区域结构的代表性:

# 示例:每个子区域保留70%的相机,可调整prop参数(0-1之间)
downsampled_data <- camera_data %>%
  group_by(sub_region) %>%
  slice_sample(prop = 0.7) %>%
  ungroup()

基于检测率的定向下采样

如果想优先保留检测率高的相机,或者剔除检测率极低的,可按检测率排序后采样:

# 示例:每个子区域保留检测率前60%的相机
downsampled_data <- camera_data %>%
  group_by(sub_region) %>%
  arrange(desc(species_detection_rate)) %>%
  slice_head(prop = 0.6) %>%
  ungroup()

3. 评估下采样后的检测率变化

对比原始数据和下采样数据的子区域检测率,判断是否在可接受范围内:

# 计算原始数据各子区域的平均检测率和相机数量
original_summary <- camera_data %>%
  group_by(sub_region) %>%
  summarise(avg_detection_original = mean(species_detection_rate),
            total_cameras_original = n())

# 计算下采样后的数据指标
downsampled_summary <- downsampled_data %>%
  group_by(sub_region) %>%
  summarise(avg_detection_down = mean(species_detection_rate),
            total_cameras_down = n())

# 合并对比,计算检测率下降百分比
comparison <- left_join(original_summary, downsampled_summary, by = "sub_region") %>%
  mutate(detection_drop_percent = (avg_detection_original - avg_detection_down)/avg_detection_original * 100)

如果detection_drop_percent在你设定的阈值内(比如低于10%),说明对应比例的相机减少是可行的。

4. 批量测试不同采样比例

写个循环自动测试多个保留比例,快速找到最优解:

# 定义要测试的保留比例
test_props <- c(0.9, 0.8, 0.7, 0.6, 0.5)
results <- list()

for(p in test_props){
  # 按当前比例下采样
  temp_down <- camera_data %>%
    group_by(sub_region) %>%
    slice_sample(prop = p) %>%
    ungroup()
  
  # 统计下采样后的指标
  temp_summary <- temp_down %>%
    group_by(sub_region) %>%
    summarise(avg_detection = mean(species_detection_rate),
              camera_count = n())
  
  # 和原始数据合并对比
  temp_compare <- left_join(temp_summary, original_summary, by = "sub_region") %>%
    mutate(drop_rate = (avg_detection_original - avg_detection)/avg_detection_original * 100,
           sample_proportion = p)
  
  results[[as.character(p)]] <- temp_compare
}

# 合并所有结果,查看不同比例下的平均检测率下降情况
all_results <- bind_rows(results)
all_results %>%
  group_by(sample_proportion) %>%
  summarise(mean_drop_rate = mean(drop_rate, na.rm = TRUE))

内容的提问来源于stack exchange,提问作者Sierra McMurry

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 01:57:50