You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中按同观测条件移除未识别昆虫数据行?

用dplyr处理昆虫观测数据集的清洗需求

需求拆解

  • 分组依据:按SiteID、Year、Round、Treatment、Plant_species这5个字段分组
  • 规则1:如果组内存在已识别的Solitary bee,直接删掉组内所有Unidentified Hymenoptera行
  • 规则2:如果组内同属(genus)有已识别的具体物种,删掉组内类似Chelostoma spp.这种未识别种级的行

示例数据集

先构造一份符合需求的示例数据,方便测试代码:

library(dplyr)

obs_data <- tibble(
  SiteID = c("Site1", "Site1", "Site1", "Site2", "Site2", "Site3", "Site3", "Site3"),
  Year = c(2021, 2021, 2021, 2021, 2021, 2021, 2021, 2021),
  Round = c("Round3", "Round3", "Round3", "Round3", "Round3", "Round3", "Round3", "Round3"),
  Treatment = c("TreatmentP", "TreatmentP", "TreatmentP", "TreatmentP", "TreatmentP", "TreatmentP", "TreatmentP", "TreatmentP"),
  Plant_species = c("plant2", "plant2", "plant2", "plant2", "plant2", "plant3", "plant3", "plant3"),
  Species = c("Osmia bicornis", "Unidentified Hymenoptera", "Chelostoma florisomne", 
              "Unidentified Hymenoptera", "Fly", 
              "Chelostoma spp.", "Chelostoma florisomne", "Unidentified Hymenoptera"),
  Type = c("Solitary bee", "Hymenoptera", "Solitary bee", 
           "Hymenoptera", "Diptera", 
           "Solitary bee", "Solitary bee", "Hymenoptera")
)

代码实现(分步解释)

第一步:处理规则1

先给每组标记是否存在已识别的Solitary bee,再过滤掉需要删除的行:

# 加载字符串处理工具包
library(stringr)

# 处理规则1
step1_clean <- obs_data %>%
  group_by(SiteID, Year, Round, Treatment, Plant_species) %>%
  # 标记当前组是否有已识别的Solitary bee
  mutate(has_solitary = any(Type == "Solitary bee")) %>%
  # 过滤:如果是Unidentified Hymenoptera且组内有Solitary bee,就删掉
  filter(!(Species == "Unidentified Hymenoptera" & has_solitary)) %>%
  ungroup()

第二步:处理规则2

需要先从物种名里提取属名,再标记同属是否有已识别物种,最后过滤:

# 处理规则2
final_clean <- step1_clean %>%
  # 提取属名(物种名的第一个单词)
  mutate(genus = word(Species, 1),
         # 标记是否是未识别种级(以spp.结尾)
         is_spp = str_detect(Species, "spp\\.$")) %>%
  # 按分组+属名再次分组,判断同属有没有已识别物种
  group_by(SiteID, Year, Round, Treatment, Plant_species, genus) %>%
  mutate(has_identified_sp = any(!is_spp)) %>%
  # 过滤:如果是spp.且同属有已识别物种,就删掉
  filter(!(is_spp & has_identified_sp)) %>%
  # 删掉临时生成的辅助字段
  select(-has_solitary, -genus, -is_spp, -has_identified_sp) %>%
  ungroup()

一步到位的合并写法

如果不想分步,也可以把两个规则放到同一个管道里:

final_clean <- obs_data %>%
  group_by(SiteID, Year, Round, Treatment, Plant_species) %>%
  mutate(has_solitary = any(Type == "Solitary bee")) %>%
  filter(!(Species == "Unidentified Hymenoptera" & has_solitary)) %>%
  mutate(genus = word(Species, 1),
         is_spp = str_detect(Species, "spp\\.$")) %>%
  group_by(genus, .add = TRUE) %>%
  mutate(has_identified_sp = any(!is_spp)) %>%
  filter(!(is_spp & has_identified_sp)) %>%
  select(-has_solitary, -genus, -is_spp, -has_identified_sp) %>%
  ungroup()

验证结果

运行代码后,示例数据的处理结果符合需求:

  • Site1的Unidentified Hymenoptera被移除(组内有Solitary bee)
  • Site2的Unidentified Hymenoptera保留(组内无Solitary bee)
  • Site3的Chelostoma spp.被移除(同属有已识别的Chelostoma florisomne)

内容的提问来源于stack exchange,提问作者Vevey

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 11:13:17