如何在R中按同观测条件移除未识别昆虫数据行?
用dplyr处理昆虫观测数据集的清洗需求
需求拆解
- 分组依据:按
SiteID、Year、Round、Treatment、Plant_species这5个字段分组 - 规则1:如果组内存在已识别的Solitary bee,直接删掉组内所有
Unidentified Hymenoptera行 - 规则2:如果组内同属(genus)有已识别的具体物种,删掉组内类似
Chelostoma spp.这种未识别种级的行
示例数据集
先构造一份符合需求的示例数据,方便测试代码:
library(dplyr) obs_data <- tibble( SiteID = c("Site1", "Site1", "Site1", "Site2", "Site2", "Site3", "Site3", "Site3"), Year = c(2021, 2021, 2021, 2021, 2021, 2021, 2021, 2021), Round = c("Round3", "Round3", "Round3", "Round3", "Round3", "Round3", "Round3", "Round3"), Treatment = c("TreatmentP", "TreatmentP", "TreatmentP", "TreatmentP", "TreatmentP", "TreatmentP", "TreatmentP", "TreatmentP"), Plant_species = c("plant2", "plant2", "plant2", "plant2", "plant2", "plant3", "plant3", "plant3"), Species = c("Osmia bicornis", "Unidentified Hymenoptera", "Chelostoma florisomne", "Unidentified Hymenoptera", "Fly", "Chelostoma spp.", "Chelostoma florisomne", "Unidentified Hymenoptera"), Type = c("Solitary bee", "Hymenoptera", "Solitary bee", "Hymenoptera", "Diptera", "Solitary bee", "Solitary bee", "Hymenoptera") )
代码实现(分步解释)
第一步:处理规则1
先给每组标记是否存在已识别的Solitary bee,再过滤掉需要删除的行:
# 加载字符串处理工具包 library(stringr) # 处理规则1 step1_clean <- obs_data %>% group_by(SiteID, Year, Round, Treatment, Plant_species) %>% # 标记当前组是否有已识别的Solitary bee mutate(has_solitary = any(Type == "Solitary bee")) %>% # 过滤:如果是Unidentified Hymenoptera且组内有Solitary bee,就删掉 filter(!(Species == "Unidentified Hymenoptera" & has_solitary)) %>% ungroup()
第二步:处理规则2
需要先从物种名里提取属名,再标记同属是否有已识别物种,最后过滤:
# 处理规则2 final_clean <- step1_clean %>% # 提取属名(物种名的第一个单词) mutate(genus = word(Species, 1), # 标记是否是未识别种级(以spp.结尾) is_spp = str_detect(Species, "spp\\.$")) %>% # 按分组+属名再次分组,判断同属有没有已识别物种 group_by(SiteID, Year, Round, Treatment, Plant_species, genus) %>% mutate(has_identified_sp = any(!is_spp)) %>% # 过滤:如果是spp.且同属有已识别物种,就删掉 filter(!(is_spp & has_identified_sp)) %>% # 删掉临时生成的辅助字段 select(-has_solitary, -genus, -is_spp, -has_identified_sp) %>% ungroup()
一步到位的合并写法
如果不想分步,也可以把两个规则放到同一个管道里:
final_clean <- obs_data %>% group_by(SiteID, Year, Round, Treatment, Plant_species) %>% mutate(has_solitary = any(Type == "Solitary bee")) %>% filter(!(Species == "Unidentified Hymenoptera" & has_solitary)) %>% mutate(genus = word(Species, 1), is_spp = str_detect(Species, "spp\\.$")) %>% group_by(genus, .add = TRUE) %>% mutate(has_identified_sp = any(!is_spp)) %>% filter(!(is_spp & has_identified_sp)) %>% select(-has_solitary, -genus, -is_spp, -has_identified_sp) %>% ungroup()
验证结果
运行代码后,示例数据的处理结果符合需求:
- Site1的
Unidentified Hymenoptera被移除(组内有Solitary bee) - Site2的
Unidentified Hymenoptera保留(组内无Solitary bee) - Site3的
Chelostoma spp.被移除(同属有已识别的Chelostoma florisomne)
内容的提问来源于stack exchange,提问作者Vevey
相关产品推荐
相关产品推荐

