如何在R中用mutate()生成无标点小写列验证物种识别准确性
解决iNaturalist观测数据标准化与识别正确性判断问题
针对你需要处理iNaturalist导出的2万条观测数据、标准化文本列并标记识别是否正确的需求,以下是完整的实现步骤:
1. 加载必要工具包
首先确保安装并加载dplyr包(用于数据处理管道和mutate函数):
install.packages("dplyr") # 首次使用需安装 library(dplyr)
2. 标准化三列文本
利用mutate批量处理species_guess、taxon_species_name、common_name三列,统一转为小写并去除所有标点、合并多余空格:
spiranthes <- spiranthes %>% mutate( # 标准化species_guess列 standardized_species_guess = gsub('[[:punct:] ]+', ' ', tolower(species_guess)) %>% str_squish(), # 标准化taxon_species_name列 standardized_taxon = gsub('[[:punct:] ]+', ' ', tolower(taxon_species_name)) %>% str_squish(), # 标准化common_name列 standardized_common = gsub('[[:punct:] ]+', ' ', tolower(common_name)) %>% str_squish() )
tolower():将整列文本转为全小写gsub('[[:punct:] ]+', ' ', ...):匹配所有标点和连续空格,替换为单个空格str_squish():去除字符串首尾空格并合并中间多余空格(比单纯用gsub更简洁)
3. 标记识别正确性
添加correct_identification列,判断标准化后的species_guess是否与标准化后的学名或俗名匹配:
spiranthes <- spiranthes %>% mutate( correct_identification = case_when( standardized_species_guess == standardized_taxon ~ TRUE, standardized_species_guess == standardized_common ~ TRUE, TRUE ~ FALSE ) )
如果需要支持模糊子串匹配,可替换为grepl()函数:
spiranthes <- spiranthes %>% mutate( correct_identification = case_when( grepl(standardized_species_guess, standardized_taxon, fixed = TRUE) ~ TRUE, grepl(standardized_species_guess, standardized_common, fixed = TRUE) ~ TRUE, TRUE ~ FALSE ) )
case_when():按条件分支赋值grepl(fixed = TRUE):进行精确子串匹配,避免正则模式干扰
4. 查看处理结果
运行以下代码查看处理后的数据表:
View(spiranthes)
内容的提问来源于stack exchange,提问作者Malaxis28
相关产品推荐
相关产品推荐

