You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中用mutate()生成无标点小写列验证物种识别准确性

解决iNaturalist观测数据标准化与识别正确性判断问题

针对你需要处理iNaturalist导出的2万条观测数据、标准化文本列并标记识别是否正确的需求,以下是完整的实现步骤:

1. 加载必要工具包

首先确保安装并加载dplyr包(用于数据处理管道和mutate函数):

install.packages("dplyr") # 首次使用需安装
library(dplyr)

2. 标准化三列文本

利用mutate批量处理species_guess、taxon_species_name、common_name三列,统一转为小写并去除所有标点、合并多余空格:

spiranthes <- spiranthes %>%
  mutate(
    # 标准化species_guess列
    standardized_species_guess = gsub('[[:punct:] ]+', ' ', tolower(species_guess)) %>% str_squish(),
    # 标准化taxon_species_name列
    standardized_taxon = gsub('[[:punct:] ]+', ' ', tolower(taxon_species_name)) %>% str_squish(),
    # 标准化common_name列
    standardized_common = gsub('[[:punct:] ]+', ' ', tolower(common_name)) %>% str_squish()
  )
  • tolower():将整列文本转为全小写
  • gsub('[[:punct:] ]+', ' ', ...):匹配所有标点和连续空格,替换为单个空格
  • str_squish():去除字符串首尾空格并合并中间多余空格(比单纯用gsub更简洁)

3. 标记识别正确性

添加correct_identification列,判断标准化后的species_guess是否与标准化后的学名或俗名匹配:

spiranthes <- spiranthes %>%
  mutate(
    correct_identification = case_when(
      standardized_species_guess == standardized_taxon ~ TRUE,
      standardized_species_guess == standardized_common ~ TRUE,
      TRUE ~ FALSE
    )
  )

如果需要支持模糊子串匹配,可替换为grepl()函数:

spiranthes <- spiranthes %>%
  mutate(
    correct_identification = case_when(
      grepl(standardized_species_guess, standardized_taxon, fixed = TRUE) ~ TRUE,
      grepl(standardized_species_guess, standardized_common, fixed = TRUE) ~ TRUE,
      TRUE ~ FALSE
    )
  )
  • case_when():按条件分支赋值
  • grepl(fixed = TRUE):进行精确子串匹配,避免正则模式干扰

4. 查看处理结果

运行以下代码查看处理后的数据表:

View(spiranthes)

内容的提问来源于stack exchange,提问作者Malaxis28

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 19:40:31