You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何移除R数据框annotation列中Exon条目括号内的冗余文本

R语言处理Annotation列冗余信息的实现方案

需求说明

将数据框aa的annotation列中所有包含Exon字符串的条目统一简化为Exon,其余非Exon类条目保留原有内容,输出与示例Output_required列一致的结果。


实现方案

方案1:基础R实现

方法1:字符判断法(通用易读)

# 先转字符类型做判断,需要保留factor可后续转回
aa$annotation_clean <- ifelse(
  grepl("Exon", aa$annotation), 
  "Exon", # 若需要匹配示例中的带空格Exon,可改为"Exon "
  as.character(aa$annotation)
)
# 可选:转换回factor类型
aa$annotation_clean <- as.factor(aa$annotation_clean)

方法2:直接修改Factor Levels(效率更高,适合大数据集)

因为原数据的annotation列为factor类型,直接修改标签无需逐行判断:

levels(aa$annotation)[grepl("Exon", levels(aa$annotation))] <- "Exon"

方案2:Tidyverse体系实现

适合搭配管道流处理分析流程:

library(dplyr)
library(stringr)

aa <- aa %>%
  mutate(
    annotation_clean = case_when(
      str_detect(annotation, "Exon") ~ "Exon",
      TRUE ~ as.character(annotation)
    ) %>% as.factor() # 不需要保留factor可删除这行
  )

结果验证

处理完成后可执行以下代码验证与预期输出的匹配度:

# 忽略首尾空格的匹配验证
table(trimws(aa$annotation_clean) == trimws(aa$Output_required), useNA = "ifany")

内容的提问来源于stack exchange,提问作者PesKchan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 22:57:04