You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用数据框中的字典替换字符串中的特定双词分类名称

批量替换字符串列表中的双词分类名称

问题背景

现有存储为数据框(ddf)的双词分类名称字典,包含old列(待替换的旧名称)和new列(替换后的新名称)。需要在字符串列表(listn)中精准查找并替换这些旧名称:旧名称是字符串的一部分,可能重复出现或完全不出现,前后会伴随不同长度的其他字符串。

示例数据

# 示例字符串列表,实际数据超过30万条
listn <- c("AB001440.1.1538 Pseudomonas coronafaciens pv. atropurpurea", "HG530070.1.1349 Trueperella pyogenes",                       
          "ET631036.837.2346 Jonquetella anthropi", "AB001448.1.1538 Pseudomonas savastanoi pv. phaseolicola",   
          "HG530249.1.1462 Paucibacter toxinivorans", "HG530235.1.1493 Paucibacter toxinivorans",                  
          "AB001781.1.1507 Chlamydia psittaci", "AB001785.1.1507 Chlamydia felis",                            
          "AB001804.1.1507 Chlamydia psittaci", "AB001794.1.1507 Chlamydia psittaci")   

# 示例替换字典,实际包含超过400条记录
ddf <- data.frame(old = c("Lactobacillus casei", "Trueperella pyogenes", "Pseudomonas savastanoi"), 
                  new = c("Newbacillus casei", "Newperella pyogenes", "Newudomonas savastanoi"))

尝试过的无效方法

以下两种方法均未实现预期的批量替换效果:

# 使用stringr::str_replace_all(错误用法)
listn2 <- stringr::str_replace_all(string = listn,
                                   pattern = ddf$old,
                                   replacement = ddf$new)  

# 使用基础R的gsub(错误用法)
list2 <- gsub(ddf$old, ddf$new, listn)

此外,尝试ChatGPT生成的脚本也无法解决问题(脚本仅支持单词替换,不兼容双词名称):

# 示例输入数据
selected_strings <- c("John likes apples", "Mary eats bananas", "David enjoys grapes")
dictionary <- data.frame(old_name = c("John", "Mary", "David"),
                         new_name = c("Peter", "Alice", "Michael"),
                         stringsAsFactors = FALSE)

# 用于替换字符串中名称的函数(仅支持单词)
replace_names <- function(strings, dictionary) {
  # 将字符串拆分为单个单词
  words <- unlist(strsplit(strings, " "))
  
  # 使用字典替换名称
  replaced_words <- words
  for (i in seq_len(nrow(dictionary))) {
    replaced_words[words == dictionary$old_name[i]] <- dictionary$new_name[i]
  }
  
  # 重构修改后的字符串
  modified_strings <- sapply(strsplit(strings, " "), function(x) paste(replaced_words[x], collapse = " "))
  
  return(modified_strings)
}

# 调用函数替换名称
modified_strings <- replace_names(selected_strings, dictionary)

# 打印修改后的字符串
print(modified_strings)

有效解决方案

方法1:stringr批量替换(推荐,适合大规模数据)

核心是将字典转换为命名向量,这是stringr::str_replace_all支持批量替换的标准用法,效率极高,适合处理30万条级别的数据。

# 将字典转换为命名向量:旧名称为键,新名称为值
name_mapping <- setNames(ddf$new, ddf$old)

# 执行批量替换
listn_updated <- stringr::str_replace_all(listn, name_mapping)

方法2:基础R循环替换(适合小批量数据)

若不想依赖第三方包,可使用gsub循环逐条处理替换规则:

listn_updated <- listn
for (i in seq_len(nrow(ddf))) {
  listn_updated <- gsub(pattern = ddf$old[i], 
                        replacement = ddf$new[i], 
                        x = listn_updated)
}

特殊情况处理

如果旧名称中包含正则特殊字符(如.、()等),需关闭正则匹配以避免错误:

# 方法1添加fixed参数
listn_updated <- stringr::str_replace_all(listn, stringr::fixed(name_mapping))

# 方法2添加fixed参数
listn_updated <- gsub(pattern = ddf$old[i], 
                      replacement = ddf$new[i], 
                      x = listn_updated,
                      fixed = TRUE)

内容的提问来源于stack exchange,提问作者mschmidt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 02:05:54