You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中替换文本列内两位字母加两位数字的术语

解决方法:替换文本中特定格式的术语

你可以通过正则表达式结合R的字符串处理函数实现需求,以下是两种常用方案:

方案1:使用tidyverse的stringr包(适配你的dplyr工作流)

利用str_replace_all函数批量替换所有符合格式的术语:

library(tidyverse)

# 构造示例数据
df_new <- data.frame(
  given_info=c('SA12 is given','he has his sa12',
         'she will get Sr15','why not having an ra31',
         'his tA23 is missing', 'pa12 is given'))

# 执行替换
df_new <- df_new %>%
  mutate(given_info = str_replace_all(given_info, "\\b[a-zA-Z]{2}\\d{2}\\b", "document"))

# 查看结果
df_new %>% select(given_info)

正则说明:

  • \\b:匹配单词边界,确保替换的是独立术语(不会误匹配包含该格式的长单词)
  • [a-zA-Z]{2}:匹配任意大小写的两位字母
  • \\d{2}:匹配两位数字

方案2:使用base R的gsub函数

如果不想加载额外包,直接用base R的gsub也能实现:

# 构造示例数据
df_new <- data.frame(
  given_info=c('SA12 is given','he has his sa12',
         'she will get Sr15','why not having an ra31',
         'his tA23 is missing', 'pa12 is given'))

# 执行替换(两种写法效果一致)
# 写法1:明确匹配大小写字母
df_new$given_info <- gsub("\\b[a-zA-Z]{2}\\d{2}\\b", "document", df_new$given_info)

# 写法2:通过ignore.case参数忽略大小写
# df_new$given_info <- gsub("\\b[a-z]{2}\\d{2}\\b", "document", df_new$given_info, ignore.case = TRUE)

# 查看结果
df_new %>% select(given_info)

两种方案执行后,都会得到你期望的结果:

given_info
1          document is given
2        he has his document
3      she will get document
4 why not having an document
5    his document is missing
6          document is given

内容的提问来源于stack exchange,提问作者Alex

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 17:40:41