You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用gsub提取字符串中的3位大写国家代码

问题:提取字符串中的3位大写国家代码

我有若干字符串,需要从中提取格式为3位大写字母的国家代码。

示例输入字符串:

mystring <- c(
  "Bloggs, Joe GBR London (1)/Bloggs, Joe London (2)",
  "Bloggs, Joe London (1)/Bloggs, Joe  GBR London (2)",
  "Bloggs, Joe London (1)/Bloggs, Joe London (2)",
  "Bloggs, Joe GBR London (1)/Bloggs, Joe GBR London (2)",
  "Bloggs, J-S GBR London (1)/Bloggs, J-S GBR London (2)"
)

期望输出:

GBR/
/GBR
/
GBR/GBR
GBR/GBR

无国家代码时留空即可,后续可自行处理

我尝试过两种方法都没成功:

  • 移除所有非大写字母的字符,但会保留姓名、地点中的大写字母,代码如下:
gsub("[^A-Z\"]", "", mystring)
  • 尝试移除非首尾大写的字母,也因姓名问题失败。

我觉得保留所有3位连续大写字母的思路可行,但写不出正确代码,初步思路如下,求正确写法或更稳健的实现方法:

gsub("[^A-Z$]{3}", "", mystring)

解决方案

可以用正则表达式精准匹配连续3位大写字母,结合字符串分割与拼接实现需求:

方法1:使用stringr包提取拼接(清晰易维护)

library(stringr)

# 定义处理单条字符串的函数
extract_country_codes <- function(x) {
  # 按/分割字符串为前后两部分
  parts <- str_split(x, "/")[[1]]
  # 对每一部分提取独立的3位大写字母,无匹配则留空
  codes <- sapply(parts, function(p) {
    code <- str_extract(p, "\\b[A-Z]{3}\\b")
    ifelse(is.na(code), "", code)
  })
  # 用/拼接结果
  paste(codes, collapse = "/")
}

# 批量处理所有字符串
sapply(mystring, extract_country_codes)

关键正则说明:

  • \\b:单词边界,确保匹配的是独立的3位大写字母(避免误匹配姓名/地名中连续3个大写的片段,比如姓名缩写)
  • [A-Z]{3}:精准匹配连续3位大写字母

方法2:基础R纯gsub实现(无需额外包)

如果不想加载第三方包,可通过多次正则替换实现:

# 第一步:提取每段中的3位大写代码,保留分隔符结构
temp <- gsub(".*?(\\b[A-Z]{3}\\b)?.*?(/|$)", "\\1\\2", mystring)
# 第二步:补全无代码位置的空值,对齐格式
result <- ifelse(grepl("^[^/]", temp) & !grepl("/$", temp), paste0(temp, "/"), temp)
result <- ifelse(grepl("/$", temp) & nchar(temp) == 1, temp, result)

内容的提问来源于stack exchange,提问作者Joe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 16:28:21