如何用Base R或stringr从指定格式字符串提取姓名与尊称?
解决方案
使用stringr包
方法1:正则精准提取
直接用正则表达式分别匹配姓名(前两个单词)和尊称(剩余部分):
library(stringr) my_str <- c("John Smith The Last", "Jane Smith The Best") # 提取姓名 names <- str_extract(my_str, "^\\w+ \\w+") # 提取尊称 titles <- str_extract(my_str, "(?<=^\\w+ \\w+ ).*") # 输出结果 names #> [1] "John Smith" "Jane Smith" titles #> [1] "The Last" "The Best"
方法2:拆分后重组姓名
基于你尝试过的str_split,拆分后将前两个元素合并为姓名:
split_result <- str_split(my_str, " ", n = 3) names <- sapply(split_result, function(x) paste(x[1], x[2])) titles <- sapply(split_result, function(x) x[3])
使用Base R
方法1:strsplit+列表处理
拆分后按需提取并重组内容,即使尊称包含更多单词也适用:
my_str <- c("John Smith The Last", "Jane Smith The Best") split_result <- strsplit(my_str, " ", fixed = TRUE) # 提取姓名:合并前两个元素 names <- sapply(split_result, function(x) paste(x[1], x[2])) # 提取尊称:合并除前两个外的所有元素 titles <- sapply(split_result, function(x) paste(x[-c(1,2)], collapse = " "))
方法2:regmatches正则提取
用Base R原生的正则匹配工具实现提取:
# 匹配姓名的正则 name_pattern <- "^\\w+ \\w+" names <- regmatches(my_str, regexpr(name_pattern, my_str)) # 匹配尊称的正则(需开启perl模式支持正向预查) title_pattern <- "(?<=^\\w+ \\w+ ).*" titles <- regmatches(my_str, regexpr(title_pattern, my_str, perl = TRUE))
内容的提问来源于stack exchange,提问作者maurobio
相关产品推荐
相关产品推荐

