如何在R语言中从指定字符串提取带逗号的目标子串
在R语言中提取「城市, 州」格式的子串
我拥有如下字符串列表:
Woman Killed in Uber Accident in Goleta, California Sacramento, California Man Dies in Lyft Accident希望提取出类似「Goleta, California」格式的子串,请问在R语言中该如何实现?
这其实是个典型的正则表达式匹配场景,在R里有几种简洁的实现方式,我给你详细说明:
方法一:用stringr包(推荐,语法更直观)
stringr是tidyverse家族里专门处理字符串的包,用它来提取匹配内容非常方便:
# 先安装包(如果还没装的话) # install.packages("stringr") library(stringr) # 定义你的输入文本 input_text <- "Woman Killed in Uber Accident in Goleta, California Sacramento, California Man Dies in Lyft Accident" # 提取符合格式的子串 location_matches <- str_extract_all(input_text, "[A-Z][a-z]+,\\s[A-Z][a-z]+")[[1]] # 查看结果 print(location_matches)
运行后你会得到结果:
[1] "Goleta, California" "Sacramento, California"
正则表达式说明
[A-Z][a-z]+:匹配首字母大写、后续为小写字母的单词(对应城市/州的单个单词名称),\\s:匹配逗号加空格,对应格式里的分隔部分- 整体组合起来就精准匹配「城市名, 州名」的格式
方法二:用基础R函数(无需额外装包)
如果你不想安装新包,用基础R的gregexpr和regmatches也能实现:
input_text <- "Woman Killed in Uber Accident in Goleta, California Sacramento, California Man Dies in Lyft Accident" # 提取匹配内容 location_matches <- regmatches(input_text, gregexpr("[A-Z][a-z]+,\\s[A-Z][a-z]+", input_text))[[1]] print(location_matches)
扩展:匹配多单词的城市名
如果你的文本里有像「San Francisco, California」这种多单词的城市名,可以把正则表达式调整为:
# 适配多单词城市名的正则 location_matches <- str_extract_all(input_text, "[A-Z][a-z]+(\\s[A-Z][a-z]+)*,\\s[A-Z][a-z]+")[[1]]
这个表达式里的(\\s[A-Z][a-z]+)*表示可以匹配0个或多个「空格+大写开头单词」的部分,完美兼容多单词城市名。
内容的提问来源于stack exchange,提问作者Jane
相关产品推荐
相关产品推荐

