You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中从指定字符串提取带逗号的目标子串

在R语言中提取「城市, 州」格式的子串

我拥有如下字符串列表:

Woman Killed in Uber Accident in Goleta, California Sacramento, California Man Dies in Lyft Accident

希望提取出类似「Goleta, California」格式的子串,请问在R语言中该如何实现?

这其实是个典型的正则表达式匹配场景,在R里有几种简洁的实现方式,我给你详细说明:

方法一:用stringr包(推荐,语法更直观)

stringr是tidyverse家族里专门处理字符串的包,用它来提取匹配内容非常方便:

# 先安装包(如果还没装的话)
# install.packages("stringr")
library(stringr)

# 定义你的输入文本
input_text <- "Woman Killed in Uber Accident in Goleta, California Sacramento, California Man Dies in Lyft Accident"

# 提取符合格式的子串
location_matches <- str_extract_all(input_text, "[A-Z][a-z]+,\\s[A-Z][a-z]+")[[1]]

# 查看结果
print(location_matches)

运行后你会得到结果:

[1] "Goleta, California"    "Sacramento, California"

正则表达式说明

  • [A-Z][a-z]+:匹配首字母大写、后续为小写字母的单词(对应城市/州的单个单词名称)
  • ,\\s:匹配逗号加空格,对应格式里的分隔部分
  • 整体组合起来就精准匹配「城市名, 州名」的格式

方法二:用基础R函数(无需额外装包)

如果你不想安装新包,用基础R的gregexpr和regmatches也能实现:

input_text <- "Woman Killed in Uber Accident in Goleta, California Sacramento, California Man Dies in Lyft Accident"

# 提取匹配内容
location_matches <- regmatches(input_text, gregexpr("[A-Z][a-z]+,\\s[A-Z][a-z]+", input_text))[[1]]

print(location_matches)

扩展:匹配多单词的城市名

如果你的文本里有像「San Francisco, California」这种多单词的城市名,可以把正则表达式调整为:

# 适配多单词城市名的正则
location_matches <- str_extract_all(input_text, "[A-Z][a-z]+(\\s[A-Z][a-z]+)*,\\s[A-Z][a-z]+")[[1]]

这个表达式里的(\\s[A-Z][a-z]+)*表示可以匹配0个或多个「空格+大写开头单词」的部分,完美兼容多单词城市名。

内容的提问来源于stack exchange,提问作者Jane

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:07:12