求移除字符向量中单词右侧至少2个空格后数字与逗号的正则表达式
问题:提取Azure Speech服务区域名称的正则表达式
我抓取了支持微软Speech服务的区域表格,得到了如下R语言字符向量:
region <- c("southafricanorth 6", "eastasia 5", "southeastasia 1,2,3,4,5", "australiaeast 1,2,3,4", "centralindia 1,2,3,4,5", "japaneast 2,5", "japanwest", "koreacentral 2", "canadacentral 1", "northeurope 1,2,4,5", "westeurope 1,2,3,4,5", "francecentral", "germanywestcentral", "norwayeast", "switzerlandnorth 6", "switzerlandwest", "uksouth 1,2,3,4", "uaenorth 6", "brazilsouth 6", "centralus", "eastus 1,2,3,4,5", "eastus2 1,2,4,5", "northcentralus 4,6", "southcentralus 1,2,3,4,5,6", "westcentralus 5", "westus 2,5", "westus2 1,2,4,5", "westus3" )
我需要一个正则表达式,移除区域名称右侧至少2个空格后的所有数字和逗号,例如将westus2 1,2,4,5处理为westus2。我尝试了gsub("\s{2,}\d+.*", "", region)但没有效果,请问正确的正则表达式是什么?
解决方法
你之前的正则失效主要有两个原因:
- R语言字符串中,反斜杠需要双重转义,
\s要写成\\s,\d要写成\\d - 匹配数字和逗号的组合,用
[\\d,]+比\\d+.*更精准,能直接覆盖数字、逗号的任意排列
正确的代码如下:
clean_region <- gsub("\\s{2,}[\\d,]+", "", region)
如果需要更严谨的匹配(比如处理末尾可能的逗号),可以用:
clean_region <- gsub("\\s{2,}(?:\\d+,)*\\d*", "", region)
运行后,所有带后缀数字逗号的区域名称都会被清理成纯区域名,符合需求。
内容的提问来源于stack exchange,提问作者Howard Baik
相关产品推荐
相关产品推荐

