R语言中提取字符串内LA实例的正则表达式问题
问题描述
现有如下数据:
text_string <- structure(list(text_string = c("A Nanny-Back Up Care and Staffing Company-San Diego, OC, LA, San Francisco, Portland, Las Vegas, Phoenix, Seattle, Denver and NY. @jefffoes", "Creative Producer of @crwnmag LA-NY-TX dereksith@googke.com Founded @marcusharper", "daily elements for life and style texas transplant in california LA lauren@gmail.com read my blog + shop my instagram", "LIVE, LAUGH, LOVE")), class = "data.frame", row.names = c(NA, -4L))
需求是捕获每个字符串中的"LA"实例,生成新字段存储该内容,预期前三条能匹配到"LA",最后一条无匹配。
尝试了以下代码,但新字段返回的是原字段副本,未达到预期:
text_string_new <- text_string %>% mutate(new_field = str_replace(string = text_string, pattern = "(LA)(\\b|,)", replacement = "\\1"))
解决方案
你使用的str_replace是替换类函数,它仅会将匹配到的内容替换为指定值(此处是把LA或LA,替换为LA),无法实现提取匹配内容的需求。要完成捕获"LA"的目标,需使用字符串提取类函数:
- 提取单个"LA"实例(每条取第一个匹配结果)
使用str_extract搭配正则模式,可直接提取字符串中符合要求的"LA",无匹配时返回NA:
library(dplyr) library(stringr) text_string_new <- text_string %>% mutate(new_field = str_extract(text_string, "(LA)(?=\\b|,)"))
这里采用正向预查(?=\\b|,),确保匹配的"LA"后紧跟单词边界或逗号,避免误匹配LIVE、LAUGH中的子串"LA"。
- 提取所有"LA"实例(处理单条字符串含多个匹配的情况)
若需要提取每条字符串中所有符合要求的"LA",可使用str_extract_all,结果为列表格式,也可通过str_c将多个结果合并为单个字符串:
text_string_new <- text_string %>% mutate(new_field = str_extract_all(text_string, "(LA)(?=\\b|,)") %>% map_chr(str_c, collapse = ", "))
若希望无匹配时返回NA而非空字符串,可结合case_when处理:
text_string_new <- text_string %>% mutate(new_field = case_when( str_detect(text_string, "(LA)(?=\\b|,)") ~ str_extract_all(text_string, "(LA)(?=\\b|,)") %>% map_chr(str_c, collapse = ", "), TRUE ~ NA_character_ ))
内容的提问来源于stack exchange,提问作者wizkids121
相关产品推荐
相关产品推荐

