如何从R语言DataFrame字符串列中提取匹配州名的内容?
解决R语言提取文本中美国州名的问题
错误原因
你遇到的报错是因为str_match要求输入的string和pattern长度匹配,而state.name是包含50个州名的向量,quote只有3个字符串,R无法完成长度循环匹配,因此抛出错误。
正确提取方法
将所有州名合并成一个正则表达式的备选模式,用str_extract从每个文本中匹配并提取州名:
library(tidyverse) # 示例数据集 df <- tibble( num = c(11,12,13), quote = c( "In Ohio, there are plenty of hobos", "Georgia, where the peaches are peachy", "Oregon, no, we did not die of dysentery" ) ) # 生成匹配所有州名的正则模式 state_pattern <- str_c(state.name, collapse = "|") # 新增列提取州名 df <- df %>% mutate(state = str_extract(quote, state_pattern)) # 输出结果 df
运行后会得到包含提取出的州名的新列:
# A tibble: 3 × 3 num quote state <dbl> <chr> <chr> 1 11 In Ohio, there are plenty of hobos Ohio 2 12 Georgia, where the peaches are peachy Georgia 3 13 Oregon, no, we did not die of dysentery Oregon
补充说明
- 如果某条文本包含多个州名,
str_extract只会返回第一个匹配到的州名;若要提取所有匹配的州名,可改用str_extract_all,结果会是列表格式。 - 若需要更精确的匹配(比如避免州名作为其他单词的一部分被提取),可以给正则模式加上单词边界:
state_pattern <- str_c("\\b", state.name, "\\b", collapse = "|")。
内容的提问来源于stack exchange,提问作者Robert Frey
相关产品推荐
相关产品推荐

