You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从R语言DataFrame字符串列中提取匹配州名的内容?

解决R语言提取文本中美国州名的问题

错误原因

你遇到的报错是因为str_match要求输入的string和pattern长度匹配,而state.name是包含50个州名的向量,quote只有3个字符串,R无法完成长度循环匹配,因此抛出错误。

正确提取方法

将所有州名合并成一个正则表达式的备选模式,用str_extract从每个文本中匹配并提取州名:

library(tidyverse)

# 示例数据集
df <- tibble(
  num = c(11,12,13),
  quote = c(
    "In Ohio, there are plenty of hobos",
    "Georgia, where the peaches are peachy",
    "Oregon, no, we did not die of dysentery"
  )
)

# 生成匹配所有州名的正则模式
state_pattern <- str_c(state.name, collapse = "|")

# 新增列提取州名
df <- df %>% 
  mutate(state = str_extract(quote, state_pattern))

# 输出结果
df

运行后会得到包含提取出的州名的新列:

# A tibble: 3 × 3
    num quote                                      state 
  <dbl> <chr>                                     <chr> 
1    11 In Ohio, there are plenty of hobos        Ohio  
2    12 Georgia, where the peaches are peachy     Georgia
3    13 Oregon, no, we did not die of dysentery   Oregon

补充说明

  • 如果某条文本包含多个州名,str_extract只会返回第一个匹配到的州名;若要提取所有匹配的州名,可改用str_extract_all,结果会是列表格式。
  • 若需要更精确的匹配(比如避免州名作为其他单词的一部分被提取),可以给正则模式加上单词边界:state_pattern <- str_c("\\b", state.name, "\\b", collapse = "|")。

内容的提问来源于stack exchange,提问作者Robert Frey

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 04:03:14