You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中循环/迭代调用str_extract提取国家相关表述?

提取文本中每个国家的首次提及内容

我有一个大型文本数据集,想要识别每个文档中提及的国家。有时文本中会出现“afghanistan”,有时会出现“afghan”,但二者指代同一国家,因此我只想提取二者中的首次提及内容。

现有定义的模式与文本

模式向量

pattern <- c("afghanistan|afghan", "algeria|algerian", "albania|albanian", "angola|angolan", "argentina|argentine")

待处理文本

text <- c("the first stop on the trip is afghanistan, where he will meet the afghan president", 
          "then he will leave afghanistan and head to argentina", 
          "meetings with the afghan president in afghanistan should last 1 hour, and meetings with the argentine president in argentina should last 2 hours")

期望结果

每个文本对应的提取结果如下:

c("afghanistan")
c("afghanistan", "argentina")
c("afghan", "argentine")

之前的尝试问题

  • 用str_extract_all() + unique()处理时,会将同一国家的不同表述(如afghanistan和afghan)视为两个独立结果,导致重复统计。
  • 尝试map()、mapply()的多种用法,结果常为包含character(0)的无效列表。
  • 编写的for循环返回全NA向量:
country <- as.character(1:length(pattern)) #占位向量

for(i in 1:length(pattern)){
    country[i] = str_extract(text, pattern[i])
}

解决方案

使用purrr的迭代函数结合stringr的提取功能,对每条文本逐一匹配所有模式,过滤掉无匹配的NA值,即可得到每个国家的首次提及内容:

代码实现

library(stringr)
library(purrr)

# 定义单文本处理函数:匹配所有模式,提取首次出现内容并过滤NA
extract_first_country <- function(txt) {
  map_chr(pattern, ~str_extract(txt, .x)) %>% 
    na.omit() %>% 
    as.character()
}

# 对所有文本应用函数
result <- map(text, extract_first_country)

# 查看结果
result

输出结果

[[1]]
[1] "afghanistan"

[[2]]
[1] "afghanistan" "argentina"  

[[3]]
[1] "afghan"    "argentine"

如果需要转换为数据框格式,可使用:

library(dplyr)
library(tibble)

tibble(text = text, countries = result)

内容的提问来源于stack exchange,提问作者user14517212

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 23:42:29