You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于文本列添加匹配关键词列表列的R语言技术问询

解决R语言数据框提取匹配关键词并生成列表列的问题

嘿,这个需求其实很容易实现,咱们用R的工具就能搞定!我给你两种方案,一种是用tidyverse生态的工具,另一种是纯基础R的写法,你看哪种顺手:

第一步:先准备好你的初始数据

首先咱们先把你给出的数据定义好,确保格式正确:

df <- data.frame(
  text = c("This string is not that long", 
           "This string is a bit longer but still not that long", 
           "This one just helps with the example"),
  stringsAsFactors = FALSE
)
keywords <- c("not that long", "This string", "example", "helps")

方案一:用tidyverse的purrr和stringr包

这个方案代码更简洁易读,适合习惯tidyverse风格的用户:

library(purrr)
library(stringr)

# 新增keywords列,每行存储匹配到的关键词列表
df$keywords <- map(df$text, ~ keywords[str_detect(.x, fixed(keywords))])

这里的关键细节:

  • map()函数会遍历df$text的每一行文本
  • str_detect(.x, fixed(keywords))会检查每个关键词是否精确匹配当前行的文本(用fixed()是为了避免关键词里的空格或特殊字符被当成正则表达式处理)
  • 最后通过keywords[匹配结果]筛选出当前行匹配到的所有关键词,组成一个向量存入对应的行

方案二:纯基础R实现(无需额外加载包)

如果不想加载额外的包,用基础R的lapply和grepl也能完成:

# 基础R版本的实现
df$keywords <- lapply(df$text, function(x) {
  # 遍历每个关键词,检查是否在当前文本中
  match_flag <- sapply(keywords, function(k) grepl(fixed(k), x))
  # 筛选出匹配的关键词
  keywords[match_flag]
})

验证结果

现在咱们查看生成的keywords列,就得到你想要的结果啦:

print(df$keywords)
# 输出结果:
# [[1]]
# [1] "not that long" "This string"  
# 
# [[2]]
# [1] "not that long" "This string"  
# 
# [[3]]
# [1] "example" "helps"  

内容的提问来源于stack exchange,提问作者Saleem Khan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:48:46