You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言DataFrame中匹配单词列表并生成匹配词新列

嘿,这个需求很常见,我来给你几个实用的解法!要从你的单词列表里匹配DataFrame的text列,并把匹配到的动物名存入新列,我们可以结合R的dplyr和stringr包来实现,同时还能处理大小写不一致、完整单词匹配这些细节问题。

先明确前提

你的单词列表和text列里的动物名存在大小写差异(比如列表里的Lion、tiger,text里的lion、Tiger),所以我们优先做大小写不敏感的完整单词匹配,避免误匹配子字符串(比如不会把dogs里的dog误提出来,除非你需要处理复数)。

方法1:提取小写匹配结果(简洁直接)

先加载需要的包,然后构建正则表达式,提取所有匹配的单词:

library(dplyr)
library(stringr)

# 你的原始数据
list_of_words <- c("tiger","elephant","rabbit", "hen", "dog", "Lion", "camel", "horse")
df <- tibble::tibble(page=c(12,6,9,18,2,15,81,65), text=c("I have two pets: a dog and a hen", "lion and Tiger are dangerous animals", "I have tried to ride a horse", "Why elephants are so big in size", "dogs are very loyal pets", "I saw a tiger in the zoo", "the lion was eating a buffalo", "parrot and crow are very clever birds"))

# 构建匹配正则:单词边界 + 小写单词列表 + 单词边界,用|分隔
match_pattern <- str_c("\\b", str_to_lower(list_of_words), "\\b", collapse = "|")

# 提取匹配结果并存入新列
df_with_matches <- df %>%
  mutate(
    # 提取所有匹配的单词,返回列表
    matched_animals_list = str_extract_all(str_to_lower(text), match_pattern),
    # 把列表转为逗号分隔的字符串,方便查看
    matched_animals = str_c(matched_animals_list, collapse = ", ")
  )

这样处理后,matched_animals列会显示每行text里匹配到的动物名,比如第二行的结果是lion, tiger。

方法2:保留原文本的大小写

如果你想保留text里动物名的原始大小写(比如提取Tiger而不是tiger),可以调整正则的匹配模式:

# 构建带大小写不敏感标记的正则
match_pattern <- str_c("\\b(", str_c(list_of_words, collapse = "|"), ")\\b")

df_with_matches <- df %>%
  mutate(
    matched_animals_list = str_extract_all(text, regex(match_pattern, ignore_case = TRUE)),
    matched_animals = str_c(matched_animals_list, collapse = ", ")
  )

这时候第二行的结果会是lion, Tiger,和原文本的大小写一致。

可选:处理复数形式

如果你的text里有动物的复数(比如dogs),而单词列表里是单数(dog),想要匹配的话,可以修改正则允许末尾的s:

match_pattern <- str_c("\\b", str_to_lower(list_of_words), "s?\\b", collapse = "|")

df_with_matches <- df %>%
  mutate(
    matched_animals_list = str_extract_all(str_to_lower(text), match_pattern),
    # 把提取到的复数转成单数
    matched_animals_clean = str_remove(unlist(matched_animals_list), "s$"),
    matched_animals = str_c(unique(matched_animals_clean), collapse = ", ")
  )

这样第五行的dogs会被匹配并转为dog,存入新列。

结果示例

处理后的DataFrame会新增matched_animals列,比如:

pagetextmatched_animals
12I have two pets: a dog and a hendog, hen
6lion and Tiger are dangerous animalslion, tiger
2dogs are very loyal petsdog

内容的提问来源于stack exchange,提问作者Fraxxx

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:41:37