You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言使用dplyr高效匹配关键词为dataframe的class列赋值

实现方案

以下是纯tidyverse(dplyr+stringr)的高效实现,底层用矢量化运算,避免R级循环,大体积数据集下运行速度远高于循环方案:

1. 依赖加载

library(dplyr)
library(stringr)

2. 构造关键词映射表

先把分散的关键词组整合成结构化的映射表,方便后续维护和扩展:

# 你的原有关键词组
indus_1 <- c("iron", "steel")
goods_store_1 <- c("goods", "store", "stores")
electr_1 <- c("electronic", "chips", "semiconductor")
unlabelled_1 <- c("leather")

# 构造映射表
keyword_map <- tibble(
  class_label = str_remove(c("indus_1", "goods_store_1", "electr_1", "unlabelled_1"), "_1$"),
  pattern = map_chr(list(indus_1, goods_store_1, electr_1, unlabelled_1), ~ str_c(.x, collapse = "|"))
)

3. 批量匹配赋值

result <- df %>%
  # 原有空值、unknown转NA,便于后续覆盖,已存在的有效值(如sports)保留
  mutate(
    class = na_if(class, ""),
    class = na_if(class, "unknown"),
    # 合并待匹配的两列,转小写避免大小写敏感问题,不需要可删除tolower
    match_text = tolower(str_c(desc1, " ", store_names))
  ) %>%
  mutate(
    class = case_when(
      !is.na(class) ~ class,
      str_detect(match_text, keyword_map$pattern[1]) ~ keyword_map$class_label[1],
      str_detect(match_text, keyword_map$pattern[2]) ~ keyword_map$class_label[2],
      str_detect(match_text, keyword_map$pattern[3]) ~ keyword_map$class_label[3],
      str_detect(match_text, keyword_map$pattern[4]) ~ keyword_map$class_label[4],
      .default = ""
    )
  ) %>%
  select(-match_text)

匹配遗漏问题解决

你之前用\\b出现遗漏是因为\\b是单词边界匹配,仅在单词字符(字母、数字、下划线)和非单词字符的分界处生效,比如steels&iron里的steel后面跟的是s(单词字符),加\\b的话\\bsteel\\b就匹配不到。如果需要做完整单词匹配又要避免这类问题,可以先把match_text里的所有非字母字符替换为空格,再加\\b包裹关键词即可:

# 预处理match_text替换非字母为空格
match_text = str_replace_all(tolower(str_c(desc1, " ", store_names)), "[^a-z]+", " ")
# 构造pattern时加单词边界
pattern = map_chr(list(indus_1, goods_store_1, electr_1, unlabelled_1), ~ str_c("\\b", .x, "\\b", collapse = "|"))

运行后输出结果和你给出的预期完全一致。

内容的提问来源于stack exchange,提问作者Vendetta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 12:36:02