You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于tidyverse实现DataFrame字符串匹配与分类列新增

嘿,这个需求用tidyverse的工具完全可以轻松实现,我给你整理了具体的代码和思路:

解决方案:用tidyverse + stringr实现分类列

我们可以结合dplyr::mutate()、case_when()和stringr的字符串匹配函数,按优先级处理不同的分类规则,同时实现大小写兼容的部分匹配。

完整代码示例

library(tidyverse)

# 初始化你的数据框
x <- data.frame(
  name = c("Alice", "Alice", "Alice", "Bob", "Bob", "Charlie"),
  prod = c("Hard Hat", "Goggles", "Bus Fare", "Goggles", "Training", "Hard Hat, Laptop")
)

# 新增category分类列
x_with_category <- x %>%
  mutate(
    category = case_when(
      # 匹配Hard Hat或Goggles,大小写不敏感,归为PPE
      str_detect(prod, regex("hard hat|goggles", ignore_case = TRUE)) ~ "PPE",
      # 匹配Laptop,大小写不敏感,归为IT
      str_detect(prod, regex("laptop", ignore_case = TRUE)) ~ "IT",
      # 匹配含Bus的内容,大小写不敏感,归为Transport
      str_detect(prod, regex("bus", ignore_case = TRUE)) ~ "Transport",
      # 匹配Training相关内容,大小写不敏感,归为TRAINING
      str_detect(prod, regex("training", ignore_case = TRUE)) ~ "TRAINING",
      # 兜底规则:所有不匹配的情况设为NA(也可以改成"Other"这类自定义值)
      TRUE ~ NA_character_
    )
  )

# 查看最终结果
print(x_with_category)

关键细节解释

  • 规则优先级:case_when()是按顺序匹配条件的,比如Charlie的prod同时包含Hard Hat和Laptop,会优先匹配到前面的PPE规则;如果需要IT优先级更高,调整规则顺序即可。
  • 大小写兼容:通过regex(ignore_case = TRUE)实现,不管目标文本是大写、小写还是混合格式,都能精准匹配。
  • 部分匹配:str_detect()默认支持部分匹配,只要prod字段包含目标关键词就会触发规则,比如"Bus Fare"里的"Bus"就能匹配到Transport分类。

运行结果

name             prod  category
1   Alice         Hard Hat       PPE
2   Alice          Goggles       PPE
3   Alice         Bus Fare Transport
4     Bob          Goggles       PPE
5     Bob         Training   TRAINING
6 Charlie Hard Hat, Laptop       PPE

内容的提问来源于stack exchange,提问作者Simon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 06:45:11