如何用R语言通过字符串匹配完成数据框的混乱数据分类?
R语言实现字符串匹配分类混乱公司数据
实现思路
核心逻辑:先提取分类规则(分类名称-匹配关键词),再对每个公司名称做不区分大小写的字符串匹配,找到包含的关键词后,将对应分类填充到company_type列。若存在多个匹配关键词,取第一个对应的分类(可通过调整规则顺序改变优先级)。
代码实现
1. 构造初始数据框
先还原你的初始数据结构:
# 创建初始数据框 df <- data.frame( company_name = c("John landscaping", "constructing pros", "Brother Lawn care", "Top cleaning", "Emmas paints"), categories = c("Landscaping", "Landscaping", "Cleaning", "Painting", "Construction"), search = c("lawn", "landscape", "clean", "paint", "construct"), company_type = rep(NA, 5), stringsAsFactors = FALSE )
2. 提取去重的分类规则
从数据中提取唯一的分类-关键词对应关系,避免重复匹配:
# 提取去重后的分类规则表 category_rules <- unique(df[, c("categories", "search")])
3. 执行匹配并填充分类
使用stringr包处理字符串匹配,结合sapply批量处理每个公司名称:
library(stringr) # 定义匹配函数:输入公司名称,返回对应的分类 get_company_type <- function(name, rules) { # 不区分大小写检查公司名是否包含关键词 matches <- str_detect(tolower(name), tolower(rules$search)) # 返回第一个匹配的分类,无匹配则返回NA if (any(matches)) { return(rules$categories[matches][1]) } else { return(NA) } } # 批量填充company_type列 df$company_type <- sapply(df$company_name, get_company_type, rules = category_rules)
4. 查看结果
运行后打印数据框即可得到目标结果:
print(df)
输出内容:
company_name categories search company_type 1 John landscaping Landscaping lawn Landscaping 2 constructing pros Landscaping landscape Construction 3 Brother Lawn care Cleaning clean Landscaping 4 Top cleaning Painting paint Cleaning 5 Emmas paints Construction construct Painting
替代方案(dplyr管道风格)
如果你习惯使用dplyr的管道语法,可采用以下写法:
library(dplyr) library(stringr) category_rules <- df %>% select(categories, search) %>% distinct() df <- df %>% rowwise() %>% mutate(company_type = category_rules$categories[str_detect(tolower(company_name), tolower(category_rules$search))][1]) %>% ungroup()
内容的提问来源于stack exchange,提问作者Linda
相关产品推荐
相关产品推荐

