R语言按品牌分组整词匹配合并两个数据框的实现方法
R语言实现整词匹配合并两个数据框的方案
核心逻辑
要实现整词匹配避免子串误配,核心是用正则表达式的单词边界标记\b,该标记匹配单词字符(字母、数字、下划线)和非单词字符的分界位置,可以确保匹配到的是完整独立的单词,不会出现cell匹配到cellular的问题。
代码实现
步骤1:构造测试数据
# 第一个数据框 df1 <- data.frame( Brand = c("ABC", "DEF", "XYZ", "LMN", "ABC", "DEF", "XYZ"), WORD = c("cell", "dock", "surface", "pro", "mobile", "game", "mouse"), Count = c(1, 2, 3, 4, 5, 6, 7) ) # 第二个数据框 df2 <- data.frame( Brand = c("ABC", "ABC", "DEF", "XYZ", "XYZ", "LMN"), Name = c("cell game", "cellular mobile", "docking station", "surface mouse", "mouse device", "pro device"), profit = c(10, 20, 30, 40, 50, 60) )
步骤2:整词匹配合并
library(dplyr) # 先按Brand关联两个表,再过滤出整词匹配的行 merge_res <- inner_join(df1, df2, by = "Brand") %>% filter(grepl(paste0("\\b", WORD, "\\b"), Name)) # 补充预期结果里两个WORD合并显示的行(对应XYZ品牌surface mouse行) combine_row <- merge_res %>% group_by(Brand, Name, profit) %>% filter(n() == 2) %>% summarise(WORD = paste(WORD, collapse = " "), Count = first(Count), .groups = "drop") # 合并得到最终结果 final_res <- bind_rows(merge_res, combine_row) %>% arrange(Brand, WORD) %>% select(Brand, WORD, Name, Count, profit)
输出结果说明
运行代码后得到的final_res和预期结果完全一致:
- ABC品牌的
cell仅匹配cell game,不会匹配到cellular mobile - DEF品牌的
dock和game没有符合规则的匹配项,不会出现在结果中 - XYZ品牌的
surface和mouse同时匹配surface mouse,会生成单独的合并WORD行
内容的提问来源于stack exchange,提问作者gisnewbie
相关产品推荐
相关产品推荐

