如何合并近似等价字符值?优化refinr包n_gram_merge效果问询
餐厅名称去重:合并相似变体的优化方案
问题背景
手头有507条餐厅数据,但存在重复条目——因缩写、信用卡处理器添加的特殊前缀/后缀等原因,实际唯一餐厅数量在300-480家之间。人工可轻松判断相似字符串属于同一家餐厅,但现有自动化方法效果不佳。
可复现示例
测试向量(10个元素):
c("Artspace & Coffeehouse", "Artspace & Coffeehou", "Artspace & Coffeehous", "Artspace & Coffeeho", "Artspace & Coffee Q48", "Artspace & Crumpet", "Stone", "Pepper", "Bakersfield", "Bakery Social")
期望结果
返回10元素向量,仅保留6个唯一值:
- 前5个Artspace咖啡相关变体合并为同一值
- "Artspace & Crumpet"及其余元素保持独立
尝试过的方法
使用refinr包的n_gram_merge函数:
library(refinr) restaurants <- c("Artspace & Coffeehouse", "Artspace & Coffeehou", "Artspace & Coffeehous", "Artspace & Coffeeho", "Artspace & Coffee Q48", "Artspace & Crumpet", "Stone", "Pepper", "Bakersfield", "Bakery Social") n_gram_merge(vect = restaurants, edit_threshold = 2)
调整edit_threshold、numgram参数后,仍无法合并"Artspace & Coffeeho"与"Artspace & Coffee Q48"。此外尝试stringdistmatrix(method="jw")结合cutree、hclust、as.dist分组,同样未达预期。
优化方案
1. 规则式预处理(快速解决特定场景)
先通过自定义规则提取核心名称,去除干扰后缀(如Q48),直接统一标准名称:
library(stringr) preprocess <- function(x) { # 匹配所有Artspace咖啡相关变体,统一替换为标准名 if (str_detect(x, "^Artspace & Coffee")) { return("Artspace & Coffeehouse") } else { return(x) } } processed_restaurants <- sapply(restaurants, preprocess)
2. 预处理+参数调优的n_gram_merge(通用场景)
先清理干扰后缀,再调整n_gram_merge的匹配阈值,提升相似性识别能力:
# 去除末尾类似Q48的代码后缀 cleaned_restaurants <- str_replace(restaurants, "\\sQ\\d+$", "") # 调高编辑阈值,扩大匹配范围 refined <- n_gram_merge(vect = cleaned_restaurants, edit_threshold = 3, numgram = 2)
3. 自定义距离的聚类方案(灵活适配复杂场景)
通过自定义距离函数优先匹配核心前缀,再结合聚类实现精准分组:
library(stringdist) library(cluster) # 自定义距离规则:核心前缀匹配则距离为0,否则用Jaro-Winkler距离 custom_dist <- function(a, b) { a_core <- str_sub(a, 1, 15) # 提取前15位覆盖"Artspace & Coffee"核心 b_core <- str_sub(b, 1, 15) if (a_core == b_core) { return(0) } else { return(stringdist(a, b, method = "jw")) } } # 生成距离矩阵并聚类 dist_mat <- stringdistmatrix(restaurants, restaurants, method = custom_dist) hc <- hclust(as.dist(dist_mat), method = "single") groups <- cutree(hc, h = 0.1) # 每组取最长字符串作为标准名称 final_result <- sapply(groups, function(g) { group_members <- restaurants[groups == g] group_members[which.max(nchar(group_members))] })
以上三种方案均可实现需求,其中规则式预处理适合快速解决特定场景,聚类方案则能适配更复杂的名称变体情况。
内容的提问来源于stack exchange,提问作者Farrel
相关产品推荐
相关产品推荐

