You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何合并近似等价字符值?优化refinr包n_gram_merge效果问询

餐厅名称去重:合并相似变体的优化方案

问题背景

手头有507条餐厅数据,但存在重复条目——因缩写、信用卡处理器添加的特殊前缀/后缀等原因,实际唯一餐厅数量在300-480家之间。人工可轻松判断相似字符串属于同一家餐厅,但现有自动化方法效果不佳。

可复现示例

测试向量(10个元素):

c("Artspace & Coffeehouse", "Artspace & Coffeehou", "Artspace & Coffeehous", "Artspace & Coffeeho", "Artspace & Coffee Q48", "Artspace & Crumpet", "Stone", "Pepper", "Bakersfield", "Bakery Social")

期望结果

返回10元素向量,仅保留6个唯一值:

  • 前5个Artspace咖啡相关变体合并为同一值
  • "Artspace & Crumpet"及其余元素保持独立

尝试过的方法

使用refinr包的n_gram_merge函数:

library(refinr)
restaurants <- c("Artspace & Coffeehouse", "Artspace & Coffeehou", "Artspace & Coffeehous", "Artspace & Coffeeho", "Artspace & Coffee Q48", "Artspace & Crumpet", "Stone", "Pepper", "Bakersfield", "Bakery Social")  
n_gram_merge(vect = restaurants,  edit_threshold = 2)  

调整edit_threshold、numgram参数后,仍无法合并"Artspace & Coffeeho"与"Artspace & Coffee Q48"。此外尝试stringdistmatrix(method="jw")结合cutree、hclust、as.dist分组,同样未达预期。

优化方案

1. 规则式预处理(快速解决特定场景)

先通过自定义规则提取核心名称,去除干扰后缀(如Q48),直接统一标准名称:

library(stringr)

preprocess <- function(x) {
  # 匹配所有Artspace咖啡相关变体,统一替换为标准名
  if (str_detect(x, "^Artspace & Coffee")) {
    return("Artspace & Coffeehouse")
  } else {
    return(x)
  }
}

processed_restaurants <- sapply(restaurants, preprocess)

2. 预处理+参数调优的n_gram_merge(通用场景)

先清理干扰后缀,再调整n_gram_merge的匹配阈值,提升相似性识别能力:

# 去除末尾类似Q48的代码后缀
cleaned_restaurants <- str_replace(restaurants, "\\sQ\\d+$", "")

# 调高编辑阈值,扩大匹配范围
refined <- n_gram_merge(vect = cleaned_restaurants, edit_threshold = 3, numgram = 2)

3. 自定义距离的聚类方案(灵活适配复杂场景)

通过自定义距离函数优先匹配核心前缀,再结合聚类实现精准分组:

library(stringdist)
library(cluster)

# 自定义距离规则:核心前缀匹配则距离为0,否则用Jaro-Winkler距离
custom_dist <- function(a, b) {
  a_core <- str_sub(a, 1, 15) # 提取前15位覆盖"Artspace & Coffee"核心
  b_core <- str_sub(b, 1, 15)
  if (a_core == b_core) {
    return(0)
  } else {
    return(stringdist(a, b, method = "jw"))
  }
}

# 生成距离矩阵并聚类
dist_mat <- stringdistmatrix(restaurants, restaurants, method = custom_dist)
hc <- hclust(as.dist(dist_mat), method = "single")
groups <- cutree(hc, h = 0.1)

# 每组取最长字符串作为标准名称
final_result <- sapply(groups, function(g) {
  group_members <- restaurants[groups == g]
  group_members[which.max(nchar(group_members))]
})

以上三种方案均可实现需求,其中规则式预处理适合快速解决特定场景,聚类方案则能适配更复杂的名称变体情况。


内容的提问来源于stack exchange,提问作者Farrel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 15:30:53