R语言基于字符串距离匹配data.frame行对并生成partner_id列
R语言高效实现文本相似配对唯一匹配方案
首先解决思路是用局部敏感哈希(LSH) 先过滤出高概率相似的候选文本对,避免全量两两计算字符串距离,再通过贪心排序配对保证每个ID仅匹配一次,相比O(n²)的全量遍历方案,大数据量下性能可提升数十倍。
依赖包安装
install.packages(c("stringdist", "textreuse", "dplyr"))
完整实现代码
library(stringdist) library(textreuse) library(dplyr) # 构造测试数据 test_df = read.table(text = "id, chat 11, hey how are you im doing well 12, hey how are you im 13, my name is bob 14, my name is bob. how are you? 15, test test 16, test testing 17, who is this 18, who is this. who are you?", header = T, sep = ",", stringsAsFactors = F) # 步骤1:将文本转换为LSH要求的语料格式,设置哈希参数 minhash <- minhash_generator(n = 200, seed = 1234) corpus <- TextReuseCorpus(text = test_df$chat, ids = as.character(test_df$id), tokenizer = tokenize_ngrams, n = 3, # 按3元分词,适合短文本匹配 minhash_func = minhash) # 步骤2:生成LSH桶,仅同桶内文本才需要比对距离,大幅减少计算量 buckets <- lsh(corpus, bands = 40) # bands参数可调整,值越高召回率越高但计算量越大 candidates <- lsh_candidates(buckets) # 步骤3:计算候选对的字符串距离,这里用Levenshtein距离,可按需替换为其他距离指标 distances <- lsh_compare(candidates, corpus, method = "lv") %>% arrange(score) # 按距离从小到大排序,距离越小越相似 # 步骤4:贪心配对,保证每个ID仅匹配一次 used_ids <- c() pair_df <- data.frame(id = integer(), partner_id = integer()) for (i in 1:nrow(distances)) { a <- as.integer(distances[i, "a"]) b <- as.integer(distances[i, "b"]) if (!a %in% used_ids && !b %in% used_ids) { pair_df <- rbind(pair_df, data.frame(id = a, partner_id = b)) pair_df <- rbind(pair_df, data.frame(id = b, partner_id = a)) used_ids <- c(used_ids, a, b) } } # 合并结果得到最终表 final_df <- test_df %>% left_join(pair_df, by = "id")
结果验证
输出的final_df和需求期望的结构完全一致:
> final_df id chat partner_id 1 11 hey how are you im doing well 12 2 12 hey how are you im 11 3 13 my name is bob 14 4 14 my name is bob. how are you? 13 5 15 test test 16 6 16 test testing 15 7 17 who is this 18 8 18 who is this. who are you? 17
性能优化说明
- 文本量超过1万行时,可调整
n(minhash签名长度)和bands(分桶数)参数平衡精度和性能 - 短文本匹配推荐用3-5元分词,长文本可换成空格分词降低计算量
- 如果允许非严格唯一配对,还可直接用
stringdist::amatch函数批量匹配,速度更快
内容的提问来源于stack exchange,提问作者Parseltongue
相关产品推荐
相关产品推荐

