R语言大数据框基于Jaro-Winkler相似度的模糊左连接实现问题
解决方案
核心采用stringdist包的amatch函数实现,该方案不会生成全量笛卡尔积,内存占用极低,可轻松适配10万+2.5万规模的数据匹配需求。
依赖安装
若未安装对应依赖先执行安装:
install.packages(c("tidyverse", "stringdist"))
实现代码
library(tidyverse) library(stringdist) # 示例数据 df1 <- data.frame( id = c(1, 2, 3, 4, 5, 6, 7, 8, 9, 10), join_column = c("alice123burgerstorechicago", "alicewonderland", "bubbletea45london", "blueonion", "chandle34song", "crazyjoeohio", "donaldduckshop123", "dartcommunitygermany", "evergreen78hall", "exittheroom15florida")) df2 <- data.frame( id = c(15, 16, 18, 20), join_column = c("aliceburgerstorechicag", "bubbletealndon", "crazyjoeohio178", "exittheroom25florid")) # 执行模糊匹配,返回df2中匹配项的索引 match_index <- amatch( x = df1$join_column, table = df2$join_column, method = "jw", # 指定使用Jaro-Winkler距离 maxDist = 0.2, # 对应Jaro-Winkler相似度≥0.8的要求(相似度=1-距离) p = 0.1 # Jaro-Winkler算法的前缀权重,默认值符合通用场景需求 ) # 拼接生成最终结果 result <- df1 %>% mutate( joined_with_id = df2$id[match_index], joined_with_string = df2$join_column[match_index] )
结果验证
上述代码输出的result与你给出的预期target完全一致,且运行过程中不会出现内存不足的问题。
内容的提问来源于stack exchange,提问作者crazy-wasserratte
相关产品推荐
相关产品推荐

