You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于字符串距离的不等长数据集匹配R语言实现咨询

R语言模糊匹配实现方案

1. 依赖包准备

需要用到tidyverse做数据处理,stringdist做字符串距离计算,加载代码:

# 首次运行需先安装依赖
install.packages(c("tidyverse", "stringdist"))
# 加载包
library(tidyverse)
library(stringdist)

2. 构造示例数据集

# 数据集A
dataset_a <- tibble(
  name = c("Sally", "Peter", "Joe", "Samantha", "Kyle", "Kieran", "Molly"),
  age = c(22, 35, 57, 33, 30, 41, 28)
)
# 数据集B
dataset_b <- tibble(
  name = c("Samanta", "Peter", "Joey", "Samantha"),
  company = c("A", "B", "C", "A")
)

3. 匹配逻辑实现

match_result <- dataset_a %>%
  rename(name_a = name) %>% # 重命名避免列名冲突
  crossing(dataset_b %>% rename(name_b = name)) %>% # 生成所有姓名配对组合
  mutate(distance = stringdist(name_a, name_b, method = "jw")) %>% # 计算Jaro-Winkler距离
  filter(distance < 0.1) %>% # 过滤非相关匹配,阈值可按需调整
  arrange(distance) %>% # 按距离升序排序
  select(name_a, name_b, age, company, distance) # 调整列顺序匹配预期输出

输出结果

打印match_result即可得到符合要求的结果:

# A tibble: 4 × 5
  name_a   name_b    age company distance
  <chr>    <chr>   <dbl> <chr>      <dbl>
1 Peter    Peter      35 B         0     
2 Samantha Samantha   33 A         0     
3 Samantha Samanta    33 A         0.0417
4 Joe      Joey       57 C         0.0833

如果需要调整匹配的严格程度,修改filter里的阈值即可,阈值越大允许的匹配误差越大。

内容的提问来源于stack exchange,提问作者teogj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 14:57:01