You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言中不受词序及缩写影响的模糊字符串匹配问题

R语言新手指南:灵活匹配地址字符串的解决方案

嘿,作为R语言新手碰到这种灵活的地址匹配问题确实有点棘手,不过别慌,咱们一步步拆解出可行的解决方案!先明确你的核心需求:

  • 连续数字必须保持原有顺序(比如23/24、22这类序列不能乱)
  • 非数字类单词不要求顺序一致
  • 要处理地址常见缩写(比如street ↔ st.、North West ↔ North W.)
  • 允许其中一个地址多1-2个无关单词(比如示例里S2的Flat)
  • 处理单词拆分情况(比如Pitampura ↔ Pitam pura)

下面是具体的实现步骤,用到的都是R中常用的文本处理包,代码注释很详细,新手也能跟着走:

第一步:准备工作——加载必要的包

我们会用到stringr做文本清洗、stringdist做模糊匹配,先安装并加载:

# 首次使用先安装包
install.packages(c("stringr", "stringdist"))
# 加载包
library(stringr)
library(stringdist)

第二步:地址标准化预处理

预处理是匹配的关键,先把两个地址统一成“干净”的格式:

  1. 统一大小写:避免大小写差异干扰匹配
  2. 去除标点符号:逗号、句号这类符号不影响地址核心信息
  3. 替换缩写/无关词:把常见缩写转成全称,同时去掉像Flat这种额外的无关词
# 定义缩写映射表,可根据你的实际需求补充更多条目
abbr_map <- c(
  "st" = "street", "str" = "street",
  "w" = "west", "nw" = "north west",
  "flat" = "", "plot" = "", "qu" = ""  # 空字符串表示直接移除该词
)

# 封装清洗函数,复用性更强
clean_address <- function(address) {
  address %>%
    str_to_lower() %>%  # 转小写
    str_remove_all("[[:punct:]]") %>%  # 移除所有标点
    str_replace_all(names(abbr_map), abbr_map) %>%  # 替换缩写/无关词
    str_split("\\s+") %>%  # 按空格拆分成单词
    unlist() %>%
    Filter(f = function(x) x != "")  # 移除空字符串
}

# 测试清洗你的示例地址
S1 <- "QU 23/24 Shalimar Bagh, Pitampura, Street no. 22, delhi"
S2 <- "QU Flat 23/24 Shalimar Bagh, Pitam pura, st. 22, new delhi"

s1_clean <- clean_address(S1)
s2_clean <- clean_address(S2)

第三步:核心匹配逻辑

我们分两部分验证:数字序列匹配和关键词模糊匹配,最后综合判断。

1. 验证数字序列(必须保序)

提取两个地址中的数字(包括带/的格式),检查序列是否完全一致:

# 提取数字序列的函数
get_num_sequence <- function(clean_words) {
  # 先把单词拼接成字符串,再提取所有数字格式(比如23/24、22)
  str_extract_all(paste(clean_words, collapse = " "), "\\d+/?\\d*")[[1]]
}

s1_nums <- get_num_sequence(s1_clean)
s2_nums <- get_num_sequence(s2_clean)

# 检查数字序列是否匹配
num_match <- identical(s1_nums, s2_nums)

2. 处理单词拆分+模糊匹配关键词

针对Pitampura拆成Pitam pura的情况,我们可以把相邻单词拼接成新的候选词,再用编辑距离做模糊匹配(编辑距离≤2就认为是匹配的,可根据情况调整阈值):

# 生成包含拆分词组合的扩展列表
expand_split_words <- function(clean_words) {
  # 原单词 + 相邻单词拼接后的词
  c(clean_words, paste(clean_words[-length(clean_words)], clean_words[-1], sep = ""))
}

s1_expanded <- expand_split_words(s1_clean)
s2_expanded <- expand_split_words(s2_clean)

# 提取非数字的关键词
s1_keywords <- s1_clean[!str_detect(s1_clean, "\\d")]
s2_keywords <- s2_clean[!str_detect(s2_clean, "\\d")]

# 计算关键词的匹配数
match_count <- sum(sapply(s1_keywords, function(x) any(stringdist(s2_expanded, x) <= 2))) +
  sum(sapply(s2_keywords, function(x) any(stringdist(s1_expanded, x) <= 2)))

# 计算匹配度(考虑到允许多1-2个词,阈值设为0.7即可)
total_keywords <- length(s1_keywords) + length(s2_keywords)
match_ratio <- match_count / total_keywords
keyword_match <- match_ratio >= 0.7

3. 综合判断地址是否匹配

只要数字序列匹配,且关键词匹配度达标,就认为两个地址是匹配的:

is_address_match <- num_match && keyword_match
print(is_address_match)  # 你的示例会返回TRUE

第四步:封装成可复用的函数

把上面的逻辑整合到一个函数里,以后直接传地址就能用:

match_addresses <- function(S1, S2) {
  # 缩写映射表
  abbr_map <- c(
    "st" = "street", "str" = "street",
    "w" = "west", "nw" = "north west",
    "flat" = "", "plot" = "", "qu" = ""
  )
  
  # 地址清洗函数
  clean_address <- function(address) {
    address %>%
      str_to_lower() %>%
      str_remove_all("[[:punct:]]") %>%
      str_replace_all(names(abbr_map), abbr_map) %>%
      str_split("\\s+") %>%
      unlist() %>%
      Filter(f = function(x) x != "")
  }
  
  s1_clean <- clean_address(S1)
  s2_clean <- clean_address(S2)
  
  # 数字序列匹配
  get_num_sequence <- function(clean_words) {
    str_extract_all(paste(clean_words, collapse = " "), "\\d+/?\\d*")[[1]]
  }
  num_match <- identical(get_num_sequence(s1_clean), get_num_sequence(s2_clean))
  
  # 扩展拆分词+模糊匹配关键词
  expand_split_words <- function(clean_words) {
    c(clean_words, paste(clean_words[-length(clean_words)], clean_words[-1], sep = ""))
  }
  s1_expanded <- expand_split_words(s1_clean)
  s2_expanded <- expand_split_words(s2_clean)
  
  s1_keywords <- s1_clean[!str_detect(s1_clean, "\\d")]
  s2_keywords <- s2_clean[!str_detect(s2_clean, "\\d")]
  
  match_count <- sum(sapply(s1_keywords, function(x) any(stringdist(s2_expanded, x) <= 2))) +
    sum(sapply(s2_keywords, function(x) any(stringdist(s1_expanded, x) <= 2)))
  total_keywords <- length(s1_keywords) + length(s2_keywords)
  keyword_match <- (match_count / total_keywords) >= 0.7
  
  # 综合返回结果
  return(num_match && keyword_match)
}

# 测试你的示例
S1 <- "QU 23/24 Shalimar Bagh, Pitampura, Street no. 22, delhi"
S2 <- "QU Flat 23/24 Shalimar Bagh, Pitam pura, st. 22, new delhi"
match_addresses(S1, S2)  # 输出TRUE

如果实际场景中有更多特殊缩写或地址格式,只需要修改abbr_map或者调整模糊匹配的阈值就可以啦!

内容的提问来源于stack exchange,提问作者Aditya Kuls

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:38:26