You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何实现两个DataFrame间文本字符串的非对称部分匹配?

解决方案(R语言)

先准备测试用的示例数据,方便验证效果:

# 示例df1:用户自由填写的地点数据
df1 <- data.frame(
  Location = c(
    "Alice: London/Liverpool",
    "David: I am based in Cardiff",
    "Bob: No fixed address",
    "Charlie: Manchester, UK",
    "Eve: Edinburgh"
  ),
  stringsAsFactors = FALSE
)

# 示例df2:标准英国城镇列表
df2 <- data.frame(
  Town = c("London", "Liverpool", "Cardiff", "Manchester", "Edinburgh", "Nottingham"),
  stringsAsFactors = FALSE
)

核心实现代码

解决关键是用单词边界正则确保只匹配完整城镇名(避免子串误匹配,比如不让"No"匹配"Nottingham"),再遍历提取第一个匹配项:

library(stringr)
library(purrr)

# 给每个城镇名加上单词边界,生成正则匹配规则
town_patterns <- str_c("\\b", df2$Town, "\\b")

# 定义单条文本的匹配函数
extract_first_match <- function(text) {
  # 找出所有匹配的城镇索引
  match_idx <- which(str_detect(text, town_patterns))
  if (length(match_idx) > 0) {
    df2$Town[match_idx[1]]  # 返回第一个匹配的城镇
  } else {
    "-"  # 无匹配时填充'-'
  }
}

# 给df1新增精准匹配列
df1$Location_precise <- map_chr(df1$Location, extract_first_match)

运行结果

执行后df1的输出如下:

print(df1)
#                     Location Location_precise
# 1    Alice: London/Liverpool           London
# 2 David: I am based in Cardiff          Cardiff
# 3        Bob: No fixed address               -
# 4        Charlie: Manchester, UK      Manchester
# 5               Eve: Edinburgh        Edinburgh

适配特殊场景

如果df2里有带空格、特殊字符的城镇名(比如"Newcastle upon Tyne"),可以用str_escape转义特殊字符,避免正则报错:

town_patterns <- str_c("\\b", str_escape(df2$Town), "\\b")

内容的提问来源于stack exchange,提问作者Edward Blackburn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 02:13:22