You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中基于字符串匹配关联Dataframe的方法求助

基于字符串匹配的高效左关联实现(dplyr 1.1+)

需要将两个业务大表做左关联,关联规则是df_1里的关键词小字符串,是否出现在df_2的大字符串字段中。之前用tidyr::crossing先做全量交叉再过滤的方法,会生成超大中间表,性能极差,现在想用dplyr 1.1版本的join_by结合str_detect来实现高效关联。

示例数据

df_1 <- data.frame(index = 1:5,
                   keyword = c("john", "ella", "mil", "nin", "billi"))

df_2 <- data.frame(index_2 = 1001:1008,
                   name = c("John Coltrane", "Ella Fitzgerald", "Miles Davis", "Billie Holliday", 
                            "Nina Simone", "Bob Smith", "John Brown", "Tony Montana"))

# 预期输出结果
df_results_i_want <- data.frame(index = c(1, 1:5),
                                keyword = c("john", "john", "ella", "mil", "nin", "billi"),
                                index_2 = c(1001, 1007, 1002, 1003, 1005, 1004),
                                name    = c("John Coltrane", "John Brown", "Ella Fitzgerald", 
                                            "Miles Davis", "Nina Simone", "Billie Holliday"))

高效实现代码

library(tidyverse)

df_results <- df_1 |>
  left_join(df_2, 
            join_by(str_detect(name, fixed(keyword, ignore_case = TRUE))))

说明

  • dplyr 1.1.0及以后版本,支持在join_by()中直接使用逻辑表达式作为关联条件,无需生成全量笛卡尔积
  • 这里用fixed(keyword, ignore_case = TRUE)是为了做大小写不敏感的精确子串匹配,既避免正则表达式特殊字符的干扰,匹配效率也更高
  • 这种方式会在关联过程中直接过滤匹配项,不会产生超大中间表,相比crossing+过滤的方案,内存占用和运行速度都有质的提升

内容的提问来源于stack exchange,提问作者Alan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 03:50:41