R语言双数据框无顺序短语匹配及匹配率统计实现方法
R语言双数据框短语匹配占比计算实现
需求说明
将df_1的col_1列与df_2的col_2列按逗号拆分后进行短语匹配,匹配规则如下:
- 不区分短语先后顺序
- 短语完全一致判定为匹配
最终输出包含col_1、col_2、分数格式匹配占比、数值化百分比匹配率的结果数据框df_out。
实现代码
依赖包
# 需提前安装加载以下依赖包 library(dplyr) library(purrr) library(stringr) library(tidyr)
示例输入数据
col_1 = c("inside the world,worldwide web,google chrome app","world health organisation, from country", "team work","size of the country, bigger compared to other") df_1 = data.frame(col_1) col_2 = c("team work,in the company", "size of the country","inside the world,worldwide web,google chrome app", "google chrome app,worldwide web,inside the world","inside the world,google chrome app", "web worldwide","world health organisation, from country","from country", "size of the country, bigger compared to other","country from", "world health organisation,country from") df_2 = data.frame(col_2)
核心处理逻辑
df_out <- crossing(df_1, df_2) %>% # 拆分短语并去除前后多余空格 mutate( col1_phrases = map(col_1, ~ str_split(.x, ",")[[1]] %>% str_trim()), col2_phrases = map(col_2, ~ str_split(.x, ",")[[1]] %>% str_trim()), # 计算匹配数、总短语数 match_num = map2_int(col1_phrases, col2_phrases, ~ length(intersect(.x, .y))), total_num = map_int(col1_phrases, length), # 生成要求的两个匹配率字段 match_percentage = paste0(match_num, "/", total_num), match_numeric_percentage = round(match_num / total_num * 100, 4) ) %>% # 筛选输出需要的字段 select(col_1, col_2, match_percentage, match_numeric_percentage)
规则调整说明
- 若需要不区分大小写匹配,可在拆分短语后加上
str_to_lower()统一转小写,示例:str_split(.x, ",")[[1]] %>% str_trim() %>% str_to_lower() - 若需要匹配率分母取col_2的短语总数,将
total_num的计算逻辑改为map_int(col2_phrases, length)即可 - 若需要调整百分比保留的小数位数,修改
round函数的第二个参数即可
内容的提问来源于stack exchange,提问作者san1
相关产品推荐
相关产品推荐

