如何基于文本匹配关联人物与高校数据集?R语言代码纠错
问题原因
直接用left_join返回空结果的核心原因是:left_join基于字段值完全相等的逻辑做关联,但你的需求是从biography的长文本里模糊匹配college_records_ex中的高校名称,两个数据集不存在能直接完全匹配的关联字段,自然无法匹配出结果。
正确实现思路与代码
R语言版本(基于dplyr、stringr)
- 加载依赖包
library(dplyr) library(stringr)
- 生成高校名称匹配规则,从biography中提取匹配的高校名
# 把所有高校名拼接成正则匹配串 college_match_str <- str_c(college_records_ex$college_name, collapse = "|") # 提取高校名后再关联 final_result <- people_records_ex %>% mutate(matched_college = str_extract(biography, college_match_str)) %>% left_join(college_records_ex, by = c("matched_college" = "college_name"))
Python版本(基于pandas)
import pandas as pd import re # 生成高校名称匹配正则串 college_pattern = '|'.join(college_records_ex['college_name']) # 从biography中提取匹配的高校名,再做左关联 people_records_ex['matched_college'] = people_records_ex['biography'].str.extract(f'({college_pattern})') final_result = people_records_ex.merge(college_records_ex, left_on='matched_college', right_on='college_name', how='left')
注意:如果存在高校简称/全称不统一的情况(比如"北大"和"北京大学"),需要先统一名称映射规则,否则会出现匹配遗漏。
内容的提问来源于stack exchange,提问作者wizkids121
相关产品推荐
相关产品推荐

