You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于文本匹配关联人物与高校数据集?R语言代码纠错

问题原因

直接用left_join返回空结果的核心原因是:left_join基于字段值完全相等的逻辑做关联,但你的需求是从biography的长文本里模糊匹配college_records_ex中的高校名称,两个数据集不存在能直接完全匹配的关联字段,自然无法匹配出结果。

正确实现思路与代码

R语言版本(基于dplyr、stringr)

  1. 加载依赖包
library(dplyr)
library(stringr)
  1. 生成高校名称匹配规则,从biography中提取匹配的高校名
# 把所有高校名拼接成正则匹配串
college_match_str <- str_c(college_records_ex$college_name, collapse = "|")

# 提取高校名后再关联
final_result <- people_records_ex %>%
  mutate(matched_college = str_extract(biography, college_match_str)) %>%
  left_join(college_records_ex, by = c("matched_college" = "college_name"))

Python版本(基于pandas)

import pandas as pd
import re

# 生成高校名称匹配正则串
college_pattern = '|'.join(college_records_ex['college_name'])

# 从biography中提取匹配的高校名,再做左关联
people_records_ex['matched_college'] = people_records_ex['biography'].str.extract(f'({college_pattern})')
final_result = people_records_ex.merge(college_records_ex, left_on='matched_college', right_on='college_name', how='left')

注意:如果存在高校简称/全称不统一的情况(比如"北大"和"北京大学"),需要先统一名称映射规则,否则会出现匹配遗漏。

内容的提问来源于stack exchange,提问作者wizkids121

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 14:15:41