如何检测两个DataFrame中的匹配字符串片段并提取该片段?
问题:匹配并提取两个DataFrame中的共同名称片段
输入数据
DF_1
id x 1 eu continent hamburg 2 asia singapore 3 austrlia hedland
DF_2
name germany hamburg singapore china west australia hedland
需求
检测两个DataFrame中存在匹配的名称片段,提取该片段后合并输出,期望结果:
id x name 1 eu continent hamburg hamburg 2 asia singapore singapore 3 austrlia hedland hedland
尝试用str.contains但需要遍历完整字符串,寻求实现方法。
解决方案
可以通过拆分文本为单词集合,利用集合交集快速定位匹配片段,再合并结果。以下是基于pandas的实现代码:
import pandas as pd # 构造输入数据 df1 = pd.DataFrame({ 'id': [1,2,3], 'x': ['eu continent hamburg', 'asia singapore', 'austrlia hedland'] }) df2 = pd.DataFrame({ 'name': ['germany hamburg', 'singapore china', 'west australia hedland'] }) # 将df2每行文本转为单词集合 df2['word_set'] = df2['name'].str.split().apply(set) # 定义函数:从df1的x字段中找出与对应df2行匹配的单词 def get_matched_word(x_str, target_words): x_words = set(x_str.split()) matched = x_words & target_words return next(iter(matched), None) # 取首个匹配项(假设每行仅一个匹配) # 按索引遍历处理每行,提取匹配片段 df1['name'] = [get_matched_word(df1.loc[i, 'x'], df2.loc[i, 'word_set']) for i in df1.index] # 输出结果 print(df1.to_string(index=False))
代码说明
- 把两个DataFrame的文本字段拆分为单词集合,利用集合交集操作快速定位共同片段;
- 基于输入数据的行对应关系(df1的第n行对应df2的第n行),按索引逐行处理;
- 若存在多个匹配单词,可修改逻辑返回所有匹配项,当前代码默认取首个匹配片段。
内容的提问来源于stack exchange,提问作者Tmiskiewicz
相关产品推荐
相关产品推荐

