You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何检测两个DataFrame中的匹配字符串片段并提取该片段?

问题:匹配并提取两个DataFrame中的共同名称片段

输入数据

DF_1

id  x
1   eu continent hamburg
2   asia singapore
3   austrlia hedland

DF_2

name
germany hamburg
singapore china
west australia hedland

需求

检测两个DataFrame中存在匹配的名称片段,提取该片段后合并输出,期望结果:

id  x                     name
1   eu continent hamburg hamburg
2   asia singapore       singapore
3   austrlia hedland     hedland

尝试用str.contains但需要遍历完整字符串,寻求实现方法。

解决方案

可以通过拆分文本为单词集合,利用集合交集快速定位匹配片段,再合并结果。以下是基于pandas的实现代码:

import pandas as pd

# 构造输入数据
df1 = pd.DataFrame({
    'id': [1,2,3],
    'x': ['eu continent hamburg', 'asia singapore', 'austrlia hedland']
})

df2 = pd.DataFrame({
    'name': ['germany hamburg', 'singapore china', 'west australia hedland']
})

# 将df2每行文本转为单词集合
df2['word_set'] = df2['name'].str.split().apply(set)

# 定义函数:从df1的x字段中找出与对应df2行匹配的单词
def get_matched_word(x_str, target_words):
    x_words = set(x_str.split())
    matched = x_words & target_words
    return next(iter(matched), None)  # 取首个匹配项(假设每行仅一个匹配)

# 按索引遍历处理每行,提取匹配片段
df1['name'] = [get_matched_word(df1.loc[i, 'x'], df2.loc[i, 'word_set']) for i in df1.index]

# 输出结果
print(df1.to_string(index=False))

代码说明

  1. 把两个DataFrame的文本字段拆分为单词集合,利用集合交集操作快速定位共同片段;
  2. 基于输入数据的行对应关系(df1的第n行对应df2的第n行),按索引逐行处理;
  3. 若存在多个匹配单词,可修改逻辑返回所有匹配项,当前代码默认取首个匹配片段。

内容的提问来源于stack exchange,提问作者Tmiskiewicz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 23:59:55