You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于字符串部分匹配实现两个DataFrame的内连接

Pandas实现基于字符串前缀/包含匹配的内连接

问题场景

现有两个DataFrame,需要通过temp的message字段以temp_truncated的message字段开头(或包含该字符串)的规则完成内连接,得到目标结果。

原DataFrame定义:

import pandas as pd
import numpy as np

temp = pd.DataFrame(np.array([['I am feeling very well',1],['It is hard to believe this happened',0],
                              ['What is love?',1], ['No new friends',0],
                              ['I love this show',1],['Amazing day today',1]]),
                    columns = ['message','sentiment'])

temp_truncated = pd.DataFrame(np.array([['I am feeling very',1],['It is hard to believe',1],
                                        ['What is',1], ['Amazing day',1]]),
                              columns = ['message','cutoff'])

解决方案

由于Pandas原生merge仅支持精确匹配,我们可以通过交叉连接+条件筛选实现模糊匹配的内连接:

# 1. 添加临时键实现交叉连接,生成所有可能的记录组合
temp['temp_key'] = 1
temp_truncated['temp_key'] = 1
cross_merged = pd.merge(temp, temp_truncated, on='temp_key')

# 2. 筛选符合前缀匹配的记录(如需包含匹配,替换为str.contains即可)
filtered = cross_merged[cross_merged['message_x'].str.startswith(cross_merged['message_y'])]

# 3. 整理结果列与索引,得到目标格式
final_result = filtered.rename(columns={'message_x': 'message'})[['message', 'sentiment', 'cutoff']]
final_result = final_result.reset_index(drop=True)

输出结果

执行上述代码后,final_result即为需求中的目标DataFrame:

message sentiment cutoff
0              I am feeling very well         1      1
1  It is hard to believe this happened         0      1
2                      What is love?         1      1
3                   Amazing day today         1      1

扩展说明

如果需要改为包含匹配而非前缀匹配,只需将筛选条件中的str.startswith替换为str.contains:

filtered = cross_merged[cross_merged['message_x'].str.contains(cross_merged['message_y'])]

内容的提问来源于stack exchange,提问作者DarknessPlusPlus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 14:50:44