You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas技术问题:跨DataFrame列子串位置查找及产品子串匹配结果全为False的排查

问题排查与解决方案

你遇到的问题核心在于错误地将整个df2['Product2'] Series转换成了字符串,而不是逐行对应匹配两个DataFrame的元素。让我拆解一下问题,再给出修正方案:

问题根源分析

你的代码里这一行是关键错误:

df1['productMatch?'] = np.where(df1['Product1'].str.find(str(df2['Product2'])) !=-1, 'True', 'False')

当你执行str(df2['Product2'])时,得到的不是逐行的Product2值,而是整个Series的字符串化结果,类似这样:

0      PPP
1       Ph
2    Printer
3      Desk
Name: Product2, dtype: object

所以df1['Product1'].str.find(...)实际上是在每个Product1字符串里查找这个长文本,显然找不到匹配项,结果自然全是-1,最终所有匹配结果都是False。

修正方案

我们需要实现逐行对应匹配:检查df1每一行的Product1是否包含df2对应行的Product2子串。这里提供两种高效的实现方式:

方法1:使用apply逐行处理(直观易懂)

import pandas as pd
import numpy as np

firstProductSet = {'Product1':['Computer','Phone','Printer','Desk'],'Price1':[1200,800,200,350]}
df1 = pd.DataFrame(firstProductSet,columns= ['Product1', 'Price1'])
secondProductSet = {'Product2': ['PPP','Ph','Printer','Desk'],'Price2':[900,800,300,350]}
df2 = pd.DataFrame(secondProductSet,columns= ['Product2', 'Price2'])

# 逐行检查df2的Product2是否是df1对应行Product1的子串
df1['productMatch?'] = df1.apply(lambda row: df2.loc[row.name, 'Product2'] in row['Product1'], axis=1)

print(df1)

方法2:向量化字符串匹配(性能更优)

利用Pandas的向量化str.contains方法,直接传入df2['Product2']的数值数组:

df1['productMatch?'] = df1['Product1'].str.contains(df2['Product2'].values)

运行结果验证

执行修正后的代码,输出的df1会完全符合你的预期:

Product1  Price1  productMatch?
0  Computer    1200          False
1     Phone     800           True
2    Printer     200           True
3       Desk     350           True

这样就完美实现了你想要的子串匹配效果啦!

内容的提问来源于stack exchange,提问作者Kalia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 05:57:37