Pandas技术问题:跨DataFrame列子串位置查找及产品子串匹配结果全为False的排查
问题排查与解决方案
你遇到的问题核心在于错误地将整个df2['Product2'] Series转换成了字符串,而不是逐行对应匹配两个DataFrame的元素。让我拆解一下问题,再给出修正方案:
问题根源分析
你的代码里这一行是关键错误:
df1['productMatch?'] = np.where(df1['Product1'].str.find(str(df2['Product2'])) !=-1, 'True', 'False')
当你执行str(df2['Product2'])时,得到的不是逐行的Product2值,而是整个Series的字符串化结果,类似这样:
0 PPP 1 Ph 2 Printer 3 Desk Name: Product2, dtype: object
所以df1['Product1'].str.find(...)实际上是在每个Product1字符串里查找这个长文本,显然找不到匹配项,结果自然全是-1,最终所有匹配结果都是False。
修正方案
我们需要实现逐行对应匹配:检查df1每一行的Product1是否包含df2对应行的Product2子串。这里提供两种高效的实现方式:
方法1:使用apply逐行处理(直观易懂)
import pandas as pd import numpy as np firstProductSet = {'Product1':['Computer','Phone','Printer','Desk'],'Price1':[1200,800,200,350]} df1 = pd.DataFrame(firstProductSet,columns= ['Product1', 'Price1']) secondProductSet = {'Product2': ['PPP','Ph','Printer','Desk'],'Price2':[900,800,300,350]} df2 = pd.DataFrame(secondProductSet,columns= ['Product2', 'Price2']) # 逐行检查df2的Product2是否是df1对应行Product1的子串 df1['productMatch?'] = df1.apply(lambda row: df2.loc[row.name, 'Product2'] in row['Product1'], axis=1) print(df1)
方法2:向量化字符串匹配(性能更优)
利用Pandas的向量化str.contains方法,直接传入df2['Product2']的数值数组:
df1['productMatch?'] = df1['Product1'].str.contains(df2['Product2'].values)
运行结果验证
执行修正后的代码,输出的df1会完全符合你的预期:
Product1 Price1 productMatch? 0 Computer 1200 False 1 Phone 800 True 2 Printer 200 True 3 Desk 350 True
这样就完美实现了你想要的子串匹配效果啦!
内容的提问来源于stack exchange,提问作者Kalia
相关产品推荐
相关产品推荐

