使用DataFrame.append存问题?嵌套apply生成相似度数据优化方案
问题详情
我用嵌套apply()遍历两个DataFrame,目的是计算新物品与库存物品的相似度得分并生成结构符合要求的新DataFrame,但遇到了以下问题:
初始问题
运行代码时收到警告:
/tmp/ipykernel_126/2064736442.py:1: FutureWarning: The frame.append method is deprecated and will be removed from pandas in a future version. Use pandas.concat instead.
不知道怎么在这个场景下用pandas.concat,而且当前代码生成的DataFrame列显示异常:
1 2 3 4 5 6 7 8 9 10 ... 90 91 92 93 94 95 96 97 98 99 1 Ne... Ne... Ne... Ne... Ne... Ne... Ne... Ne... Ne... Ne... ... Ne... Ne... Ne... Ne... Ne... Ne... Ne... Ne... Ne... Ne...
我想要生成的目标DataFrame结构如下:
test.head() new item inventory item similarity 0 paintbrush broom 22 1 paintbrush mop 15 2 paintbrush bucket 45 3 paintbrush flashlight 4
更新后的问题
我改用字典存储数据,消除了警告、输出格式正确且速度提升,但只能看到一行数据。想问用update()是不是添加多行数据的正确方式?相关代码与输出如下:
dict = {} temp = df_test.apply( lambda x: item_test.apply( lambda y: dict.update( { "New Item": x["ITEM_DESC"], "Inventory Item": y["Item_Desc"], "Similarity": fuzz.ratio( str(x["ITEM_DESC"]).lower(), str(y["Item_Desc"]).lower() ), }, ignore_index=True, ), axis=1, ), axis=1, ) dict
输出:
{'New Item': 'mop', 'Inventory Item': 'paintbrush', 'Similarity': 23, 'ignore_index': True}
原始DataFrame结构
df_test.head():
item 0 paintbrush 1 mop #2 2 red bucket 3 o-light flashlight
item_test.head():
item_desc 0 broom 1 mop 2 bucket 3 flashlight
内容的提问来源于stack exchange,提问作者Zachqwerty
相关产品推荐
相关产品推荐

