You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Pandas实现热门商品推荐系统:现有代码优化方案咨询

问题描述

现有两个DataFrame:

df1

cust_id           product_list
1                 ['phone', 'tv']
2                 ['ball', 'bat']
3                 ['bat']
4                 ['ball', 'bat', 'phone', 'tv']
5                 ['tv']
6                 ['bat']
7                 ['ball', 'bat', 'phone', 'tv']       
8                 ['phone', 'tv']

df2

product          support
ball             0.18
bat              0.29
phone            0.24
tv               0.29   

需要生成如下目标DataFrame:

预期输出

cust_id      product_list             recommended_dictionary
0   1            [phone, tv]              {'bat': 0.29, 'ball': 0.18}
1   2            [ball, bat]              {'tv': 0.29, 'phone': 0.24}
2   3            [bat]                    {'tv': 0.29, 'phone': 0.24, 'ball': 0.18}
3   4            [ball, bat, phone, tv]   {}
4   5            [tv]                     {'bat': 0.29, 'phone': 0.24, 'ball': 0.18}
5   6            [bat]                    {'tv': 0.29, 'phone': 0.24, 'ball': 0.18}
6   7            [ball, bat, phone, tv]   {}
7   8            [phone, tv]              {'bat': 0.29, 'ball': 0.18}
现有可行代码

已实现的可正常运行代码如下:

创建热门度字典

popularity_dict = dict(df2.values)

定义推荐函数:按热门度推荐用户未购买的商品

def f(x):
    out = {}
    dif = [i for i in popularity_dict.keys() if i not in x]
    for i in dif:
        out[i] = popularity_dict[i]
        
    return out

计算推荐字典

df1['popularity_based_recommended_dictionary'] = df1['product_list'].apply(f)
更优实现方式

可以从效率和代码简洁性两方面优化:

优化后的代码

# 1. 构建热门度字典(这部分和原代码一致)
popularity_dict = dict(df2.values)

# 2. 简化推荐逻辑:用集合查找+字典推导式
def recommend_unowned(x):
    owned = set(x)
    return {prod: score for prod, score in popularity_dict.items() if prod not in owned}

# 3. 应用到DataFrame
df1['recommended_dictionary'] = df1['product_list'].apply(recommend_unowned)

优化点说明

  1. 成员检查效率提升:将用户已购商品列表转为集合set(x),集合的成员查找操作是O(1),远快于列表的O(n),当商品数量或用户规模较大时,性能提升明显。
  2. 代码简洁性提升:用字典推导式直接构建结果字典,替代原有的列表推导+循环赋值,代码更紧凑且可读性更强。
  3. 可选:如果追求极致简洁,也可以用lambda表达式替代函数定义:
df1['recommended_dictionary'] = df1['product_list'].apply(
    lambda x: {p: s for p, s in popularity_dict.items() if p not in set(x)}
)

内容的提问来源于stack exchange,提问作者Danish

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 17:40:38