You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用pandas提取客户满足推荐关联规则的连续浏览产品对

连续浏览有效推荐产品对提取方案

核心匹配规则

按客户维度将浏览记录按时间戳升序排列,取相邻两次浏览的产品配对:若后一次浏览的产品ID 精确存在于 前一次浏览产品的推荐商品列表中,则该配对为有效记录。

注意:禁止直接对横杠拼接的推荐字符串做子串匹配,必须拆分后做全值精确匹配,避免出现ID部分重合导致的误判(比如推荐ID为11时误匹配111、推荐ID为111时误匹配11类问题)

实现步骤

  • 读取原始浏览数据表,确保字段包含customer_id、product_id、recommendation_items、time_stamp
  • 按customer_id分组,组内记录按time_stamp从小到大排序,还原客户真实浏览顺序
  • 对每条记录,将横杠分隔的recommendation_items字段通过split('-')拆分为列表,再转为集合类型,用于后续精确匹配
  • 遍历每个客户的排序后浏览序列,将相邻记录两两配对:前序记录的product_id为product_id_1,后序记录的product_id为product_id_2
  • 校验配对有效性:判断product_id_2是否存在于前序记录拆分好的推荐商品集合中,保留符合条件的配对
  • 整理有效配对,输出仅包含customer_id、product_id_1、product_id_2三个字段的结果表

可运行Python代码(基于Pandas)

import pandas as pd

# 1. 构造样例输入数据
sample_data = [
    {"customer_id": 1, "product_id": 123, "recommendation_items": "111-222-333-444", "time_stamp": "2024-01-01 10:00:00"},
    {"customer_id": 1, "product_id": 111, "recommendation_items": "123-555", "time_stamp": "2024-01-01 10:05:00"},
    {"customer_id": 2, "product_id": 213, "recommendation_items": "234-345-456", "time_stamp": "2024-01-01 11:00:00"},
    {"customer_id": 2, "product_id": 987, "recommendation_items": "213-777", "time_stamp": "2024-01-01 11:10:00"}
]
df = pd.DataFrame(sample_data)

# 2. 数据预处理
# 时间戳转时间类型,保证排序准确
df["time_stamp"] = pd.to_datetime(df["time_stamp"])
# 拆分推荐列表为集合,做精确匹配(核心:解决子串误匹配问题)
df["rec_set"] = df["recommendation_items"].apply(lambda x: set(x.split("-")))
# 按客户分组、组内按时间升序排序
df = df.sort_values(by=["customer_id", "time_stamp"], ascending=[True, True]).reset_index(drop=True)

# 3. 相邻记录配对+匹配
valid_pairs = []
for cust_id, group in df.groupby("customer_id", group_keys=False):
    # 组内记录已经按时间排序,错位拼接相邻记录
    for i in range(len(group) - 1):
        prev_row = group.iloc[i]
        curr_row = group.iloc[i+1]
        # 精确判断:后一个浏览的商品是否在前一个的推荐集合里,统一转字符串避免类型不匹配
        if str(curr_row["product_id"]) in prev_row["rec_set"]:
            valid_pairs.append({
                "customer_id": cust_id,
                "product_id_1": prev_row["product_id"],
                "product_id_2": curr_row["product_id"]
            })

# 4. 输出结果
result_df = pd.DataFrame(valid_pairs)
print(result_df)

代码运行输出

customer_id  product_id_1  product_id_2
0            1           123           111

和预期结果完全一致,不会出现子串误匹配问题。

内容的提问来源于stack exchange,提问作者RajeshM

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 21:39:05