基于多条件创建用户忠诚度标签并更新Pandas DataFrame列
完整实现方案
核心思路是先对每条结账确认记录按用户+时间顺序编号,再根据编号匹配对应标签,因为结账标签优先级更高,放在前两个标签赋值之后执行即可覆盖冲突值。
步骤1:先按用户和时间排序(保证结账次数统计顺序正确)
import pandas as pd import numpy as np # 先对数据按用户ID和时间戳升序排序,确保结账顺序统计准确 test_df = test_df.sort_values(by=['user_id', 'timestamp']).reset_index(drop=True)
步骤2:保留你已有的前两个标签赋值逻辑
test_df['loyalty'] = np.where((test_df['session'] > 0) & ((test_df['type'] != 'checkout:confirmation')), 'frequent_visitor', None) test_df.loc[test_df['session'] == 0, 'loyalty'] = 'first_time_visitor'
步骤3:统计每个用户的结账次数顺序,赋值对应标签
# 给每个用户的结账确认记录按时间顺序编号,从0开始计数 test_df['checkout_seq'] = test_df[test_df['type'] == 'checkout:confirmation'].groupby('user_id').cumcount() # 按次数赋值标签 test_df.loc[test_df['checkout_seq'] == 0, 'loyalty'] = 'first_time_customer' test_df.loc[test_df['checkout_seq'] == 1, 'loyalty'] = 'repeat_customer' test_df.loc[test_df['checkout_seq'] >= 2, 'loyalty'] = 'loyal_customer' # 可选:删除中间辅助列 test_df.drop('checkout_seq', axis=1, inplace=True)
最终输出结果示例
| user_id | timestamp | type | session | count_session_products | loyalty |
|---|---|---|---|---|---|
| 9EPWZVMNP6D6KWX | 1612139269 | productDetails | 0 | 4 | first_time_visitor |
| 9EPWZVMNP6D6KWX | 1612139579 | checkout:confirmation | 2 | 0 | first_time_customer |
| 9EPWZVMNP6D6KWX | 1612139665 | productDetails | 1 | 1 | frequent_visitor |
| 9EPWZVMNP6D6KWX | 1612141096 | checkout:confirmation | 3 | 4 | repeat_customer |
| 9EPWZVMNP6D6KWX | 1612143046 | productList | 4 | 2 | frequent_visitor |
| 9EPWZVMNP6D6KWX | 1612143729 | checkout:confirmation | 5 | 2 | loyal_customer |
内容的提问来源于stack exchange,提问作者Salaaned
相关产品推荐
相关产品推荐

