You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python(pandas)用loc()做客户画像如何实现同user_id标签统一赋值

问题根因

你现有逻辑是按单条商品交易记录匹配标签,只有用户购买了对应部门商品的记录行会命中非Other标签,其余行自然被标记为Other,不符合「用户级统一标签」的需求。

解决思路

先在用户维度完成标签判定(只要该用户满足属性要求+曾经买过对应部门的商品,就统一打对应标签),再将标签映射回全量交易记录即可,同时要提前定义标签优先级避免同一用户命中多个标签。

实现代码

步骤1:聚合生成用户维度特征表

# 按user_id聚合,提取每个用户的固定属性,以及对应部门的购买标记
user_df = final_df.groupby('user_id').agg(
    # 同一个用户的固定属性取第一条即可
    age = ('age', 'first'),
    age_range = ('age_range', 'first'),
    income = ('income', 'first'),
    parental_status = ('parental_status', 'first'),
    # 标记用户是否购买过对应画像要求的部门商品
    has_high_earner_dept = ('department_id', lambda x: x.isin([1,4,7,19,16]).any()),
    has_young_single_dept = ('department_id', lambda x: x.isin([1,4,7,19]).any()),
    has_young_parent_dept = ('department_id', lambda x: x.isin([4,13,16,17,18]).any()),
    has_over60_dept = ('department_id', lambda x: x.isin([1,4,11,12,15,20]).any())
).reset_index()

步骤2:为每个用户打统一标签

按你给出的标签顺序设置优先级,优先级高的标签先判定,已命中高优先级标签的用户不会再被低优先级标签覆盖:

# 初始化标签为Other
user_df['customer_profile'] = 'Other'

# 优先级1:Higher earner
user_df.loc[
    (user_df['age_range'].isin(['40-49', '50-59'])) & 
    (user_df['income'] >= 400000) & 
    (user_df['has_high_earner_dept']) & 
    (user_df['parental_status'] == 'Parent'),
    'customer_profile'
] = 'Higher earner'

# 优先级2:Young single adult
user_df.loc[
    (user_df['age'] <= 39) & 
    (user_df['income'] <= 199999) & 
    (user_df['has_young_single_dept']) & 
    (user_df['parental_status'] == 'Non-parent') &
    (user_df['customer_profile'] == 'Other'),
    'customer_profile'
] = 'Young single adult'

# 优先级3:Young parent
user_df.loc[
    (user_df['age_range'].isin(['20-29', '30-39'])) & 
    (user_df['income'] <= 199999) & 
    (user_df['has_young_parent_dept']) & 
    (user_df['parental_status'] == 'Parent') &
    (user_df['customer_profile'] == 'Other'),
    'customer_profile'
] = 'Young parent'

# 优先级4:Over 60
user_df.loc[
    (user_df['age'] >= 60) & 
    (user_df['income'] <= 199999) & 
    (user_df['has_over60_dept']) & 
    (user_df['parental_status'] == 'Parent') &
    (user_df['customer_profile'] == 'Other'),
    'customer_profile'
] = 'Over 60'

步骤3:将用户标签映射回原交易表

# 左连接将用户标签同步到所有交易记录
final_df = final_df.merge(
    user_df[['user_id', 'customer_profile']],
    on='user_id',
    how='left',
    suffixes=('', '_new')
)

# 替换原标签列,删除冗余字段
final_df['customer_profile'] = final_df['customer_profile_new']
final_df.drop('customer_profile_new', axis=1, inplace=True)

额外说明

该方案仅需在10万+的用户维度做计算,远比对3000万条交易记录逐行计算效率更高,内存占用也更低。如果你需要调整标签优先级,只需要调整各标签的判定顺序即可。

内容的提问来源于stack exchange,提问作者merlin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.23 23:54:01