You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于CLV分数创建自定义聚类名称列(K-Means聚类场景)

解决方案:基于CLV分数自定义客户聚类名称

核心思路

先将带格式的CLV分数转换为可排序的数值类型,根据CLV从高到低的顺序为每个原始聚类分配用户自定义名称,最后按CLV降序排序DataFrame以匹配期望输出。


1. 数据初始化与预处理

构造原始DataFrame,并将带千分位逗号的CLV Score转换为整数,确保后续排序逻辑正确:

import pandas as pd

# 原始数据
df = pd.DataFrame({
    'Cluster': [2, 3, 1],
    'CLV Score': ['571,038', '1,474,358', '568,211']
})

# 转换CLV Score为整数:去除逗号后转类型
df['CLV Score'] = df['CLV Score'].str.replace(',', '').astype(int)

2. 自定义聚类名称映射

用户可根据需求自定义名称列表,列表顺序直接对应CLV从高到低的聚类:

# 自定义名称列表:按CLV降序排列,依次对应最高、次高、最低CLV的聚类
custom_cluster_names = ['Gold Customer', 'Silver Customer', 'Dormant Customer']

# 按CLV降序排序,为每个Cluster分配对应名称
df_sorted = df.sort_values(by='CLV Score', ascending=False)
df_sorted['Cluster Name'] = custom_cluster_names

3. 恢复CLV格式(可选)

如果需要将CLV Score转回带千分位的字符串格式,执行以下代码:

# 恢复千分位格式
df_sorted['CLV Score'] = df_sorted['CLV Score'].apply(lambda x: f"{x:,}")

# 重置索引(可选,让索引从0开始)
df_sorted = df_sorted.reset_index(drop=True)

最终结果

执行后df_sorted的输出与期望一致:

ClusterCLV ScoreCluster Name
31,474,358Gold Customer
2571,038Silver Customer
1568,211Dormant Customer

关键说明

  • 自定义灵活性:只需修改custom_cluster_names列表的内容和顺序,即可调整聚类名称的映射关系
  • 预处理必要性:必须将CLV Score转换为数值类型,否则字符串排序会出现逻辑错误(比如"1,474,358"作为字符串会被判定为小于"571,038")
  • 排序逻辑:通过sort_values按CLV降序排列,确保名称与CLV高低一一对应

内容的提问来源于stack exchange,提问作者Jonathan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 01:20:38