You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无目标变量分类:基于5个月还款数据集预测客户行为

Alright, let's break down how to solve this unlabeled customer behavior classification problem—since we don't have a predefined "good/bad" target variable, we'll rely on unsupervised learning and business logic to get the job done. Here's a step-by-step approach tailored to your dataset:

Step 1: Clean & Structure Your Raw Data

First, let's organize your messy input into a structured format that's easier to work with. Your dataset has 9 customers with two key features:

  • Monthly payment amount: 1271, 1172, 3038, 4057, 5059, 4151, 911, 1780, 1368
  • Customer group (likely payment timing): D-3, D+6, D-3, D+6, D+3, D, D, D+15, D+3

First, we need to translate the group labels into numerical values for modeling. Based on common credit industry conventions, I'll assume:

  • D-3 = Paid 3 days early → coded as -3
  • D = Paid on time → coded as 0
  • D+3 = Paid 3 days late → coded as 3
  • D+6 = Paid 6 days late → coded as 6
  • D+15 = Paid 15 days late → coded as 15

If your group definitions are different, adjust these codes accordingly—this is critical for accurate results.

Step 2: Define "Good/Bad" Behavior Using Business Logic

Since we don't have labeled data, we need to align on what counts as good vs. bad behavior with your business team. For example:

  • Good behavior: Customers who pay early (D-3) or on time (D)
  • Bad behavior: Customers who pay significantly late (D+6, D+15)
  • Neutral: Customers with minor delays (D+3)—we'll let clustering help decide where these fall
Step 3: Apply Unsupervised Classification Methods

We'll use clustering algorithms to group similar customers, then map those clusters to our good/bad definitions.

3.1 K-Means Clustering (Quick & Practical)

K-Means is great for splitting data into distinct groups. We'll use it to split your 9 customers into 2 clusters (good/bad). Here's a Python implementation:

import pandas as pd
from sklearn.cluster import KMeans
from sklearn.preprocessing import StandardScaler

# Build structured DataFrame
data = {
    "monthly_payment": [1271, 1172, 3038, 4057, 5059, 4151, 911, 1780, 1368],
    "group_code": [-3, 6, -3, 6, 3, 0, 0, 15, 3]
}
df = pd.DataFrame(data)

# Standardize features (monthly payment varies widely, so scaling helps)
scaler = StandardScaler()
scaled_features = scaler.fit_transform(df)

# Run K-Means with 2 clusters
kmeans = KMeans(n_clusters=2, random_state=42)
df["behavior_cluster"] = kmeans.fit_predict(scaled_features)

# Analyze cluster characteristics
print("Cluster Summary:")
print(df.groupby("behavior_cluster").agg({
    "monthly_payment": ["mean", "min", "max"],
    "group_code": ["mean", "min", "max"]
}))

Interpreting the Results:

  • Look at the average group_code for each cluster:
    • Cluster with a lower average (e.g., negative/close to 0) will correspond to good behavior (early/on-time payers)
    • Cluster with a higher average (positive, large numbers) will correspond to bad behavior (late payers)

3.2 Hierarchical Clustering (For Visual Validation)

If you want to see how customers are grouped step-by-step, hierarchical clustering with a dendrogram is perfect. It lets you manually cut the tree to form your 2 groups:

from scipy.cluster.hierarchy import dendrogram, linkage
import matplotlib.pyplot as plt

# Generate linkage matrix
linked = linkage(scaled_features, 'ward')

# Plot dendrogram
plt.figure(figsize=(10, 5))
dendrogram(
    linked,
    orientation='top',
    labels=[f"Customer {i+1}" for i in df.index],
    distance_sort='descending',
    show_leaf_counts=True
)
plt.title("Customer Behavior Hierarchical Clustering")
plt.xlabel("Customer")
plt.ylabel("Distance")
plt.show()

Interpreting the Results:

  • Look for a clear split in the dendrogram that separates customers with early/on-time payments from those with late payments. This visual can help you validate the K-Means results.
Step 4: Refine & Assign Final Labels

Once you have your clusters, cross-reference them with your business logic:

  • If a cluster has mostly D-3 and D customers, label all of them as Good
  • If a cluster has mostly D+6 and D+15 customers, label them as Bad
  • For D+3 customers who fall into either cluster, use the cluster's overall behavior to assign their label (e.g., if they're grouped with on-time payers, count them as Good; if with late payers, count as Bad)
Key Notes to Remember
  • Always confirm the group definitions (D-3, D+6, etc.) with your business team—wrong assumptions here will break your classification.
  • Since this is unsupervised, the results are a starting point. You may need to adjust features (e.g., add payment frequency if you had it) or cluster numbers if the initial split doesn't make sense.

内容的提问来源于stack exchange,提问作者Melwyn J

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 07:46:35