无目标变量分类:基于5个月还款数据集预测客户行为
Alright, let's break down how to solve this unlabeled customer behavior classification problem—since we don't have a predefined "good/bad" target variable, we'll rely on unsupervised learning and business logic to get the job done. Here's a step-by-step approach tailored to your dataset:
First, let's organize your messy input into a structured format that's easier to work with. Your dataset has 9 customers with two key features:
- Monthly payment amount: 1271, 1172, 3038, 4057, 5059, 4151, 911, 1780, 1368
- Customer group (likely payment timing): D-3, D+6, D-3, D+6, D+3, D, D, D+15, D+3
First, we need to translate the group labels into numerical values for modeling. Based on common credit industry conventions, I'll assume:
D-3= Paid 3 days early → coded as-3D= Paid on time → coded as0D+3= Paid 3 days late → coded as3D+6= Paid 6 days late → coded as6D+15= Paid 15 days late → coded as15
If your group definitions are different, adjust these codes accordingly—this is critical for accurate results.
Since we don't have labeled data, we need to align on what counts as good vs. bad behavior with your business team. For example:
- Good behavior: Customers who pay early (
D-3) or on time (D) - Bad behavior: Customers who pay significantly late (
D+6,D+15) - Neutral: Customers with minor delays (
D+3)—we'll let clustering help decide where these fall
We'll use clustering algorithms to group similar customers, then map those clusters to our good/bad definitions.
3.1 K-Means Clustering (Quick & Practical)
K-Means is great for splitting data into distinct groups. We'll use it to split your 9 customers into 2 clusters (good/bad). Here's a Python implementation:
import pandas as pd from sklearn.cluster import KMeans from sklearn.preprocessing import StandardScaler # Build structured DataFrame data = { "monthly_payment": [1271, 1172, 3038, 4057, 5059, 4151, 911, 1780, 1368], "group_code": [-3, 6, -3, 6, 3, 0, 0, 15, 3] } df = pd.DataFrame(data) # Standardize features (monthly payment varies widely, so scaling helps) scaler = StandardScaler() scaled_features = scaler.fit_transform(df) # Run K-Means with 2 clusters kmeans = KMeans(n_clusters=2, random_state=42) df["behavior_cluster"] = kmeans.fit_predict(scaled_features) # Analyze cluster characteristics print("Cluster Summary:") print(df.groupby("behavior_cluster").agg({ "monthly_payment": ["mean", "min", "max"], "group_code": ["mean", "min", "max"] }))
Interpreting the Results:
- Look at the average
group_codefor each cluster:- Cluster with a lower average (e.g., negative/close to 0) will correspond to good behavior (early/on-time payers)
- Cluster with a higher average (positive, large numbers) will correspond to bad behavior (late payers)
3.2 Hierarchical Clustering (For Visual Validation)
If you want to see how customers are grouped step-by-step, hierarchical clustering with a dendrogram is perfect. It lets you manually cut the tree to form your 2 groups:
from scipy.cluster.hierarchy import dendrogram, linkage import matplotlib.pyplot as plt # Generate linkage matrix linked = linkage(scaled_features, 'ward') # Plot dendrogram plt.figure(figsize=(10, 5)) dendrogram( linked, orientation='top', labels=[f"Customer {i+1}" for i in df.index], distance_sort='descending', show_leaf_counts=True ) plt.title("Customer Behavior Hierarchical Clustering") plt.xlabel("Customer") plt.ylabel("Distance") plt.show()
Interpreting the Results:
- Look for a clear split in the dendrogram that separates customers with early/on-time payments from those with late payments. This visual can help you validate the K-Means results.
Once you have your clusters, cross-reference them with your business logic:
- If a cluster has mostly
D-3andDcustomers, label all of them as Good - If a cluster has mostly
D+6andD+15customers, label them as Bad - For
D+3customers who fall into either cluster, use the cluster's overall behavior to assign their label (e.g., if they're grouped with on-time payers, count them as Good; if with late payers, count as Bad)
- Always confirm the group definitions (
D-3,D+6, etc.) with your business team—wrong assumptions here will break your classification. - Since this is unsupervised, the results are a starting point. You may need to adjust features (e.g., add payment frequency if you had it) or cluster numbers if the initial split doesn't make sense.
内容的提问来源于stack exchange,提问作者Melwyn J

