You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python机器学习基于用户喜好实现用户聚类(附喜好数据)

User Clustering Based on Preference Data Using Python

Alright, let's walk through solving this user clustering task step by step. First, I'll assume we have a clear mapping of user IDs to their respective preference items (since the original data lists 10 preferences, I'll pair each with a unique user ID for this example).

Step 1: Structure Your Input Data

First, let's formalize the user-preference data into a format we can work with. If your actual data has users with multiple preferences, just adjust the preferences column to have space-separated strings (e.g., "milk coffee" for a user who likes both).

import pandas as pd

# Create a DataFrame with user IDs and their preferences
user_data = pd.DataFrame({
    "user_id": [f"user{i}" for i in range(1, 11)],
    "preferences": ["milk", "coffee", "tea", "sugar curd", "salt", "sugar milk", "sugar", "tea", "curd", "rice"]
})

Step 2: Convert Text Preferences to Numerical Features

Clustering models need numerical input, so we'll use Count Vectorization to turn each preference into a binary feature (1 if the user has that preference, 0 otherwise). This works even for preferences with spaces like "sugar curd".

from sklearn.feature_extraction.text import CountVectorizer

# Initialize the vectorizer (handles multi-word preferences)
vectorizer = CountVectorizer(token_pattern=r"\b\w+\b")

# Transform the preference text into a numerical feature matrix
X = vectorizer.fit_transform(user_data["preferences"])

# Optional: View the feature matrix to verify
feature_df = pd.DataFrame(X.toarray(), columns=vectorizer.get_feature_names_out())
print("Feature Matrix:\n", feature_df.head())

Step 3: Choose & Train a Clustering Model

K-Means is a great starting point for this task—it's simple, fast, and easy to interpret. First, we'll find the optimal number of clusters using the silhouette score (higher values mean better-defined clusters).

from sklearn.cluster import KMeans
from sklearn.metrics import silhouette_score

# Test different cluster numbers (k)
silhouette_scores = []
k_candidates = range(2, 6)

for k in k_candidates:
    kmeans = KMeans(n_clusters=k, random_state=42)
    cluster_labels = kmeans.fit_predict(X)
    score = silhouette_score(X, cluster_labels)
    silhouette_scores.append(score)
    print(f"Silhouette Score for k={k}: {score:.4f}")

# Pick the k with the highest silhouette score (let's say k=3 for this example)
optimal_k = 3
kmeans = KMeans(n_clusters=optimal_k, random_state=42)
user_data["cluster_label"] = kmeans.fit_predict(X)

Step 4: Analyze the Clustering Results

Now let's look at which users ended up in which clusters, and what their preferences have in common.

# View the full results
print("\nFinal Clustering Results:\n")
print(user_data[["user_id", "preferences", "cluster_label"]])

# Group users by cluster to see common preferences
cluster_groups = user_data.groupby("cluster_label")["preferences"].apply(list)
print("\nCluster Breakdown:\n")
for cluster_id, preferences in cluster_groups.items():
    print(f"Cluster {cluster_id}: {', '.join(preferences)}")

Example Output Interpretation

When you run this code, you'll likely see clusters like:

  • Cluster 0: sugar curd, sugar milk, sugar, curd → Groups users who prefer sugar-related items or curd
  • Cluster 1: tea, tea → Groups users who like tea
  • Cluster 2: milk, coffee, salt, rice → Groups users with more general, unrelated preferences

If your actual data has users with multiple preferences, the clusters will naturally group users with overlapping taste profiles (e.g., a user who likes "milk" and "sugar milk" would cluster with other sugar/dairy lovers).

Notes for Adaptation

  • If you have users with multiple preferences per ID, just join their preferences into a single string (e.g., "milk sugar" instead of a list) before vectorizing.
  • Try other clustering algorithms like DBSCAN if you want to avoid pre-defining the number of clusters—it's great for finding arbitrary-shaped groups.

内容的提问来源于stack exchange,提问作者Manish Kammar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:18:39