使用Python机器学习基于用户喜好实现用户聚类(附喜好数据)
Alright, let's walk through solving this user clustering task step by step. First, I'll assume we have a clear mapping of user IDs to their respective preference items (since the original data lists 10 preferences, I'll pair each with a unique user ID for this example).
Step 1: Structure Your Input Data
First, let's formalize the user-preference data into a format we can work with. If your actual data has users with multiple preferences, just adjust the preferences column to have space-separated strings (e.g., "milk coffee" for a user who likes both).
import pandas as pd # Create a DataFrame with user IDs and their preferences user_data = pd.DataFrame({ "user_id": [f"user{i}" for i in range(1, 11)], "preferences": ["milk", "coffee", "tea", "sugar curd", "salt", "sugar milk", "sugar", "tea", "curd", "rice"] })
Step 2: Convert Text Preferences to Numerical Features
Clustering models need numerical input, so we'll use Count Vectorization to turn each preference into a binary feature (1 if the user has that preference, 0 otherwise). This works even for preferences with spaces like "sugar curd".
from sklearn.feature_extraction.text import CountVectorizer # Initialize the vectorizer (handles multi-word preferences) vectorizer = CountVectorizer(token_pattern=r"\b\w+\b") # Transform the preference text into a numerical feature matrix X = vectorizer.fit_transform(user_data["preferences"]) # Optional: View the feature matrix to verify feature_df = pd.DataFrame(X.toarray(), columns=vectorizer.get_feature_names_out()) print("Feature Matrix:\n", feature_df.head())
Step 3: Choose & Train a Clustering Model
K-Means is a great starting point for this task—it's simple, fast, and easy to interpret. First, we'll find the optimal number of clusters using the silhouette score (higher values mean better-defined clusters).
from sklearn.cluster import KMeans from sklearn.metrics import silhouette_score # Test different cluster numbers (k) silhouette_scores = [] k_candidates = range(2, 6) for k in k_candidates: kmeans = KMeans(n_clusters=k, random_state=42) cluster_labels = kmeans.fit_predict(X) score = silhouette_score(X, cluster_labels) silhouette_scores.append(score) print(f"Silhouette Score for k={k}: {score:.4f}") # Pick the k with the highest silhouette score (let's say k=3 for this example) optimal_k = 3 kmeans = KMeans(n_clusters=optimal_k, random_state=42) user_data["cluster_label"] = kmeans.fit_predict(X)
Step 4: Analyze the Clustering Results
Now let's look at which users ended up in which clusters, and what their preferences have in common.
# View the full results print("\nFinal Clustering Results:\n") print(user_data[["user_id", "preferences", "cluster_label"]]) # Group users by cluster to see common preferences cluster_groups = user_data.groupby("cluster_label")["preferences"].apply(list) print("\nCluster Breakdown:\n") for cluster_id, preferences in cluster_groups.items(): print(f"Cluster {cluster_id}: {', '.join(preferences)}")
Example Output Interpretation
When you run this code, you'll likely see clusters like:
- Cluster 0: sugar curd, sugar milk, sugar, curd → Groups users who prefer sugar-related items or curd
- Cluster 1: tea, tea → Groups users who like tea
- Cluster 2: milk, coffee, salt, rice → Groups users with more general, unrelated preferences
If your actual data has users with multiple preferences, the clusters will naturally group users with overlapping taste profiles (e.g., a user who likes "milk" and "sugar milk" would cluster with other sugar/dairy lovers).
Notes for Adaptation
- If you have users with multiple preferences per ID, just join their preferences into a single string (e.g.,
"milk sugar"instead of a list) before vectorizing. - Try other clustering algorithms like DBSCAN if you want to avoid pre-defining the number of clusters—it's great for finding arbitrary-shaped groups.
内容的提问来源于stack exchange,提问作者Manish Kammar

