如何读取多Steam用户CSV文件并开展k-modes聚类分析?
Great question! When working with thousands of user-specific CSV files, the key is to first aggregate all the data into a single structured dataset where each row represents a user, and columns represent categorical features relevant for clustering. Here's a step-by-step guide to achieve this with Python:
1. Setup Required Libraries
First, make sure you have the necessary packages installed:
pip install pandas kmodes
2. Load and Aggregate All CSV Files
Assuming your CSV files are stored in a single directory, we'll iterate through each file, extract the user's game genre data, and compile it into a unified DataFrame.
Key Notes:
- Adjust the user ID extraction logic based on your filename format (or use a
user_idcolumn if present in each CSV). - Handle multi-genre entries (e.g., "Action, Adventure") by splitting them into individual genres.
import pandas as pd import os # Define the directory containing your CSV files DATA_DIR = "/path/to/your/steam_user_csvs" user_records = [] for filename in os.listdir(DATA_DIR): if not filename.endswith(".csv"): continue # Skip non-CSV files # Extract user ID (modify this based on your filename structure) # Example: if filename is "user_12345.csv", split to get "12345" user_id = filename.split("_")[-1].replace(".csv", "") # Read the user's CSV user_df = pd.read_csv(os.path.join(DATA_DIR, filename)) # Process genres: split multi-genre strings and collect all genres for the user all_genres = [] for genre_entry in user_df["genre"].dropna(): # Replace "genre" with your actual column name # Split comma-separated genres and strip whitespace genres = [g.strip() for g in genre_entry.split(",")] all_genres.extend(genres) # Create a genre count dictionary for the user genre_counts = pd.Series(all_genres).value_counts().to_dict() # Add user ID and genre counts to our records list user_record = {"user_id": user_id} user_record.update(genre_counts) user_records.append(user_record) # Convert to a DataFrame, filling missing genres with 0 (no purchases in that genre) full_user_data = pd.DataFrame(user_records).fillna(0)
3. Transform Data for k-Modes Clustering
k-Modes works best with categorical features. We'll cover two common approaches to structure your data:
Option A: Binary Genre Presence/Absence
This creates a binary feature for each genre (1 = user purchased at least one game in the genre, 0 = no purchases):
# Convert counts to binary values binary_user_data = full_user_data.set_index("user_id").applymap(lambda x: 1 if x > 0 else 0).reset_index()
Option B: Top N Genres per User
This uses the user's most frequent genres as categorical features (good if you want to focus on primary preferences):
def get_top_n_genres(row, n=3): # Exclude user_id column and filter genres with at least one purchase genre_counts = row.drop("user_id") top_genres = genre_counts[genre_counts > 0].sort_values(ascending=False).head(n).index.tolist() # Fill empty slots with "None" if user has fewer than n genres while len(top_genres) < n: top_genres.append("None") return pd.Series(top_genres, index=[f"top_{i+1}_genre" for i in range(n)]) # Get top 3 genres for each user top_genres_features = full_user_data.apply(get_top_n_genres, axis=1) top_genres_user_data = pd.concat([full_user_data[["user_id"]], top_genres_features], axis=1)
4. Run k-Modes Clustering
Now we can apply the k-Modes algorithm to our structured data. First, choose an optimal number of clusters (k) using the elbow method (plot cost vs k to find where the cost plateaus).
Example with Binary Data:
from kmodes.kmodes import KModes # Prepare data (drop user_id for clustering) X = binary_user_data.set_index("user_id") # Initialize k-Modes (adjust k to your optimal value) km = KModes(n_clusters=5, init="Huang", n_init=5, verbose=1) cluster_labels = km.fit_predict(X) # Add cluster labels back to the original DataFrame binary_user_data["cluster"] = cluster_labels
Example with Top Genres Data:
X = top_genres_user_data.set_index("user_id") km = KModes(n_clusters=5, init="Huang", n_init=5, verbose=1) cluster_labels = km.fit_predict(X) top_genres_user_data["cluster"] = cluster_labels
5. Analyze Cluster Results
- Inspect Centroids: The cluster centroids show the most representative values for each feature in a cluster. For binary data, centroids are 0/1; for top genres, they're the most common genres in each cluster:
print("Cluster Centroids:") print(km.cluster_centroids_) - Profile Clusters: Group users by cluster and analyze their genre preferences to understand distinct user segments (e.g., "Action Gamers", "Indy Simulation Fans").
Pro Tips
- Normalize Genres: Ensure genre names are consistent (e.g., lowercase all entries, fix typos like "Acton" → "Action").
- Handle Large Datasets: If you have tens of thousands of users, consider using
dask.dataframefor parallel processing to avoid memory issues. - Optimize k: Use the elbow method to find the best
k: calculate the cost (dissimilarity) for k values from 2 to 10, then plot the results to find the "elbow" point.
内容的提问来源于stack exchange,提问作者Maz

