You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何读取多Steam用户CSV文件并开展k-modes聚类分析?

Steam User Clustering with k-Modes: Combining Multiple CSV Files

Great question! When working with thousands of user-specific CSV files, the key is to first aggregate all the data into a single structured dataset where each row represents a user, and columns represent categorical features relevant for clustering. Here's a step-by-step guide to achieve this with Python:

1. Setup Required Libraries

First, make sure you have the necessary packages installed:

pip install pandas kmodes

2. Load and Aggregate All CSV Files

Assuming your CSV files are stored in a single directory, we'll iterate through each file, extract the user's game genre data, and compile it into a unified DataFrame.

Key Notes:

  • Adjust the user ID extraction logic based on your filename format (or use a user_id column if present in each CSV).
  • Handle multi-genre entries (e.g., "Action, Adventure") by splitting them into individual genres.
import pandas as pd
import os

# Define the directory containing your CSV files
DATA_DIR = "/path/to/your/steam_user_csvs"

user_records = []

for filename in os.listdir(DATA_DIR):
    if not filename.endswith(".csv"):
        continue  # Skip non-CSV files
    
    # Extract user ID (modify this based on your filename structure)
    # Example: if filename is "user_12345.csv", split to get "12345"
    user_id = filename.split("_")[-1].replace(".csv", "")
    
    # Read the user's CSV
    user_df = pd.read_csv(os.path.join(DATA_DIR, filename))
    
    # Process genres: split multi-genre strings and collect all genres for the user
    all_genres = []
    for genre_entry in user_df["genre"].dropna():  # Replace "genre" with your actual column name
        # Split comma-separated genres and strip whitespace
        genres = [g.strip() for g in genre_entry.split(",")]
        all_genres.extend(genres)
    
    # Create a genre count dictionary for the user
    genre_counts = pd.Series(all_genres).value_counts().to_dict()
    
    # Add user ID and genre counts to our records list
    user_record = {"user_id": user_id}
    user_record.update(genre_counts)
    user_records.append(user_record)

# Convert to a DataFrame, filling missing genres with 0 (no purchases in that genre)
full_user_data = pd.DataFrame(user_records).fillna(0)

3. Transform Data for k-Modes Clustering

k-Modes works best with categorical features. We'll cover two common approaches to structure your data:

Option A: Binary Genre Presence/Absence

This creates a binary feature for each genre (1 = user purchased at least one game in the genre, 0 = no purchases):

# Convert counts to binary values
binary_user_data = full_user_data.set_index("user_id").applymap(lambda x: 1 if x > 0 else 0).reset_index()

Option B: Top N Genres per User

This uses the user's most frequent genres as categorical features (good if you want to focus on primary preferences):

def get_top_n_genres(row, n=3):
    # Exclude user_id column and filter genres with at least one purchase
    genre_counts = row.drop("user_id")
    top_genres = genre_counts[genre_counts > 0].sort_values(ascending=False).head(n).index.tolist()
    # Fill empty slots with "None" if user has fewer than n genres
    while len(top_genres) < n:
        top_genres.append("None")
    return pd.Series(top_genres, index=[f"top_{i+1}_genre" for i in range(n)])

# Get top 3 genres for each user
top_genres_features = full_user_data.apply(get_top_n_genres, axis=1)
top_genres_user_data = pd.concat([full_user_data[["user_id"]], top_genres_features], axis=1)

4. Run k-Modes Clustering

Now we can apply the k-Modes algorithm to our structured data. First, choose an optimal number of clusters (k) using the elbow method (plot cost vs k to find where the cost plateaus).

Example with Binary Data:

from kmodes.kmodes import KModes

# Prepare data (drop user_id for clustering)
X = binary_user_data.set_index("user_id")

# Initialize k-Modes (adjust k to your optimal value)
km = KModes(n_clusters=5, init="Huang", n_init=5, verbose=1)
cluster_labels = km.fit_predict(X)

# Add cluster labels back to the original DataFrame
binary_user_data["cluster"] = cluster_labels

Example with Top Genres Data:

X = top_genres_user_data.set_index("user_id")

km = KModes(n_clusters=5, init="Huang", n_init=5, verbose=1)
cluster_labels = km.fit_predict(X)

top_genres_user_data["cluster"] = cluster_labels

5. Analyze Cluster Results

  • Inspect Centroids: The cluster centroids show the most representative values for each feature in a cluster. For binary data, centroids are 0/1; for top genres, they're the most common genres in each cluster:
    print("Cluster Centroids:")
    print(km.cluster_centroids_)
    
  • Profile Clusters: Group users by cluster and analyze their genre preferences to understand distinct user segments (e.g., "Action Gamers", "Indy Simulation Fans").

Pro Tips

  • Normalize Genres: Ensure genre names are consistent (e.g., lowercase all entries, fix typos like "Acton" → "Action").
  • Handle Large Datasets: If you have tens of thousands of users, consider using dask.dataframe for parallel processing to avoid memory issues.
  • Optimize k: Use the elbow method to find the best k: calculate the cost (dissimilarity) for k values from 2 to 10, then plot the results to find the "elbow" point.

内容的提问来源于stack exchange,提问作者Maz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:17:08