如何基于userId和movieId将data.frame转换为评分矩阵?
Hey there! It looks like you're trying to convert your long-format DataFrame into a wide-format matrix where rows represent user IDs, columns represent movie IDs, and cells hold the corresponding ratings. The groupby approach you started with isn't the right tool for this reshaping task—let me show you the correct methods to get the result you want.
First, Let's Confirm Your Original Data
Here's how your sample DataFrame looks (I'll recreate it for clarity):
import pandas as pd Ratings = pd.DataFrame({ 'userId': [12882, 12882, 12882, 12882, 12882], 'movieId': [1, 32, 47, 50, 110], 'rating': [4.0, 3.5, 5.0, 5.0, 4.5] })
Method 1: Use pivot() (Best for Unique User-Movie Pairs)
Since your data has no duplicate ratings for the same user-movie combination, pivot() is the simplest and most efficient way to reshape your data. It directly maps rows/columns to the values you want:
# Reshape to wide-format matrix rating_matrix = Ratings.pivot( index='userId', # Rows will be user IDs columns='movieId',# Columns will be movie IDs values='rating' # Cell values are the ratings ) # Optional: If you want to clean up any columns with all missing values (not needed here) # rating_matrix = rating_matrix.dropna(axis=1, how='all') print(rating_matrix)
Output:
movieId 1 32 47 50 110 userId 12882 4.0 3.5 5.0 5.0 4.5
Method 2: Use pivot_table() (For Handling Duplicate Ratings)
If you ever encounter duplicate ratings for the same user-movie pair (e.g., a user rated the same movie multiple times), pivot_table() lets you aggregate those duplicates (like taking the average, sum, etc.):
# Use mean as the aggregation function (default if not specified) rating_matrix = Ratings.pivot_table( index='userId', columns='movieId', values='rating', aggfunc='mean' # You can also use 'sum', 'max', etc. ) print(rating_matrix)
Why groupby Didn't Work for You
Just to clarify: groupby is designed for aggregating data within groups (e.g., calculating average ratings per user, or total ratings per movie). It doesn't reshape the structure of your DataFrame, which is why you got stuck after grouping—you were using the wrong tool for this specific reshaping task.
内容的提问来源于stack exchange,提问作者Nicolas H

