如何用Pandas读取无表头类CSV并生成指定结构的透视DataFrame?
Hey there! Let's get that ratings data reshaped into the pivot table you need. Here's a step-by-step breakdown that works perfectly for your use case:
Step 1: Read the Raw Data
Since your file has no headers, we'll explicitly name the columns when loading it into Pandas to avoid confusion:
import pandas as pd # Replace 'your_ratings_file.csv' with your actual file path raw_ratings = pd.read_csv('your_ratings_file.csv', header=None, names=['user_id', 'movie_id', 'rating'])
This creates a "long-format" DataFrame where each row represents one user's rating for one movie.
Step 2: Reshape to Wide Format
To flip this into the structure you want (movie IDs as rows, user IDs as columns, ratings as cell values), use pivot_table—it's more robust than the basic pivot function because it handles duplicate user-movie rating pairs gracefully:
# Reshape the data into your desired structure wide_ratings_df = raw_ratings.pivot_table( index='movie_id', # Sets movie IDs as the Y-axis (rows) columns='user_id', # Sets user IDs as the X-axis (columns) values='rating', # Fills cells with the corresponding rating aggfunc='first' # Uses the first rating if a user rated the same movie multiple times; use 'mean' to average duplicates )
Quick breakdown of the parameters:
index='movie_id': Makes movie IDs the row labels (your Y-axis).columns='user_id': Makes user IDs the column labels (your X-axis).values='rating': Populates each cell with the rating value from the raw data.aggfunc='first': If there are duplicate entries (same user + movie), this picks the first one. If you want to average duplicates instead, swap this with'mean'.
Example Output
Suppose your raw data looks like this:
1,101,5 1,102,3 2,101,4 2,103,2
The resulting wide_ratings_df will look exactly like what you're after:
| movie_id | 1 | 2 |
|---|---|---|
| 101 | 5 | 4 |
| 102 | 3 | NaN |
| 103 | NaN | 2 |
Optional: Clean Up Missing Values
If you want to replace NaN (which means no rating exists for that user-movie pair) with a default value like 0:
wide_ratings_df = wide_ratings_df.fillna(0)
That's all there is to it! This method gives you the precise pivot table structure you need.
内容的提问来源于stack exchange,提问作者Hayden Gill

