将Pandas DataFrame的genre列拆分为多个布尔列的实现方法
Great question! Working with pipe-separated categorical columns like your genre field is a common data wrangling task, and pandas has some handy tools to make this easy. Let's tackle both of your problems step by step.
1. Get All Unique Genre Names
First, we need to extract every distinct genre from the pipe-separated strings. Here's how to do it efficiently with pandas:
First, load your CSV into a DataFrame:
import pandas as pd # Replace 'movies.csv' with your actual file path df = pd.read_csv('movies.csv')
Next, split each genre entry into a list, "explode" those lists into individual rows, then grab the unique values:
# Split genre strings into lists, then explode to create one genre per row all_genres = df['genre'].str.split('|').explode() # Get unique genres, sorted for readability unique_genres = sorted(all_genres.unique()) # Print the result to check print("Unique genres:", unique_genres)
This will give you a clean, sorted list of every genre present in your dataset. If there are missing values in the genre column, add df['genre'] = df['genre'].fillna('') first to avoid errors.
2. Create Boolean Columns for Each Genre
Once you have the unique genres, generating boolean (True/False) columns is a breeze. There are two main approaches depending on your needs:
Approach 1: Use str.get_dummies() (Quickest)
Pandas' built-in str.get_dummies() method automatically creates columns for each unique genre, with 1 indicating the genre is present and 0 otherwise. You can convert these to boolean values if you prefer:
# Create dummy columns from pipe-separated genres genre_dummies = df['genre'].str.get_dummies(sep='|') # Optional: Convert 0/1 to True/False genre_dummies = genre_dummies.astype(bool) # Merge the new columns with your original DataFrame df_with_booleans = pd.concat([df, genre_dummies], axis=1)
This will add columns named after each genre (e.g., Adventure, Science Fiction) with boolean values.
Approach 2: Manual Column Creation (Customizable)
If you want more control over column names (like using is_adventure instead of Adventure), loop through the unique genres we found earlier:
# First, make sure we have our unique_genres list from step 1 for genre in unique_genres: # Create a column like 'is_adventure' with True if the genre is present df_with_booleans[f'is_{genre.lower().replace(" ", "_")}'] = df['genre'].str.contains(genre, case=False)
This creates lowercase, snake_case column names and uses case-insensitive matching (so "adventure" and "Adventure" are treated the same). Adjust case=True if you need exact case matching.
Both methods work efficiently even with 10,000 rows—pandas is optimized for these kinds of operations.
内容的提问来源于stack exchange,提问作者EdwardC

