如何将DataFrame中列表型类别列转为编码变量?代码报错求助
Hey Danny, let's break down why your current code is failing and walk through two clean, efficient ways to get the encoded variables you need.
What's Wrong with Your Current Code?
Your loop approach has two key issues:
pd.concat()expects a sequence of DataFrames/Series (like a list of them), but you're passing individual scalar values (data.iloc[i][0]is a business ID string/number,data.iloc[i][1][j]is a category string). That's why you're getting an error.- Even if it worked, this method would be extremely slow for larger datasets and would not structure the data into proper encoded columns.
Solution 1: Use Pandas Built-ins (explode + get_dummies + groupby)
This is a pandas-native way to handle the task, easy to read and maintain:
First, let's recreate your sample DataFrame for reference:
import pandas as pd data = pd.DataFrame({ 'business_id': ['1K4qrnfyzKzGgJPBEcJaNQ', 'dTWfATVrBfKj7Vdn0qWVWg'], 'categories': [['Tiki Bars', 'Nightlife', 'Mexican', 'Restaurants', 'Bars'], ['Restaurants', 'Chinese', 'Food Court']] })
Now run this code to generate the encoded variables:
# Step 1: Explode the list column into individual rows exploded_data = data.explode('categories') # Step 2: Create one-hot encoded columns for each category dummy_data = pd.get_dummies(exploded_data, columns=['categories']) # Step 3: Group by business_id and sum to collapse back to one row per business categorical_data = dummy_data.groupby('business_id').sum().reset_index()
The result will be a DataFrame where each category is a column, with 1 if the business belongs to that category, 0 otherwise.
Solution 2: Use MultiLabelBinarizer from Scikit-learn
If you're working with machine learning workflows, this method is more direct for multi-label encoding:
from sklearn.preprocessing import MultiLabelBinarizer import pandas as pd # Initialize the binarizer mlb = MultiLabelBinarizer() # Fit-transform the categories column and convert to DataFrame encoded_categories = pd.DataFrame(mlb.fit_transform(data['categories']), columns=mlb.classes_, index=data['business_id']) # Reset index to get business_id as a column (optional, depending on your needs) categorical_data = encoded_categories.reset_index()
This will produce the same encoded output as Solution 1, and it's especially useful if you need to reuse the binarizer for other data (e.g., test sets in ML).
Final Output Example
Both methods will give you this structure:
| business_id | Bars | Chinese | Food Court | Mexican | Nightlife | Restaurants | Tiki Bars |
|---|---|---|---|---|---|---|---|
| 1K4qrnfyzKzGgJPBEcJaNQ | 1 | 0 | 0 | 1 | 1 | 1 | 1 |
| dTWfATVrBfKj7Vdn0qWVWg | 0 | 1 | 1 | 0 | 0 | 1 | 0 |
内容的提问来源于stack exchange,提问作者Danny

