You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将DataFrame中列表型类别列转为编码变量?代码报错求助

Fix: Convert List-Type Category Column to Encoded Variables in Pandas

Hey Danny, let's break down why your current code is failing and walk through two clean, efficient ways to get the encoded variables you need.

What's Wrong with Your Current Code?

Your loop approach has two key issues:

  • pd.concat() expects a sequence of DataFrames/Series (like a list of them), but you're passing individual scalar values (data.iloc[i][0] is a business ID string/number, data.iloc[i][1][j] is a category string). That's why you're getting an error.
  • Even if it worked, this method would be extremely slow for larger datasets and would not structure the data into proper encoded columns.

Solution 1: Use Pandas Built-ins (explode + get_dummies + groupby)

This is a pandas-native way to handle the task, easy to read and maintain:

First, let's recreate your sample DataFrame for reference:

import pandas as pd

data = pd.DataFrame({
    'business_id': ['1K4qrnfyzKzGgJPBEcJaNQ', 'dTWfATVrBfKj7Vdn0qWVWg'],
    'categories': [['Tiki Bars', 'Nightlife', 'Mexican', 'Restaurants', 'Bars'],
                   ['Restaurants', 'Chinese', 'Food Court']]
})

Now run this code to generate the encoded variables:

# Step 1: Explode the list column into individual rows
exploded_data = data.explode('categories')

# Step 2: Create one-hot encoded columns for each category
dummy_data = pd.get_dummies(exploded_data, columns=['categories'])

# Step 3: Group by business_id and sum to collapse back to one row per business
categorical_data = dummy_data.groupby('business_id').sum().reset_index()

The result will be a DataFrame where each category is a column, with 1 if the business belongs to that category, 0 otherwise.


Solution 2: Use MultiLabelBinarizer from Scikit-learn

If you're working with machine learning workflows, this method is more direct for multi-label encoding:

from sklearn.preprocessing import MultiLabelBinarizer
import pandas as pd

# Initialize the binarizer
mlb = MultiLabelBinarizer()

# Fit-transform the categories column and convert to DataFrame
encoded_categories = pd.DataFrame(mlb.fit_transform(data['categories']), 
                                  columns=mlb.classes_,
                                  index=data['business_id'])

# Reset index to get business_id as a column (optional, depending on your needs)
categorical_data = encoded_categories.reset_index()

This will produce the same encoded output as Solution 1, and it's especially useful if you need to reuse the binarizer for other data (e.g., test sets in ML).


Final Output Example

Both methods will give you this structure:

business_idBarsChineseFood CourtMexicanNightlifeRestaurantsTiki Bars
1K4qrnfyzKzGgJPBEcJaNQ1001111
dTWfATVrBfKj7Vdn0qWVWg0110010

内容的提问来源于stack exchange,提问作者Danny

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 06:31:27