Pandas中DataFrame堆叠方法咨询:将含列表值的DataFrame转指定格式
Hey there! Let's break down how to convert your DataFrame into the desired format. Here's a step-by-step approach using Pandas:
Step 1: Create the Sample DataFrame (for reference)
First, let's replicate your input DataFrame to work with:
import pandas as pd data = { 'id': ['id01', 'id02', 'id03', 'id04'], 'val': [['a', 'b'], ['b'], [], ['a', 'c']] } df = pd.DataFrame(data)
Step 2: Filter Out Rows with Empty Lists
We need to drop rows where the val column has an empty list. You can do this by checking the length of each list:
# Remove rows with empty lists in 'val' df_filtered = df[df['val'].apply(len) > 0]
Alternatively, using str.len() works too (Pandas handles list columns with string methods in newer versions):
df_filtered = df[df['val'].str.len() > 0]
Step 3: Convert List Column to Multiple Columns
Now, turn each element in the val list into a separate column. We'll create a new DataFrame from the list values, using the id column as the index:
# Convert list column to multiple columns result = pd.DataFrame(df_filtered['val'].tolist(), index=df_filtered['id'])
Step 4: (Optional) Reset Index to Make 'id' a Column
If you prefer id to be a regular column instead of the index, reset it:
result = result.reset_index().rename(columns={'index': 'id'})
Final Output
After running these steps, your result will look like this (with id as index):
0 1 id01 'a' 'b' id02 'b' NaN id04 'a' 'c'
Or if you reset the index:
id 0 1 0 id01 'a' 'b' 1 id02 'b' NaN 2 id04 'a' 'c'
Quick Extra Tips
- To rename numeric columns to something more meaningful (like
val1,val2):result.columns = [f'val{i+1}' for i in result.columns] - To replace
NaNvalues with empty strings (if you don't want missing values shown):result = result.fillna('')
内容的提问来源于stack exchange,提问作者Simon Krannig

