You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas多索引列Melt操作:如何将多层列DataFrame转换为指定扁平化表格格式

Solution to Reshape Multi-Index Column DataFrame to Long Format

Let's break down how to transform your four-level multi-index column DataFrame into the desired flat, long-format structure. The key is to restructure the column hierarchies and pivot/unpivot the data correctly.

Step 1: Prepare the Original DataFrame

First, let's confirm your starting DataFrame is set up correctly (I'll reuse your code with a small clarity fix):

import pandas as pd
import numpy as np

# Original code to create the multi-index DataFrame
arrays = [['Phase 1','Phase 1','Phase 1','Phase 1','Phase 1','Phase 1','Phase 1','Phase 1'], 
          ['Function A','Function A','Function A','Function A','Function A','Function A','Function A','Function A'], 
          ['Achieved on','Achieved on','Achieved on','Achieved on','Achieved on','Achieved on','Planned for', 'Due Date'], 
          ['Deliverable 1','Status','True?','Deliverable 2','Status.1','True?.1','NaN','NaN']]
tuples = list(zip(*arrays))
index = pd.MultiIndex.from_tuples(tuples, names=['first','second','third','Project'])

s1 = pd.Series(['10/10/2020','Updated','Yes','11/10/2020','Pending','','',''], index=index)
reset_df = s1.reset_index()
df = pd.DataFrame(reset_df, index=['Project A', 'Project B'], columns=index)
df2 = pd.Series(['10/10/2020','Updated','Yes','11/10/2020','Pending','','',''], index=index)
df3 = pd.Series(['06/06/2021','Issued','','','','','',''],index=index)
df = df.append([df2,df3], ignore_index=True)
df = df.drop([0,1])

# Add a Project identifier column from the index
df = df.rename_axis('Project').reset_index()

Step 2: Filter and Clean Columns

We can ignore the Planned for and Due Date columns since they aren't needed in your target output. Let's filter those out first:

# Keep only columns where the third level is 'Achieved on'
df_filtered = df.loc[:, df.columns.get_level_values('third') == 'Achieved on']

Step 3: Restructure Column Hierarchies

Next, map the fourth-level column names (like Status.1, True?.1) to their corresponding Requisite (Deliverable 1/2) and attributes (Achieved on, Status, True?):

# Define mapping for column names to (Requisite, Attribute)
col_mapping = {
    'Deliverable 1': ('Deliverable 1', 'Achieved on'),
    'Status': ('Deliverable 1', 'Status'),
    'True?': ('Deliverable 1', 'True?'),
    'Deliverable 2': ('Deliverable 2', 'Achieved on'),
    'Status.1': ('Deliverable 2', 'Status'),
    'True?.1': ('Deliverable 2', 'True?')
}

# Rebuild the column index with Requisite and Attribute levels
new_columns = pd.MultiIndex.from_tuples(
    [(first, second) + col_mapping[proj] 
     for first, second, _, proj in df_filtered.columns],
    names=['Phase', 'Function', 'Requisite', 'Attribute']
)

df_filtered.columns = new_columns

Step 4: Pivot to Long Format

Stack and unstack the data to get the flat structure:

# Stack Requisite/Attribute levels, then unstack attributes to columns
df_flat = df_filtered.stack(level=['Requisite', 'Attribute']).unstack('Attribute').reset_index()

# Clean up redundant columns and fill empty values
df_flat = df_flat.drop(columns=['level_2'])  # Remove unused 'third' level
df_flat = df_flat.fillna('')

# Reorder columns to match your target output
df_flat = df_flat[['Project', 'Phase', 'Function', 'Requisite', 'Achieved on', 'Status', 'True?']]

Final Result

Running this code produces exactly the DataFrame you wanted:

Project     Phase    Function    Requisite Achieved on   Status True?
0       2  Phase 1  Function A  Deliverable 1  10/10/2020  Updated   Yes
1       2  Phase 1  Function A  Deliverable 2  11/10/2020  Pending      
2       3  Phase 1  Function A  Deliverable 1  06/06/2021   Issued      
3       3  Phase 1  Function A  Deliverable 2                           

Alternative Flexible Approach with melt and pivot

If you need adaptability for varying column structures, use melt to unpivot all columns, then clean and pivot back:

# Start with df that has the Project column added
df_melted = df.melt(id_vars='Project', var_name='col', value_name='value')

# Split column names into components
df_melted[['Phase', 'Function', 'Third', 'Detail']] = df_melted['col'].str.split('_', n=3, expand=True)

# Filter out irrelevant rows
df_melted = df_melted[df_melted['Third'] == 'Achieved on']

# Map details to their corresponding Deliverable (forward fill to associate attributes)
df_melted['Requisite'] = df_melted['Detail'].where(df_melted['Detail'].str.startswith('Deliverable'))
df_melted['Requisite'] = df_melted.groupby('Project')['Requisite'].ffill()

# Clean attribute names (remove .1 suffixes)
df_melted['Attribute'] = df_melted.apply(
    lambda row: 'Achieved on' if row['Detail'].startswith('Deliverable') else row['Detail'].replace('.1', ''),
    axis=1
)

# Pivot to get attributes as columns
df_pivoted = df_melted.pivot(
    index=['Project', 'Phase', 'Function', 'Requisite'],
    columns='Attribute',
    values='value'
).reset_index().fillna('')

# Reorder columns to match target
df_pivoted = df_pivoted[['Project', 'Phase', 'Function', 'Requisite', 'Achieved on', 'Status', 'True?']]

This gives the same result and is easier to adjust if your multi-index structure changes later.

内容的提问来源于stack exchange,提问作者mgrijo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 23:22:29