如何使用Pandas将Dataframe转换为指定的多列分组结构?
I get it, pivot_table isn’t quite the right tool here because we’re not aggregating data—we just need to rearrange existing rows into a side-by-side multi-column format. Here are two straightforward approaches to achieve your desired output:
Approach 1: Split, Rename, and Concatenate
This method breaks down the original DataFrame by category, adjusts each subset’s columns to include the category as a top-level header, then combines them horizontally.
import pandas as pd # Create the original DataFrame df = pd.DataFrame({ 'Column A': [1, 2, 3, 4, 5, 6], 'Column B': [7, 8, 9, 10, 11, 12], 'Category': ['A', 'A', 'B', 'B', 'C', 'C'] }) # Step 1: Split into separate DataFrames for each category category_dfs = [] for category, group in df.groupby('Category'): # Remove the Category column from each group subset = group.drop('Category', axis=1) # Create a multi-level column index with the category as the top level subset.columns = pd.MultiIndex.from_tuples( [(f'Category {category}', col) for col in subset.columns] ) category_dfs.append(subset) # Step 2: Concatenate all subsets side by side result = pd.concat(category_dfs, axis=1) print(result)
Approach 2: Use Pivot with Row IDs
This method adds a row identifier per category, then uses pivot to reshape, followed by adjusting column levels.
import pandas as pd # Create original DataFrame df = pd.DataFrame({ 'Column A': [1, 2, 3, 4, 5, 6], 'Column B': [7, 8, 9, 10, 11, 12], 'Category': ['A', 'A', 'B', 'B', 'C', 'C'] }) # Step 1: Add a row counter for each category (0,1 for each group) df['row_id'] = df.groupby('Category').cumcount() # Step 2: Pivot the DataFrame to get categories as columns pivoted = df.pivot(index='row_id', columns='Category', values=['Column A', 'Column B']) # Step 3: Swap column levels to put category first, then sort columns pivoted = pivoted.swaplevel(0, 1, axis=1).sort_index(axis=1) # Step 4: Rename top-level columns to "Category X" format pivoted.columns = [('Category ' + cat, col) for cat, col in pivoted.columns] # Step 5: Drop the row_id index to match desired output result = pivoted.reset_index(drop=True) print(result)
Both approaches will produce exactly the DataFrame structure you’re looking for, with multi-level columns where the top level is the category name and the second level is the original column names.
内容的提问来源于stack exchange,提问作者ChiYuen Lok

