Python Pandas按col2拆分DataFrame并转置指定列的需求
Hey there! Let's walk through how to reshape your DataFrame exactly as you need it, and store each grouped result in a dictionary with keys matching your col2 values.
Step 1: Setup and Original Data
First, let's make sure we have the data loaded correctly (I've cleaned up the data definition a bit for readability):
import pandas as pd data = { 'col1': [1, 101, 201, 301, 2, 102, 202, 302, 3, 103, 203, 303], 'col2': [1, 1, 1, 1, 2, 2, 2, 2, 3, 3, 3, 3], 'col3': ["2015-01-15"] * 12, 'col4': ["2015-01-15", "2015-01-16", "2015-01-17", "2015-01-18"] * 3, 'col5': [0, 1, 2, 3] * 3, 'col6': [273.2, 275.9, 343, 235] * 3, 'col7': [2.8, 3.2, 7.9, 7.2] * 3 } df = pd.DataFrame(data)
Step 2: Reshape and Group the Data
We'll use groupby to split the DataFrame by col2, then reshape each group into the wide format you want. We'll store each result in a dictionary where keys are the col2 values (as strings, like '1', '2'):
# Initialize a dictionary to hold our final DataFrames result_dfs = {} # Iterate over each group in the col2 grouping for col2_val, group in df.groupby('col2'): # Create a unique key for each row by combining col1 and col5 group['row_key'] = group.apply(lambda row: f"{row['col1']}_{row['col5']}", axis=1) # Pivot the group to turn col6/col7 into wide columns using our row_key pivoted_group = group.pivot( index=['col2', 'col3'], columns='row_key', values=['col6', 'col7'] ) # Flatten the multi-level column names to match your desired format (e.g., 1_0_col6) pivoted_group.columns = [f"{key}_{col_name}" for col_name, key in pivoted_group.columns] # Reset index to make col2 and col3 regular columns instead of index levels final_group_df = pivoted_group.reset_index() # Store the result in our dictionary with col2 value as the key result_dfs[str(col2_val)] = final_group_df
Step 3: Check the Output
Now you can access each DataFrame using the col2 value as the key in the result_dfs dictionary:
# Print df['1'] print("df['1']:") print(result_dfs['1']) # Print df['2'] print("\ndf['2']:") print(result_dfs['2']) # Print df['3'] print("\ndf['3']:") print(result_dfs['3'])
Sample Output for df['1']:
col2 col3 1_0_col6 101_1_col6 201_2_col6 301_3_col6 1_0_col7 101_1_col7 201_2_col7 301_3_col7 0 1 2015-01-15 273.2 275.9 343.0 235.0 2.8 3.2 7.9 7.2
This matches exactly the format you requested! Each group retains col2 and col3, with col6 and col7 values spread into columns named using the col1_col5 prefix.
内容的提问来源于stack exchange,提问作者rnvs1116

