Python Pandas DataFrame重塑:合并对应重复列Var3/Var4至Var1/Var2
Reshape Pandas DataFrame by Combining Corresponding Columns
Alright, let's tackle this DataFrame reshaping task. The goal is to take the paired columns (Var1/Var3, Var2/Var4) and turn them into additional rows under the original Var1/Var2 headers, while keeping the Type associated correctly. Here's a straightforward approach that gets the job done:
Step-by-Step Solution
First, let's start with your original DataFrame:
import pandas as pd import numpy as np df = pd.DataFrame({'Type' : ['A', 'A', 'B'], 'Var1' : [1.0, 2.0, 3.0], 'Var2' : [21.0, 22.0, 23.0], 'Var3' : [np.nan, 4.0, 5.0], 'Var4' : [np.nan, 24.0, 25.0] })
We can break this down into three simple steps:
- Extract original valid rows: Grab the
Type,Var1, andVar2columns as the first part of our result. - Align paired columns to original structure: Take
Type,Var3, andVar4, then renameVar3toVar1andVar4toVar2so they match the first part's format. - Combine and clean: Stack the two parts together, drop rows with missing values (these are the invalid entries where both
Var3andVar4wereNaN), and reset the index for a clean output.
Here's the code that puts this all together:
# Get the original rows with Var1/Var2 part1 = df[['Type', 'Var1', 'Var2']] # Convert Var3/Var4 to match Var1/Var2 column names part2 = df[['Type', 'Var3', 'Var4']].rename(columns={'Var3': 'Var1', 'Var4': 'Var2'}) # Combine datasets, remove invalid NaN rows, and reset index result_df = pd.concat([part1, part2]).dropna().reset_index(drop=True)
Output Result
Running this code will give you exactly the target structure you wanted:
| Type | Var1 | Var2 | |
|---|---|---|---|
| 0 | A | 1.0 | 21.0 |
| 1 | A | 2.0 | 22.0 |
| 2 | A | 4.0 | 24.0 |
| 3 | B | 3.0 | 23.0 |
| 4 | B | 5.0 | 25.0 |
Why This Works
pd.concat()stacks the two DataFrames vertically, so we retain all original rows plus the convertedVar3/Var4rows.dropna()automatically removes the row where bothVar3andVar4wereNaN—since that row would have missing values in the convertedVar1/Var2columns, which is exactly the entry we don't want to keep.reset_index(drop=True)cleans up the index to be sequential, matching your target output's structure.
内容的提问来源于stack exchange,提问作者WZhao
相关产品推荐
相关产品推荐

