如何将两个DataFrame转换为共享索引与列的并集,缺失值填充NaN?
Got it, let's solve this problem! You want to make df_A and df_B share the union of their indices and columns, filling any missing values with NaN. The easiest and most straightforward way to do this in pandas is using the built-in align() method—it’s purpose-built for this exact scenario.
Step-by-Step Solution
1. Use the align() Method
The align() method will automatically align both DataFrames to the outer join (union) of their indices and columns, and let you specify a fill value for missing entries. Here's the code:
import pandas as pd import numpy as np # Align df_A and df_B to their union of indices/columns, fill missing with NaN df_A_aligned, df_B_aligned = df_A.align(df_B, join='outer', fill_value=np.nan)
Parameter Breakdown:
join='outer': Tells pandas to use the union of indices and columns (instead of inner join, which uses only the overlap).fill_value=np.nan: Sets missing values in either DataFrame toNaN(you could swap this for another value if needed, butNaNis what you requested).
2. Example to See It in Action
Let’s use sample DataFrames to demonstrate how this works:
# Create sample DataFrames df_A = pd.DataFrame({'X': [1, 2], 'Y': [3, 4]}, index=['a', 'b']) df_B = pd.DataFrame({'Y': [5, 6], 'Z': [7, 8]}, index=['b', 'c']) # Run alignment df_A_aligned, df_B_aligned = df_A.align(df_B, join='outer', fill_value=np.nan)
Result for df_A_aligned:
| X | Y | Z | |
|---|---|---|---|
| a | 1 | 3 | NaN |
| b | 2 | 4 | NaN |
| c | NaN | NaN | NaN |
Result for df_B_aligned:
| X | Y | Z | |
|---|---|---|---|
| a | NaN | NaN | NaN |
| b | NaN | 5 | 7 |
| c | NaN | 6 | 8 |
Alternative: Using pd.concat (If You Prefer)
If you want another approach, you can concatenate the DataFrames and then split them back out, though this is less efficient than align():
# Combine both DataFrames with a key to track which is which combined = pd.concat([df_A, df_B], keys=['A', 'B']) # Extract and reindex each DataFrame to the full set of columns/indices df_A_aligned = combined.loc['A'].reindex(combined.index.get_level_values(1).unique(), axis=0).reindex(combined.columns, axis=1).fillna(np.nan) df_B_aligned = combined.loc['B'].reindex(combined.index.get_level_values(1).unique(), axis=0).reindex(combined.columns, axis=1).fillna(np.nan)
But again, the align() method is the cleanest and most efficient choice here.
内容的提问来源于stack exchange,提问作者ian_chan

