使用掩码填充混合类型DataFrame子集缺失值,跨表按NaN位置补全
Let's walk through how to solve this problem step by step. We need to fill missing values (NaNs) in columns ["A", "B", "C"] of both DataFrames a and b—specifically, any position that has a NaN in either DataFrame should be replaced with the corresponding value from the other DataFrame. We'll use boolean masks to target exactly the positions that need filling.
Step 1: Define the Original DataFrames
First, let's recap the original data so we can see what we're working with:
import pandas as pd import numpy as np a = pd.DataFrame( [ ['X', 1, np.nan, 3], ['X', 4, 5, 6], ['Y', 7, 8, 9] ], columns = ["Group", "A", "B", "C"] ) b = pd.DataFrame( [ ['X', 1, 2, 3], ['X', 4, 5, np.nan], ['X', 7, 8, 9] ], columns = ["Group", "A", "B", "C"] )
If you print these, you'll spot:
ahas a NaN in row 0, columnBbhas a NaN in row 1, columnC
Step 2: Target Columns & Create Masks
We'll focus only on columns ["A", "B", "C"]. First, create boolean masks that mark where NaNs exist in each DataFrame's target columns:
# Define the columns we want to process target_cols = ["A", "B", "C"] # Mask for NaNs in DataFrame a: True = NaN exists here mask_a = a[target_cols].isna() # Mask for NaNs in DataFrame b: True = NaN exists here mask_b = b[target_cols].isna()
Each mask is a boolean DataFrame that precisely flags the positions needing attention.
Step 3: Fill NaNs Using Masks
We'll use pandas' where() method to replace NaNs. This method keeps values where the mask is False (valid values) and swaps in values from the other DataFrame where the mask is True (missing values).
Fill NaNs in a
# Replace NaNs in a with matching values from b a[target_cols] = a[target_cols].where(~mask_a, b[target_cols])
Fill NaNs in b
# Replace NaNs in b with matching values from a b[target_cols] = b[target_cols].where(~mask_b, a[target_cols])
Step 4: Verify the Result
Print the updated DataFrames to confirm the NaNs are filled:
print("Updated DataFrame a:") print(a) print("\nUpdated DataFrame b:") print(b)
Output:
Updated DataFrame a: Group A B C 0 X 1 2.0 3 1 X 4 5.0 6 2 Y 7 8.0 9 Updated DataFrame b: Group A B C 0 X 1 2 3.0 1 X 4 5 6.0 2 X 7 8 9.0
Key Notes
- This works because your two DataFrames have matching row/column indices. If they didn't, you'd want to align them first (e.g., using
mergeorreindex). - If a position has NaNs in both DataFrames, it will stay NaN—there's no valid value to replace it with!
- We only modify the
["A", "B", "C"]columns, leaving theGroupcolumn untouched as requested.
内容的提问来源于stack exchange,提问作者pault

