如何用Pandas合并CSV两列有效值生成新列?排查fillna失效问题
Hey Marco, sorry to hear you’ve been stuck on this for two days—let’s get this sorted out quickly!
The core problem here is almost certainly that the "NaN" values in your Number4 column aren’t actual missing values (like pd.NA, np.nan, or None). Instead, they’re probably stored as string literals ('NaN'), which pandas doesn’t recognize as missing when you use fillna(). That’s why your code only returns the values from Number4—it never finds any "real" NaNs to replace with Number5 values.
Here’s how to fix it:
- First, convert string-based "NaN" values to actual missing values:
Usereplace()to swap any string 'NaN' (or lowercase 'nan') entries with pandas’ native missing value marker:df['Number4'] = df['Number4'].replace({'NaN': pd.NA, 'nan': pd.NA}) - Then run your original fillna() code:
Now pandas will properly identify missing values and replace them withNumber5values:df['Number6'] = df['Number4'].fillna(df['Number5'])
How to verify the issue first:
To confirm this is the problem, run these checks on your DataFrame:
- Check unique values in
Number4to see if 'NaN' is a string:print(df['Number4'].unique()) - Count actual missing values (if this returns 0, you have string "NaN"s):
print(df['Number4'].isna().sum())
This should give you the Number6 column you’re expecting—pulling valid values from Number4 first, then falling back to Number5 whenever Number4 is missing.
内容的提问来源于stack exchange,提问作者Marco

