Pandas中合并含NaN的两列:colB为NaN时保留colA内容
Hey there! Let's solve this problem of merging colA and colB into colC while handling those NaN values properly. Here are two solid approaches:
Method 1: Vectorized Approach with np.where (Best for Large Data)
This is the most efficient way, especially if you're working with big datasets, since it uses vectorized operations instead of looping through each row.
import pandas as pd import numpy as np # Your original DataFrame df = pd.DataFrame({'ID': ['ID1', 'ID2', 'ID3'], 'colA': ['A', 'B', 'C'], 'colB': ['D', np.nan, 'E']}) # Add the new colC column df['colC'] = np.where(pd.notna(df['colB']), df['colA'] + '_' + df['colB'], df['colA']) print(df)
Output:
ID colA colB colC 0 ID1 A D A_D 1 ID2 B NaN B 2 ID3 C E C_E
Method 2: Row-wise apply (Great for Small Data/Readability)
If you want something more straightforward and easy to read (perfect for smaller datasets), use apply with a lambda function:
df['colC'] = df.apply(lambda row: f"{row['colA']}_{row['colB']}" if pd.notna(row['colB']) else row['colA'], axis=1)
This will give you exactly the same result as the first method.
How it works:
- When
colBhas a valid value (notNaN), we combinecolAandcolBwith an underscore. - When
colBisNaN, we just keep the value fromcolAwithout adding any extra characters.
内容的提问来源于stack exchange,提问作者Hardik Gupta
相关产品推荐
相关产品推荐

