如何在Pandas中根据另一列的值条件填充列缺失值?
col1 Based on col2 Values in Pandas Let's break down why your original approach didn't work, then walk through the correct ways to achieve your goal.
Why Your Code Failed
Looking at your code:
ix = list(df[df["col2"] == 0].index) df["col2"].loc[ix].fillna(0, inplace = True)
There are three key issues here:
- You're modifying the wrong column: You targeted
col2instead ofcol1—the column you actually want to fill missing values in. fillna()has no effect here: The rows wherecol2 == 0don't have NaN values incol2, so callingfillna(0)does nothing at all.- Chained indexing +
inplace=Truerisks warnings/errors:df["col2"].loc[ix]creates a potential view of the original DataFrame, and modifying it withinplace=Truecan trigger aSettingWithCopyWarning(or even fail to modify the original DataFrame in some edge cases).
Correct Solutions
Here are two straightforward, reliable ways to implement your desired logic:
1. Using np.where (Vectorized Approach)
This is a clean, concise method that applies the condition across the entire column at once:
import numpy as np import pandas as pd d = {'col1': [np.nan, 19, 32, np.nan, 54, 67], 'col2': [0, 1, 0, 1, 1, 1]} df = pd.DataFrame(d) # Apply the fill logic df['col1'] = np.where((df['col1'].isna()) & (df['col2'] == 0), 0, df['col1']) print(df)
Output:
col1 col2 0 0.0 0 1 19.0 1 2 32.0 0 3 NaN 1 4 54.0 1 5 67.0 1
The logic here: For each row, if col1 is NaN and col2 equals 0, replace col1 with 0; otherwise, keep the original col1 value.
2. Using loc with a Mask (Explicit Row/Column Targeting)
This method is more explicit and avoids any ambiguity about which rows/columns you're modifying:
import numpy as np import pandas as pd d = {'col1': [np.nan, 19, 32, np.nan, 54, 67], 'col2': [0, 1, 0, 1, 1, 1]} df = pd.DataFrame(d) # Create a mask for rows where col1 is NaN AND col2 is 0 mask = df['col1'].isna() & (df['col2'] == 0) # Assign 0 to those rows in col1 df.loc[mask, 'col1'] = 0 print(df)
This produces the exact same output as the np.where method. Using loc is generally recommended for modifying DataFrames because it's explicit and avoids unintended side effects.
内容的提问来源于stack exchange,提问作者R-Nie

