如何修改Pandas函数以支持多指标列表的NaN值填充为0
Got it, let's sort this out! The error you're seeing happens because df.isin(Indicators) creates a boolean matrix for your entire DataFrame, but df.loc needs a 1-dimensional boolean array to target the right rows. We just need to narrow the check down to the Indicator column specifically, and also make the function compatible with both single indicator strings and lists of indicators.
Modified Function
Here's the adjusted version that works for both single values and lists:
import pandas as pd import numpy as np def zerofillnaindicator(df, indicators): # Handle both single indicator strings and lists if not isinstance(indicators, list): indicators = [indicators] # Create a mask targeting rows where Indicator is in our list mask = df['Indicator'].isin(indicators) # Fill NaNs in Value column for those rows with 0 df.loc[mask, 'Value'] = df.loc[mask, 'Value'].fillna(0) return df
How to Test It
Using your sample DataFrame:
# Sample DataFrame df = pd.DataFrame({ 'ISO3': ['Australia', 'Austria', 'Belgium', 'Canada', 'Australia', 'Austria', 'Belgium', 'Canada'], 'Year': [1991]*8, 'Indicator' : ['Disaster Fatalities']*4 + ['Oil Reserves']*4, 'Value' : [np.nan, 5, np.nan, 18, np.nan, np.nan, np.nan, np.nan] }) # Test with a list of indicators df2 = zerofillnaindicator(df=df, indicators=['Disaster Fatalities', 'Oil Reserves']) print(df2)
Expected Output
ISO3 Year Indicator Value 0 Australia 1991 Disaster Fatalities 0.0 1 Austria 1991 Disaster Fatalities 5.0 2 Belgium 1991 Disaster Fatalities 0.0 3 Canada 1991 Disaster Fatalities 18.0 4 Australia 1991 Oil Reserves 0.0 5 Austria 1991 Oil Reserves 0.0 6 Belgium 1991 Oil Reserves 0.0 7 Canada 1991 Oil Reserves 0.0
Extra Tip: Avoid Modifying the Original DataFrame
If you don't want to alter your original DataFrame (which is often a good practice), add a copy step at the start of the function:
def zerofillnaindicator(df, indicators): df = df.copy() # Create a copy to avoid modifying the original if not isinstance(indicators, list): indicators = [indicators] mask = df['Indicator'].isin(indicators) df.loc[mask, 'Value'] = df.loc[mask, 'Value'].fillna(0) return df
This way, your original df stays untouched, and you get a new DataFrame with the filled values.
内容的提问来源于stack exchange,提问作者Laurens

