如何检查DataFrame列的空值并向前填充年份值?
Hey there! Let's break down why your code isn't working, and then fix it—plus I'll show you a way more efficient method built right into Pandas.
The Problem with Your Current Code
Your loop isn't detecting missing values because of how Pandas handles NaN (not-a-number) values:
- In Pandas, missing numeric values are represented as
np.nan(from NumPy), not Python'sNone - Even if you tried comparing directly to
np.nan,np.nan == np.nanreturnsFalse(this is a quirk of floating-point NaN values) - So your condition
df.iloc[i,0]==Nonenever triggers, which is why your DataFrame stays unchanged.
The Best Solution: Use Pandas' Built-in ffill() Method
Instead of writing a manual loop (which is slow for large datasets), Pandas has a dedicated method for forward-filling missing values: ffill() (short for "forward fill"). It automatically replaces each NaN with the last non-missing value above it—exactly what you need!
Here's the one-liner:
df['year'] = df['year'].ffill()
That's it! This will fill all NaN values between 2000 and 2001 with 2000, between 2001 and 2002 with 2001, and so on down to 2020.
If You Still Want to Use a Manual Loop
If you prefer sticking with your loop approach for learning purposes, you need to use Pandas' built-in function to check for missing values: pd.isna(). Here's the corrected code:
import pandas as pd size = df["year"].size val = df.iloc[0, 0] # Initialize with the first valid year for i in range(size): if pd.isna(df.iloc[i, 0]): df.iloc[i, 0] = val else: val = df.iloc[i, 0]
A couple of quick notes:
pd.isna()works for bothnp.nanandNonevalues, so it's the safest way to check for missing data in Pandas- Ensure your
yearcolumn is a numeric type (like int or float) so the assignment works correctly.
Quick Verification Check
After running either method, confirm all NaNs are gone with this line:
print(df['year'].isna().sum())
This should return 0 if all missing values were filled properly.
内容的提问来源于stack exchange,提问作者Beto

