使用.isna()与(df==np.nan).sum().sum()的差异及填充缺失值后统计异常问题
Missing Value Detection:
.isna() vs. (df == np.nan) Let’s break down your questions clearly—this is such a common pitfall when working with missing values in pandas, so it’s great you’re digging into the nuances.
1. Key Differences Between .isna() and (df == np.nan).sum().sum()
The core issue here comes down to how np.nan behaves in comparisons:
np.nancan’t be compared to anything (including itself): In NumPy,np.nan == np.nanreturnsFalse. When you rundf == np.nan, every cell in your DataFrame—even the actualnp.nanentries—will evaluate toFalse. Summing these results will always give you 0, no matter how many missing values are present. This makes this approach completely unreliable for counting missing values..isna()is purpose-built for pandas: Pandas’.isna()method was created to correctly identify all types of missing values that pandas recognizes, including:np.nan(used in float/numeric columns)pd.NA(pandas’ universal missing value marker for nullable dtypes likeInt64,string, etc.)NaT(missing datetime values)
It doesn’t rely on equality checks, so it avoids thenp.nancomparison trap entirely.
2. Explaining the Post-Fill Anomaly: .isna() Shows 33470, (df == np.nan) Shows 0
Here’s what’s going on:
- First, remember that
(df == np.nan).sum().sum()will always return 0, even if there arenp.nanvalues left. That’s just the inherent behavior ofnp.nan—it never equals anything, including itself. - The 33470 missing values detected by
.isna()are notnp.nanvalues. The most likely scenario is:df.fillna(df.mean())only fills missing values in numeric columns. If your DataFrame has non-numeric columns (likestring,datetime, or categorical columns),df.mean()ignores these (since calculating a mean doesn’t make sense for them), so their missing values remain.- Those remaining missing values are probably
pd.NA(for nullable dtypes) orNaT(for datetimes)..isna()correctly flags these as missing, but comparing them tonp.nanwith==returnsFalse(orpd.NA, which gets treated as 0 when summing), hence the 0 count from the equality check.
- A secondary edge case: If you had empty strings in
objectcolumns converted to nullablestringdtype (where empty strings becomepd.NA),.isna()would count those too, while== np.nanwould not.
内容的提问来源于stack exchange,提问作者tarun
相关产品推荐
相关产品推荐

