You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用.isna()与(df==np.nan).sum().sum()的差异及填充缺失值后统计异常问题

Missing Value Detection: .isna() vs. (df == np.nan)

Let’s break down your questions clearly—this is such a common pitfall when working with missing values in pandas, so it’s great you’re digging into the nuances.

1. Key Differences Between .isna() and (df == np.nan).sum().sum()

The core issue here comes down to how np.nan behaves in comparisons:

  • np.nan can’t be compared to anything (including itself): In NumPy, np.nan == np.nan returns False. When you run df == np.nan, every cell in your DataFrame—even the actual np.nan entries—will evaluate to False. Summing these results will always give you 0, no matter how many missing values are present. This makes this approach completely unreliable for counting missing values.
  • .isna() is purpose-built for pandas: Pandas’ .isna() method was created to correctly identify all types of missing values that pandas recognizes, including:
    • np.nan (used in float/numeric columns)
    • pd.NA (pandas’ universal missing value marker for nullable dtypes like Int64, string, etc.)
    • NaT (missing datetime values)
      It doesn’t rely on equality checks, so it avoids the np.nan comparison trap entirely.

2. Explaining the Post-Fill Anomaly: .isna() Shows 33470, (df == np.nan) Shows 0

Here’s what’s going on:

  • First, remember that (df == np.nan).sum().sum() will always return 0, even if there are np.nan values left. That’s just the inherent behavior of np.nan—it never equals anything, including itself.
  • The 33470 missing values detected by .isna() are not np.nan values. The most likely scenario is:
    • df.fillna(df.mean()) only fills missing values in numeric columns. If your DataFrame has non-numeric columns (like string, datetime, or categorical columns), df.mean() ignores these (since calculating a mean doesn’t make sense for them), so their missing values remain.
    • Those remaining missing values are probably pd.NA (for nullable dtypes) or NaT (for datetimes). .isna() correctly flags these as missing, but comparing them to np.nan with == returns False (or pd.NA, which gets treated as 0 when summing), hence the 0 count from the equality check.
  • A secondary edge case: If you had empty strings in object columns converted to nullable string dtype (where empty strings become pd.NA), .isna() would count those too, while == np.nan would not.

内容的提问来源于stack exchange,提问作者tarun

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:03:19