You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas数据输出部分NaN问题及X变量值恢复求助

Fixing NaNs in df1's '1222', '1224', '1223' Columns & Understanding Why They Appear

Hey there, let's tackle this problem step by step—first we'll unpack why those NaNs showed up in df1 even though df.head() looked perfect, then we'll get those missing values back.

Why NaNs Pop Up in Pandas DataFrames (Even When Head Looks Fine)

It's super common to see df.head() show clean data but run into NaNs later when creating a derived DataFrame like df1. Here are the most likely culprits:

  • Your df1 transformation introduced them: If you created df1 using operations like merge, pivot, groupby, or filtering, these can easily generate NaNs. For example:
    • A merge with mismatched keys (e.g., using a left join where the right DataFrame lacks some entries)
    • A pivot where certain row/column combinations don't exist in the original data
    • Filtering rows that accidentally excluded valid entries, leaving gaps
  • df.head() only shows the first 5 rows: The missing values might already exist in the original df—they're just not in the top 5 rows. df.head() doesn't give you the full picture of your dataset.
  • Index misalignment: If df1 has a different index than df, pandas will fill missing index positions with NaNs when you copy over data.
  • Hidden invalid values: Sometimes raw data has blank strings, 'NA', or 'null' that pandas didn't recognize as NaNs during initial loading. These might look fine in df.head() but turn into NaNs when you perform calculations or type conversions later.

How to Recover Values for '1222', '1224', '1223'

First, let's diagnose the root cause, then apply a fix:

Step 1: Verify if the NaNs existed in the original df

Run these commands to check if the missing values were already present before creating df1:

# Check total NaNs in each column across the entire df
print(df[['1222', '1224', '1223']].isna().sum())

# Inspect rows where these columns have NaNs (if any)
print(df[df[['1222', '1224', '1223']].isna().any(axis=1)])

If you see NaNs here, they were just hidden from df.head(). If not, the issue is in how you created df1.

Step 2: Fix based on the root cause

Case 1: NaNs came from df1 transformation (e.g., merge/pivot/filter)

  • Merge issues: Double-check your merge() parameters. If you used how='left' or how='outer', switch to how='inner' if you only want matching rows, or ensure your join keys are correct (no typos, mismatched data types).
  • Pivot issues: Add fillna() directly to your pivot call to fill missing combinations with a logical value (e.g., 0 for counts, median for numerical data):
    df1 = df.pivot(index='some_col', columns='another_col', values=['1222', '1224', '1223']).fillna(df[['1222', '1224', '1223']].median())
    
  • Filtering issues: Review your filtering logic. If you used something like df[df['some_col'] > 100], make sure you didn't accidentally exclude rows with valid values in '1222'/'1224'/'1223'.

Case 2: NaNs existed in the original df (but not in head)

Fill the missing values using a method that makes sense for your data:

  • Numerical columns: Use median (robust to outliers) or mean:
    df1['1222'] = df1['1222'].fillna(df['1222'].median())
    df1['1224'] = df1['1224'].fillna(df['1224'].mean())
    
  • Forward/backward fill: If your data is time-series or ordered, use ffill (fill with previous value) or bfill (fill with next value):
    df1[['1222', '1224', '1223']] = df1[['1222', '1224', '1223']].fillna(method='ffill')
    
  • Custom logic: If you can derive missing values from other columns, use apply() or map():
    def fill_1222(row):
        if pd.isna(row['1222']):
            return row['related_col'] * 0.8  # Example logic
        return row['1222']
    
    df1['1222'] = df1.apply(fill_1222, axis=1)
    

Case 3: Index misalignment

If df1 has a different index than df, re-align it to match:

df1 = df1.reindex(df.index)
# Or, if you're copying columns from df to df1:
df1[['1222', '1224', '1223']] = df[['1222', '1224', '1223']].values

Step 3: Double-check your fix

Run these to confirm the NaNs are gone:

print(df1[['1222', '1224', '1223']].isna().sum())
print(df1.head())

内容的提问来源于stack exchange,提问作者Chen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:25:59