You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas填充total_income空值遇FutureWarning及无限循环问题求助

问题:按分组中位数填充空值时出现FutureWarning且错误重复

我尝试填充df['total_income']列的空值,逻辑是按age_group、education和income_type分组,用对应分组的中位数填充,但运行代码时收到FutureWarning,且错误信息无限重复。

原代码

import pandas as pd
import numpy as np
df=pd.read_csv(r'C:\Users\gabri\Downloads\credit_scoring_eng.csv')
def fill_na(age_group, education, income_type):
    for i in education:
        for j in income_type:
            for f in age_group:
                df.loc[(df['total_income'].isna()) & (df['education']==i), 'total_income']=df.loc[(df['total_income'].isna())&(df['education']==i)&(df['income_type']==j)&(df['age_group']==f)].median()
    return dff
df['total_income']=fill_na(df['age_group'], df['education'], df['income_type'])
print(df.sort_values(by='total_income', ascending=False).head(numeric_only=True))

错误信息

Output exceeds the size limit. Open the full output data in a text editor
C:\Users\gabri\AppData\Local\Temp\ipykernel_3088\670034715.py:9: FutureWarning: Dropping of nuisance columns in DataFrame reductions (with 'numeric_only=None') is deprecated; in a future version this will raise TypeError.  Select only valid columns before calling the reduction.
  df.loc[(df['total_income'].isna()) & (df['education']==i), 'dob_years']=df.loc[(df['total_income'].isna())&(df['education']==i)&(df['income_type']==j)&(df['age_group']==f)].median()
C:\Users\gabri\AppData\Local\Temp\ipykernel_3088\670034715.py:9: FutureWarning: Dropping of nuisance columns in DataFrame reductions (with 'numeric_only=None') is deprecated; in a future version this will raise TypeError.  Select only valid columns before calling the reduction.
  df.loc[(df['total_income'].isna()) & (df['education']==i), 'dob_years']=df.loc[(df['total_income'].isna())&(df['education']==i)&(df['income_type']==j)&(df['age_group']==f)].median()

问题分析

  1. 三重循环导致重复执行:遍历所有age_group、education、income_type的组合,每次循环都尝试填充,造成错误信息无限重复,且逻辑冗余。
  2. 中位数计算未指定列:调用median()时未指定针对total_income列,pandas会对整个DataFrame做计算,触发FutureWarning(未来版本将强制要求明确numeric_only参数)。
  3. 填充逻辑错误:赋值时仅筛选education==i,未匹配income_type和age_group,不符合分组填充的需求;函数返回未定义的dff变量,会引发额外错误。

解决方案

使用pandas内置的groupby+transform方法,简洁高效地实现分组中位数填充,同时消除警告:

import pandas as pd
import numpy as np

# 读取数据
df = pd.read_csv(r'C:\Users\gabri\Downloads\credit_scoring_eng.csv')

# 按指定分组,用total_income的中位数填充空值
df['total_income'] = df['total_income'].fillna(
    df.groupby(['age_group', 'education', 'income_type'])['total_income'].transform('median')
)

# 查看结果
print(df.sort_values(by='total_income', ascending=False).head(numeric_only=True))

代码说明

  • groupby(['age_group', 'education', 'income_type'])['total_income']:按三个字段分组,仅针对total_income列操作,避免对非数值列计算中位数。
  • transform('median'):对每个分组计算中位数,并将结果映射回原DataFrame对应行的位置,确保空值被对应分组的中位数填充。
  • fillna():将分组计算得到的中位数填充到total_income的空值位置。

兜底方案(针对分组全为空的情况)

如果某些分组的total_income全为空,无法计算中位数,可以用全局中位数兜底:

# 先分组填充,再用全局中位数填充剩余空值
df['total_income'] = df['total_income'].fillna(
    df.groupby(['age_group', 'education', 'income_type'])['total_income'].transform('median')
).fillna(df['total_income'].median())

内容的提问来源于stack exchange,提问作者Riuk2252

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 10:20:41