You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何MinMaxScaler处理后DataFrame中部分customer_id变为NaN?如何修复?

问题原因与修复方案

为什么会出现customer_id变为NaN?

这是索引对齐错误导致的,核心逻辑是:

  • 用pd.DataFrame(scaled_features, columns=features.columns)创建新DataFrame时,numpy数组转成DataFrame会默认生成从0开始的连续整数索引。
  • 而customer_ids = df_filtered['customer_id']保留了原df_filtered的索引,如果原df_filtered的索引不是连续整数(比如之前做过数据过滤、删除行操作后没重置索引,或是自定义了非连续索引),赋值df_scaled['customer_id'] = customer_ids时,pandas会严格按索引匹配赋值,索引不对应的行就会填充NaN,最终出现部分缺失。

这种情况绝对不正常,属于数据处理中的索引对齐失误,必须修复。

修复方法

方法1:创建scaled DataFrame时保留原索引

生成缩放后的特征DataFrame时,直接指定索引为原features的索引(和customer_ids的索引完全一致),从根源避免错位:

# 分离customer_id和特征列
customer_ids = df_filtered['customer_id']
features = df_filtered.drop('customer_id', axis=1)

# 缩放特征
scaler = MinMaxScaler()
scaled_features = scaler.fit_transform(features)

# 创建DataFrame时指定原索引,确保和customer_ids索引匹配
df_scaled = pd.DataFrame(scaled_features, columns=features.columns, index=features.index)
df_scaled['customer_id'] = customer_ids

# 重排列(可选,匹配原数据列顺序)
df_scaled = df_scaled[['customer_id', 'trx_cnt', 'gtv', 'service_cnt', 'active_day_cnt', 'recency']]

方法2:先重置原数据的索引

如果原df_filtered的索引没有保留价值,可以先重置为连续整数索引,彻底消除对齐问题:

# 重置原数据索引,丢弃旧索引
df_filtered = df_filtered.reset_index(drop=True)

customer_ids = df_filtered['customer_id']
features = df_filtered.drop('customer_id', axis=1)

scaler = MinMaxScaler()
scaled_features = scaler.fit_transform(features)

df_scaled = pd.DataFrame(scaled_features, columns=features.columns)
df_scaled['customer_id'] = customer_ids

df_scaled = df_scaled[['customer_id', 'trx_cnt', 'gtv', 'service_cnt', 'active_day_cnt', 'recency']]

方法3:用pd.concat直接合并(更简洁)

直接将customer_ids和缩放后的特征按列合并,pandas会自动按索引对齐:

customer_ids = df_filtered['customer_id']
features = df_filtered.drop('customer_id', axis=1)

scaler = MinMaxScaler()
scaled_features_df = pd.DataFrame(scaler.fit_transform(features), columns=features.columns, index=features.index)

# 按列合并customer_id与缩放特征
df_scaled = pd.concat([customer_ids, scaled_features_df], axis=1)
# 按需重排列
df_scaled = df_scaled[['customer_id', 'trx_cnt', 'gtv', 'service_cnt', 'active_day_cnt', 'recency']]

验证方法

修复前可以先检查索引是否一致,确认问题根源:

# 打印两者索引,对比是否匹配
print("customer_ids索引:", customer_ids.index)
print("缩放后DataFrame默认索引:", pd.DataFrame(scaled_features, columns=features.columns).index)

修复后执行print(df_scaled['customer_id'].isna().sum()),确认缺失值数量为0即可。

内容的提问来源于stack exchange,提问作者Blaze Tama

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 07:18:37