You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas Cut分箱将NaN作为额外最大分箱时max函数取值异常问题

根因说明

你当前的有序分类定义逻辑本身是正确的:ordered=True的有序分类优先级从左到右递增,放在categories末尾的'NA'理论上应该被判定为最大值。该问题属于部分Pandas版本对多列有序分类取行级max时的兼容问题,要稳定实现需求推荐两种可落地的方案:

方案1:编码映射法(兼容性最优,逻辑完全可控)

利用有序分类自带的cat.codes属性(和分类顺序一一对应的整数编码,值越大优先级越高),取最大编码后反向映射回分类标签,完全不受Pandas版本影响:

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'col1':[10, 22, 25],
    'col2':[11,15,np.nan]
})

bins = [-float('inf'),10,20,30,float("inf")]   
labels = ['Tier 1', 'Tier 2', 'Tier 3', 'Tier 4']
# 统一提前定义分类顺序,保证全链路一致
cat_order = labels + ['NA']

# 分箱+分类转换逻辑
df['col1'] = pd.cut(pd.to_numeric(df['col1'], errors='coerce'), bins=bins, labels=labels)
df['col1'] = pd.Categorical(df['col1'], categories=cat_order, ordered=True)
df['col2'] = pd.cut(pd.to_numeric(df['col2'], errors='coerce'), bins=bins, labels=labels)
df['col2'] = pd.Categorical(df['col2'], categories=cat_order, ordered=True)

# 空白值、NaN统一转成NA
df = df.replace(r'^\s*$', np.nan, regex=True)
df.fillna('NA', inplace=True)

# 逐行取最大值逻辑
df['row_max'] = df[['col1','col2']].apply(
    lambda x: cat_order[max(x.cat.codes)], axis=1
).astype('category').cat.set_categories(cat_order, ordered=True)

print(df)

运行后row_max的输出结果为:Tier 2, Tier 3, NA,完全符合需求。

方案2:前置判断法(性能最优,适合百万级以上大数据量)

先判断行内是否存在'NA',存在则直接返回'NA',不存在再调用原生max计算,避免逐行apply的性能损耗:

has_na = (df[['col1','col2']] == 'NA').any(axis=1)
df['row_max'] = np.where(
    has_na,
    'NA',
    df[['col1','col2']].max(axis=1)
).astype('category').cat.set_categories(cat_order, ordered=True)

额外说明

如果需要后续继续对row_max做有序分类的比较运算,最后一步的astype分类转换不能省略,能保证输出的结果列和原始列的排序规则完全一致。

内容的提问来源于stack exchange,提问作者Shaun

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 19:36:03