You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

带条件的Pandas Series归一化实现求助

解决Pandas中带特殊值的列归一化问题

嘿,这个需求完全可以实现,而且用Pandas就能轻松搞定,不用绕远路!我给你两种方案,一种是分步清晰的写法,另一种是封装成通用函数的简洁版,你可以根据自己的情况选。

先明确核心需求

我们要对score1和score2列做两件事:

  • 把值为-1的替换成缺失值NaN
  • 其余非负整数归一化到0-1区间
  • 生成新列norm1和norm2,保留原列不变

方案一:分步实现(适合新手理解)

先构造你的DataFrame:

import pandas as pd

df = pd.DataFrame({'key' : [111, 222, 333, 444, 555, 666, 777, 888, 999], 
                   'score1' : [-1, 0, 2, -1, 7, 0, 15, 0, 1], 
                   'score2' : [2, 2, -1, 10, 0, 5, -1, 1, 0]})

处理score1生成norm1

# 1. 复制原列到新列,不修改原数据
df['norm1'] = df['score1'].copy()
# 2. 把所有-1替换为NaN
df.loc[df['norm1'] == -1, 'norm1'] = pd.NA
# 3. 计算有效数据(排除NaN)的最小值和最大值
min_score1 = df['norm1'].dropna().min()
max_score1 = df['norm1'].dropna().max()
# 4. 对非NaN的值做归一化
df['norm1'] = df['norm1'].apply(lambda x: (x - min_score1)/(max_score1 - min_score1) if pd.notna(x) else x)

同理处理score2生成norm2

df['norm2'] = df['score2'].copy()
df.loc[df['norm2'] == -1, 'norm2'] = pd.NA
min_score2 = df['norm2'].dropna().min()
max_score2 = df['norm2'].dropna().max()
df['norm2'] = df['norm2'].apply(lambda x: (x - min_score2)/(max_score2 - min_score2) if pd.notna(x) else x)

查看结果

运行print(df)会得到:

key  score1  score2     norm1     norm2
0  111      -1       2       NaN  0.200000
1  222       0       2  0.000000  0.200000
2  333       2      -1  0.133333       NaN
3  444      -1      10       NaN  1.000000
4  555       7       0  0.466667  0.000000
5  666       0       5  0.000000  0.500000
6  777      15      -1  1.000000       NaN
7  888       0       1  0.000000  0.100000
8  999       1       0  0.066667  0.000000

方案二:封装成通用函数(更简洁高效)

如果需要处理多列,重复写代码太麻烦,我们可以把逻辑封装成一个函数,复用性更强,还能处理边界情况(比如所有有效值都相同的情况,避免除以0错误):

def normalize_with_special_value(col, special_val=-1):
    # 复制原列
    new_col = col.copy()
    # 替换特殊值为NaN
    new_col[new_col == special_val] = pd.NA
    # 获取有效数据的min和max
    col_min = new_col.dropna().min()
    col_max = new_col.dropna().max()
    # 处理max和min相等的情况(避免除以0)
    if col_max == col_min:
        new_col = new_col.apply(lambda x: 0.0 if pd.notna(x) else x)
    else:
        new_col = new_col.apply(lambda x: (x - col_min)/(col_max - col_min) if pd.notna(x) else x)
    return new_col

调用函数生成新列

df['norm1'] = normalize_with_special_value(df['score1'])
df['norm2'] = normalize_with_special_value(df['score2'])

这样一行代码就能处理一列,非常方便!


为什么不用sklearn的MinMaxScaler?

你说得对,MinMaxScaler不适合这里的原因是:它会把-1当成正常数据纳入min/max的计算,这样归一化结果会不符合你的要求。而我们需要先排除-1(换成NaN),再基于有效数据计算归一化,自定义函数能灵活处理这种特殊规则。

内容的提问来源于stack exchange,提问作者glpsx

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:41:24