带条件的Pandas Series归一化实现求助
解决Pandas中带特殊值的列归一化问题
嘿,这个需求完全可以实现,而且用Pandas就能轻松搞定,不用绕远路!我给你两种方案,一种是分步清晰的写法,另一种是封装成通用函数的简洁版,你可以根据自己的情况选。
先明确核心需求
我们要对score1和score2列做两件事:
- 把值为
-1的替换成缺失值NaN - 其余非负整数归一化到
0-1区间 - 生成新列
norm1和norm2,保留原列不变
方案一:分步实现(适合新手理解)
先构造你的DataFrame:
import pandas as pd df = pd.DataFrame({'key' : [111, 222, 333, 444, 555, 666, 777, 888, 999], 'score1' : [-1, 0, 2, -1, 7, 0, 15, 0, 1], 'score2' : [2, 2, -1, 10, 0, 5, -1, 1, 0]})
处理score1生成norm1
# 1. 复制原列到新列,不修改原数据 df['norm1'] = df['score1'].copy() # 2. 把所有-1替换为NaN df.loc[df['norm1'] == -1, 'norm1'] = pd.NA # 3. 计算有效数据(排除NaN)的最小值和最大值 min_score1 = df['norm1'].dropna().min() max_score1 = df['norm1'].dropna().max() # 4. 对非NaN的值做归一化 df['norm1'] = df['norm1'].apply(lambda x: (x - min_score1)/(max_score1 - min_score1) if pd.notna(x) else x)
同理处理score2生成norm2
df['norm2'] = df['score2'].copy() df.loc[df['norm2'] == -1, 'norm2'] = pd.NA min_score2 = df['norm2'].dropna().min() max_score2 = df['norm2'].dropna().max() df['norm2'] = df['norm2'].apply(lambda x: (x - min_score2)/(max_score2 - min_score2) if pd.notna(x) else x)
查看结果
运行print(df)会得到:
key score1 score2 norm1 norm2 0 111 -1 2 NaN 0.200000 1 222 0 2 0.000000 0.200000 2 333 2 -1 0.133333 NaN 3 444 -1 10 NaN 1.000000 4 555 7 0 0.466667 0.000000 5 666 0 5 0.000000 0.500000 6 777 15 -1 1.000000 NaN 7 888 0 1 0.000000 0.100000 8 999 1 0 0.066667 0.000000
方案二:封装成通用函数(更简洁高效)
如果需要处理多列,重复写代码太麻烦,我们可以把逻辑封装成一个函数,复用性更强,还能处理边界情况(比如所有有效值都相同的情况,避免除以0错误):
def normalize_with_special_value(col, special_val=-1): # 复制原列 new_col = col.copy() # 替换特殊值为NaN new_col[new_col == special_val] = pd.NA # 获取有效数据的min和max col_min = new_col.dropna().min() col_max = new_col.dropna().max() # 处理max和min相等的情况(避免除以0) if col_max == col_min: new_col = new_col.apply(lambda x: 0.0 if pd.notna(x) else x) else: new_col = new_col.apply(lambda x: (x - col_min)/(col_max - col_min) if pd.notna(x) else x) return new_col
调用函数生成新列
df['norm1'] = normalize_with_special_value(df['score1']) df['norm2'] = normalize_with_special_value(df['score2'])
这样一行代码就能处理一列,非常方便!
为什么不用sklearn的MinMaxScaler?
你说得对,MinMaxScaler不适合这里的原因是:它会把-1当成正常数据纳入min/max的计算,这样归一化结果会不符合你的要求。而我们需要先排除-1(换成NaN),再基于有效数据计算归一化,自定义函数能灵活处理这种特殊规则。
内容的提问来源于stack exchange,提问作者glpsx
相关产品推荐
相关产品推荐

