为何Pandas apply在函数无return时返回NaN而非保留原值?
问题解析与解决方案
为什么会出现NaN?
当使用apply时,Pandas会将Series的每个元素传入函数,函数返回什么就会存入结果的对应位置。你的scale_positive函数在x <= 0时没有显式返回值,Python会隐式返回None。而Pandas的Series无法直接存储None(整数类型的Series会自动转为浮点类型),所以None会被转换为NaN——不是Pandas不保留原值,是你的函数根本没返回原值。
条件转换、其余元素保留原值的惯用方法
方法1:修正自定义函数,显式返回原值
给函数补上分支,确保所有情况都有返回值:
import pandas as pd def scale_positive(x): if x > 0: return x * 10 else: return x # 显式返回原值 s = pd.Series([5, -2, 0, 8, -1]) result = s.apply(scale_positive) print(result)
输出:
0 50 1 -2 2 0 3 80 4 -1 dtype: int64
方法2:使用Pandas原生的where/mask(推荐,矢量操作更高效)
where会保留满足条件的元素,替换不满足的;mask则相反,替换满足条件的元素:
import pandas as pd s = pd.Series([5, -2, 0, 8, -1]) # 保留s<=0的元素,将s>0的替换为s*10 result = s.where(s <= 0, s * 10) # 或用mask:替换s>0的元素为s*10,其余保留 # result = s.mask(s > 0, s * 10) print(result)
输出和方法1一致,且这种矢量运算比apply逐元素处理效率更高,适合大数据量场景。
方法3:结合numpy.where
利用numpy的矢量条件判断实现:
import pandas as pd import numpy as np s = pd.Series([5, -2, 0, 8, -1]) result = pd.Series(np.where(s > 0, s * 10, s), index=s.index) print(result)
针对DataFrame的情况
如果要对DataFrame的某一列做条件转换,方法类似:
import pandas as pd df = pd.DataFrame({'value': [5, -2, 0, 8, -1], 'other': [1,2,3,4,5]}) # 用apply实现 df['value'] = df['value'].apply(lambda x: x*10 if x>0 else x) # 用where实现 df['value'] = df['value'].where(df['value'] <=0, df['value']*10)
内容的提问来源于stack exchange,提问作者Джон Сноу
相关产品推荐
相关产品推荐

