使用np.where按LotArea随机值填充LotFrontage缺失值报TypeError
问题说明

数据集LotFrontage列和LotArea列存在固定关联:正常取值下,LotFrontage为对应行LotArea值的0.005%~0.01%区间内数值。
需要实现缺失值填充逻辑:LotFrontage为空时,取对应行LotArea*0.005到LotArea*0.01区间的随机值填充。
示例:图中索引1019的行
LotFrontage缺失,对应LotArea为8978,需填充8978*0.005到8978*0.01范围内的随机值。
原有实现与报错
最初编写的代码如下:
np.where(df_train[df_train["LotFrontage"].isnull()], np.random.rand(df_train['LotArea']*0.005, df_train["LotArea"]*0.01),df_train["LotFrontage"])
运行后触发报错:
Error: TypeError Traceback (most recent call last) <ipython-input-46-49a940deebcd> in <module>() ----> 1 np.random.rand(df_train['LotArea'] *0.005,df_train["LotArea"] * 0.01) mtrand.pyx in numpy.random.mtrand.RandomState.rand() mtrand.pyx in numpy.random.mtrand.RandomState.random_sample() _common.pyx in numpy.random._common.double_fill() TypeError: 'Series' object cannot be interpreted as an integer
错误原因
np.random.rand()接收的参数是整数,作用是指定生成随机数组的维度,传入两个pandas Series对象作为维度参数,必然触发类型错误。np.where的判断条件写法错误:传入了过滤缺失值后的子表,和后续待匹配的全量列长度不一致,就算解决了随机数生成的问题也会报维度对齐错误。
修正方案
要生成每行对应独立区间的随机值,正确思路是先生成0-1区间的均匀随机数,再通过线性变换把随机数映射到目标区间,再填充到缺失位置即可。
高效写法(仅对缺失值计算,推荐)
import numpy as np # 标记LotFrontage缺失的行 null_mask = df_train["LotFrontage"].isnull() # 计算缺失行对应的填充区间上下限 fill_low = df_train.loc[null_mask, "LotArea"] * 0.005 fill_high = df_train.loc[null_mask, "LotArea"] * 0.01 # 生成[fill_low, fill_high]区间的随机值 fill_values = np.random.rand(null_mask.sum()) * (fill_high - fill_low) + fill_low # 完成填充 df_train.loc[null_mask, "LotFrontage"] = fill_values
简洁一行写法
如果追求代码简短,也可以用np.where一次性完成,缺点是会给全量行都计算随机值,数据量极大时性能略差:
df_train["LotFrontage"] = np.where( df_train["LotFrontage"].isnull(), np.random.rand(len(df_train)) * df_train["LotArea"] * 0.005 + df_train["LotArea"] * 0.005, df_train["LotFrontage"] )
验证提示:填充完成后可以抽查几行之前的缺失值,确认结果落在对应LotArea计算出的区间范围内即可。
内容的提问来源于stack exchange,提问作者Tanbir Singh
相关产品推荐
相关产品推荐

