如何基于均值和标准差生成唯一随机值填充pandas列的空单元格
问题根源
你的代码问题出在np.random.randint的调用逻辑上:不指定size参数时,该函数仅会生成1个随机整数标量,调用fillna时所有空单元格都会复用这同一个值,所以才会出现重复填充的情况。
单列修复代码
import pandas as pd import os import numpy as np filename = os.path.join(os.path.dirname(__file__),'exam.csv') data = pd.read_csv(filename) math_mean = data['Math'].mean() math_std = data['Math'].std() # 直接用计算得到的真实上下限,不要手动写死固定值 low = math_mean - math_std high = math_mean + math_std # 统计当前列空值数量 na_count = data['Math'].isna().sum() # 生成和空值数量一致的随机数数组 random_grades = np.random.randint(low=low, high=high, size=na_count) # 仅给空值位置赋值对应的随机数 data.loc[data['Math'].isna(), 'Math'] = random_grades
多列扩展通用方案
如果需要对多个列执行相同的均值±标准差范围内随机填充逻辑,直接遍历目标列即可:
# 填写需要做填充的列名 fill_columns = ['Math', 'English', 'Physics', 'Chemistry'] for col in fill_columns: col_mean = data[col].mean() col_std = data[col].std() col_low = col_mean - col_std col_high = col_mean + col_std # 统计当前列空值数量 col_na_cnt = data[col].isna().sum() if col_na_cnt > 0: # 生成对应数量的随机数 col_random_vals = np.random.randint(low=col_low, high=col_high, size=col_na_cnt) # 填充空值 data.loc[data[col].isna(), col] = col_random_vals
补充说明
如果需要填充浮点型随机数而非整数,把np.random.randint替换为np.random.uniform即可,参数逻辑完全一致。
内容的提问来源于stack exchange,提问作者Robin Sage
相关产品推荐
相关产品推荐

