如何基于区域条件为Pandas DataFrame生成正态分布样本列
解决方案
步骤1:提取区域编号
先把df_conditions里的REGION_NAME转换成和df_avgs中REGION匹配的数字格式:
# 从"Region X"格式中提取数字并转为整数 df_conditions['REGION'] = df_conditions['REGION_NAME'].str.extract(r'(\d+)').astype(int)
步骤2:关联均值与标准差数据
将两个DataFrame按REGION列合并,让df_conditions的每行都能获取对应区域的均值和标准差:
# 以左连接保留df_conditions的所有行 merged_df = df_conditions.merge(df_avgs, on='REGION', how='left')
步骤3:批量生成正态分布样本
基于合并后的均值和标准差列,批量生成对应每行的随机样本并添加为新列:
import numpy as np # 生成与DataFrame长度一致的样本列 merged_df['SAMPLE'] = np.random.normal( loc=merged_df['AVERAGE'], scale=merged_df['STDEV'], size=len(merged_df) ) # 若只需给原df_conditions新增列,直接赋值即可 df_conditions['SAMPLE'] = merged_df['SAMPLE']
简化版流程(无需保留中间列)
如果不需要保留AVERAGE和STDEV列,可以用字典映射简化操作:
# 将df_avgs的均值、标准差转成按区域索引的字典 avg_map = df_avgs.set_index('REGION')['AVERAGE'].to_dict() std_map = df_avgs.set_index('REGION')['STDEV'].to_dict() # 提取区域编号 df_conditions['REGION'] = df_conditions['REGION_NAME'].str.extract(r'(\d+)').astype(int) # 生成样本并添加列 df_conditions['SAMPLE'] = np.random.normal( loc=df_conditions['REGION'].map(avg_map), scale=df_conditions['REGION'].map(std_map), size=len(df_conditions) ) # 可选:删除临时生成的REGION列 df_conditions.drop('REGION', axis=1, inplace=True)
最终效果
处理后的df_conditions会新增一列SAMPLE,每行对应其REGION_NAME所属区域的正态分布随机样本,长度与原DataFrame完全一致。
内容的提问来源于stack exchange,提问作者granger
相关产品推荐
相关产品推荐

