循环中使用loc逐行给多列赋值 是否有更简洁紧凑的替代方案?
代码优化方案
方案1:groupby聚合+表合并(推荐,性能最优)
完全抛弃逐行循环,利用pandas向量化运算特性,大数据量下运行效率远高于循环方案:
# 按ID分组计算Height字段的基础统计指标 point_stats = points.groupby('ID')['Height'].agg(['max', 'mean', 'median', 'std']).reset_index() # 计算衍生指标 point_stats['half_std'] = point_stats['std'] / 2 point_stats['max1'] = point_stats['max'] * point_stats['std'] + 15 # 按ID匹配将统计结果合并到polygon表 polygon = polygon.merge(point_stats, on='ID', how='left')
如果polygon的索引就是和points匹配的ID,可以简化赋值步骤:
point_stats = points.groupby('ID')['Height'].agg(['max', 'mean', 'median', 'std']) point_stats['half_std'] = point_stats['std'] / 2 point_stats['max1'] = point_stats['max'] * point_stats['std'] + 15 # 直接按索引对齐批量赋值多列 polygon[['max', 'mean', 'median', 'std', 'half_std', 'max1']] = point_stats
方案2:循环内批量赋值(适配必须保留循环的场景)
如果存在复杂自定义逻辑无法用聚合实现、必须保留逐行循环的,可以用列表批量传值简化重复的loc语句:
for n in polygon.index: points_in_poly = points[points['ID'] == n] # 计算所有变量 maxv = points_in_poly['Height'].max() mean = points_in_poly['Height'].mean() median = points_in_poly['Height'].median() std = points_in_poly['Height'].std() half_std = std / 2 max1 = maxv * std + 15 # 单条loc完成多列赋值,无需重复编写语句 polygon.loc[n, ['max', 'mean', 'median', 'std', 'half_std', 'max1']] = [maxv, mean, median, std, half_std, max1]
内容的提问来源于stack exchange,提问作者g123456k
相关产品推荐
相关产品推荐

