如何基于行索引统计量替换DataFrame中的数值?
按行均值规则替换DataFrame数值
原始数据构建
用户的DataFrame构建代码如下:
import pandas as pd l1 = [1,2,3,4,5,6,7,8,9,10] l2 = [11,12,13,14,15,16,17,18,19,20] index = ['FORD','GM'] df = pd.DataFrame(l1,l2).reset_index().T df.index = index
需求说明
对每行数据执行替换规则:
- 若数值小于该行均值减2,替换为
'MINI' - 否则替换为
'MEGA'
实现方案
方法1:广播运算(推荐,性能更优)
通过计算每行阈值,利用广播实现元素级对比替换:
# 计算每行的阈值(均值减2) row_thresholds = df.mean(axis=1) - 2 # 两次where实现双条件替换 result_df = df.where(df >= row_thresholds[:, None], 'MINI').where(df < row_thresholds[:, None], 'MEGA')
方法2:逐行apply处理
自定义函数处理每行数据:
def process_row(row): threshold = row.mean() - 2 return row.map(lambda x: 'MINI' if x < threshold else 'MEGA') result_df = df.apply(process_row, axis=1)
输出结果
最终result_df的内容与期望一致:
| 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | |
|---|---|---|---|---|---|---|---|---|---|---|
| FORD | MINI | MINI | MINI | MEGA | MEGA | MEGA | MEGA | MEGA | MEGA | MEGA |
| GM | MINI | MINI | MINI | MEGA | MEGA | MEGA | MEGA | MEGA | MEGA | MEGA |
内容的提问来源于stack exchange,提问作者Yash
相关产品推荐
相关产品推荐

