Pandas:根据列值条件计算行最大值并保留索引顺序
问题
现有包含col1、col2、col3、col4的Pandas DataFrame,原本通过df[['col1', 'col2', 'col3']].max(axis=1)计算每行最大值,现在需要根据col4的值做条件计算:
- 当
col4为0时,取col1的值 - 当
col4为1时,取col2和col3的最大值
同时要求保留原DataFrame的索引顺序,拆分合并数据框的方法可能打乱索引,求更优方案。
示例代码
import pandas as pd import numpy as np SIZE = 10 df = pd.DataFrame({'col1': np.random.randint(100, size=SIZE), 'col2': np.random.randint(100, size=SIZE), 'col3': np.random.randint(100, size=SIZE), 'col4': np.random.randint(2, size=SIZE)}) print(df)
示例输出
col1 col2 col3 col4 0 55 96 40 0 1 82 59 34 1 2 85 66 25 1 3 90 69 27 0 4 36 32 79 1 5 33 69 80 1 6 11 53 88 0 7 31 51 96 0 8 89 76 88 1 9 4 76 47 0
当前计算结果
0 96 1 82 2 85 3 90 4 79 5 80 6 88 7 96 8 89 9 76 dtype: int64
预期结果
0 55 # col1 1 59 # max(col2, col3) 2 66 # max(col2, col3) 3 90 # col1 4 79 # max(col2, col3) 5 80 # max(col2, col3) 6 11 # col1 7 31 # col1 8 88 # max(col2, col3) 9 4 # col1 dtype: int64
解决方案
可以用numpy.where结合Pandas的行最大值计算实现一行代码,完全保留原索引顺序:
result = pd.Series(np.where(df['col4'] == 0, df['col1'], df[['col2', 'col3']].max(axis=1)), index=df.index)
说明
np.where逐行判断条件:当col4为0时取col1对应值,否则取col2和col3的行最大值- 全程是向量化计算,不需要拆分合并数据框,严格遵循原DataFrame的索引顺序,效率远高于循环或拆分合并操作
- 如果要直接将结果添加到原DataFrame,同样一行完成:
df['result'] = np.where(df['col4'] == 0, df['col1'], df[['col2', 'col3']].max(axis=1))
内容的提问来源于stack exchange,提问作者thenewjames
相关产品推荐
相关产品推荐

