Pandas生成标记行最大值的布尔DataFrame报错解决方法
pandas逐行判断元素是否为行最大值的实现方法
问题复现
首先构造测试用DataFrame:
import pandas as pd df = pd.DataFrame({'a':[12,34,98,26],'b':[12,87,98,12],'c':[11,23,43,1]})
原始数据预览:
| 索引 | a | b | c |
|---|---|---|---|
| 0 | 12 | 12 | 11 |
| 1 | 34 | 87 | 23 |
| 2 | 98 | 98 | 43 |
| 3 | 26 | 12 | 1 |
需求:生成同形状的布尔DataFramemax_df,元素为所在行最大值时对应位置为True,否则为False,期望输出如下:
| 索引 | a | b | c |
|---|---|---|---|
| 0 | True | True | False |
| 1 | False | True | False |
| 2 | True | True | False |
| 3 | True | False | False |
使用如下代码尝试实现时触发报错:
max_df = df.eq(df.max(axis=1), axis=0)
报错信息:
ValueError: The truth value of an array with more than one element is ambiguous. Use a.any() or a.all()
报错原因
df.max(axis=1)返回的是按行计算最大值的一维Series,索引为原DataFrame的行索引(0/1/2/3)。eq方法的axis参数用于指定对齐维度:
- 传
axis=0(默认值,等价于axis='columns')时,会将传入的序列和DataFrame的列索引对齐广播 - 由于传入的行最大值Series索引是行号,和列名
a/b/c无法匹配,导致比较逻辑异常,触发真值判断报错
正确实现代码
只需要把eq方法的对齐轴改为按行对齐即可,也就是传axis=1(等价于axis='index'):
max_df = df.eq(df.max(axis=1), axis=1)
运行后得到的max_df和预期输出完全一致。
逻辑说明:
- 传入
axis=1后,pandas会将行最大值Series按行索引对齐,把每行的最大值广播到该行所有列,逐元素做相等判断 - 支持同一行存在多个并列最大值的场景(比如第0行a、b列都是12,第2行a、b列都是98,都会被正确标记为True)
内容的提问来源于stack exchange,提问作者User1917931829
相关产品推荐
相关产品推荐

