pandas 获取每行最大值对应的列名或列序号的实现方法
实现方案
前置依赖
需要提前导入pandas和numpy库:
import pandas as pd import numpy as np
步骤1:构建测试用DataFrame
基于你提供的矩阵数据生成DataFrame,你可以替换成自己的实际数据:
matrix = np.array([[0.92234683, 0.94209485, 0.90884652, 0.99763808], [0.86166401, 0.96755855, 0.9243107 , 0.94240756], [0.85457367, 0.9169915 , 0.95042024, 0.90661279], [0.83972504, 0.93902909, 0.91985442, 0.93765059], [0.84373323, 0.87762977, 0.91005636, 0.88525626]]) # 如果不需要自定义列名可以省略columns参数,默认使用0、1、2、3作为列标识 df = pd.DataFrame(matrix, columns=['col1','col2','col3','col4'])
步骤2:按需求生成结果
需求1:返回每行最大值对应的列序号(从0开始计数)
直接调用numpy的argmax方法性能最优,输出仅含1列的DataFrame:
res = pd.DataFrame(np.argmax(df.values, axis=1), columns=['max_col_index'])
输出结果示例:
max_col_index 0 3 1 1 2 2 3 1 4 2
需求2:返回每行最大值对应的列名
使用pandas自带的idxmax方法,指定axis=1按行查找即可:
res = pd.DataFrame(df.idxmax(axis=1), columns=['max_col_name'])
输出结果示例:
max_col_name 0 col4 1 col2 2 col3 3 col2 4 col3
内容的提问来源于stack exchange,提问作者JFerro
相关产品推荐
相关产品推荐

