You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas风格方法获取DataFrame每列首个非零元素的值与索引

获取DataFrame每列首个非零元素的值与索引(Pandas风格解法)

嘿,这有个非常贴合Pandas原生风格的简洁解法,完美匹配你的需求!

首先先确认你的示例数据:

import pandas as pd
df = pd.DataFrame([[0, 0, 0], [0, 10, 0], [4, 0, 0], [1, 2, 3]], columns=['first', 'second', 'third'])

输出的原始DataFrame是:

first  second  third
0      0       0      0
1      0      10      0
2      4       0      0
3      1       2      3

接下来用一行核心代码就能搞定,思路是对每一列筛选非零元素后,提取首个值和对应索引:

# 生成目标结果DataFrame
result = df.agg(
    lambda col: pd.Series(
        [col[col != 0].iloc[0], col[col != 0].index[0]] if not col[col != 0].empty else [pd.NA, pd.NA],
        index=['value', 'pos']
    )
).T

print(result)

运行后输出结果(注:你的期望示例中third列的首个非零元素应该是3而非1,大概率是笔误哦):

value  pos
first       4    2
second     10    1
third       3    3

代码细节解释

  • df.agg(...):对DataFrame的每一列执行自定义聚合操作
  • col[col != 0]:筛选当前列的所有非零元素,得到仅包含非零值的子Series
  • .iloc[0]:取子Series的第一个元素(也就是从上到下的首个非零值)
  • .index[0]:取该元素对应的行索引
  • pd.Series(..., index=['value', 'pos']):将值和索引打包成带列名的Series,方便后续整理格式
  • .T:转置结果,让原DataFrame的列名成为新结果的行索引,完全匹配你想要的输出结构
  • 额外处理了全零列的边界情况,这类列会返回pd.NA,避免代码报错

如果你的数据中不存在全零列,还可以简化成更紧凑的版本:

result = df.agg(lambda col: pd.Series([col[col!=0].iloc[0], col[col!=0].index[0]], index=['value','pos'])).T

内容的提问来源于stack exchange,提问作者Konstantin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:57:13