You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas逐行对比跨列字符串 计算值首次与末次出现月份

实现方案

核心逻辑:逐行筛选出取值不为0的列,直接取列名列表的首尾元素,即可得到对应有效字符串的首次、末次出现月份,无需嵌套np.where或编写多层循环,用pandas内置的逐行应用方法即可高效实现。


完整可运行代码

import pandas as pd

# 构造原始数据集
dictionary = {
    'Month1': ['C1','C2',0,0,'C5'],
    'Month2': ['C1','C2','C3','C4',0],
    'Month3': ['C1','C2','C3','C4',0],
    'Month4': [0,'C2','C3',0,0]
}
df = pd.DataFrame(dictionary)

# 逐行计算起止月份
def calc_month_range(row):
    # 过滤值为0的无效项,保留有效取值对应的月份列名
    valid_cols = row[row != 0].index.tolist()
    # 返回首次、末次出现的月份
    return pd.Series(
        [valid_cols[0], valid_cols[-1]],
        index=['First appear', 'last appear']
    )

# 把计算结果赋值为新列
df[['First appear', 'last appear']] = df.apply(calc_month_range, axis=1)

运行结果

执行代码后打印df,输出和预期结果完全一致:

Month1 Month2 Month3 Month4 First appear last appear
0     C1     C1     C1      0       Month1      Month3
1     C2     C2     C2     C2       Month1      Month4
2      0     C3     C3     C3       Month2      Month4
3      0     C4     C4      0       Month2      Month3
4     C5      0      0      0       Month1      Month1

补充说明

  • 该实现基于当前数据集的特征:每一行除0外仅存在唯一的有效字符串,不存在多个不同非0值,因此过滤0值后直接取列名首尾即可得到正确结果,代码简洁且运行效率高。
  • 如果后续业务场景调整,某一行出现多个不同的非0值,只需要修改calc_month_range函数内的逻辑,对非0值按取值分组后分别统计起止月份即可适配。

内容的提问来源于stack exchange,提问作者Marco Andres Alarcon Sierra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.01 02:39:50