You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas按条件替换所有列除首次出现外的2为0的实现问题咨询

实现方案

核心思路

  • 无需手动指定列名,通过apply对全表所有列批量处理,适配大尺寸DataFrame场景
  • 对每列先筛选出值为2的位置,借助duplicated(keep='first')保留第一个出现的2,其余等于2的位置全部置为0

正确可运行代码

import pandas as pd

# 构建示例数据
df = pd.DataFrame({
    'year_month': ['2018-02', '2018-03', '2018-04', '2018-05', '2018-06', '2018-07', '2018-08'],
    'adoption_1': [0, 0, 1, 1, 1, 2, 2],
    'adoption_2': [0, 0, 0, 1, 2, 2, 2],
    'adoption_3': [0, 1, 1, 1, 1, 2, 2]
})

# 核心处理逻辑
df = df.apply(lambda col: col.mask((col == 2) & col.eq(2).duplicated(keep='first'), 0))

# 设置索引
df = df.set_index('year_month')
print(df)

输出验证

运行后输出与你期望的结果完全一致:

adoption_1  adoption_2  adoption_3
year_month
2018-02              0           0           0
2018-03              0           0           1
2018-04              1           0           1
2018-05              1           1           1
2018-06              1           2           1
2018-07              2           0           2
2018-08              0           0           0

原代码问题说明

你之前尝试的写法存在两处明显问题:

  1. df.where返回的是和原表同尺寸的对象,直接作为布尔索引会触发维度匹配报错
  2. 基于shift的判断逻辑只适配特定数值前置顺序场景,通用性极低,只要2的前置值不符合规则就会失效

内容的提问来源于stack exchange,提问作者Deke Marquardt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 04:09:03