You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python pandas:正确识别dm_it完成1→0→1转换的ID问题排查

问题:如何让ID5被正确纳入目标ID列表?

需求是找出按date_collected排序后,dm_it字段经历1→0→1转变的ID,预期结果为[2,4,5],但当前代码仅返回[2,4]。

当前代码

import pandas as pd

data = {'date_collected': ['2023-09-01', '2023-09-02', '2023-09-03', '2023-09-04', '2023-09-05', '2023-09-01', '2023-09-02', '2023-09-03', '2023-09-04', '2023-09-05', '2023-09-01', '2023-09-02', '2023-09-03', '2023-09-04', '2023-09-05', '2023-09-01', '2023-09-02', '2023-09-03', '2023-09-04', '2023-09-05', '2023-09-01', '2023-09-02', '2023-09-03', '2023-09-04', '2023-09-05'],
        'id': [1, 1, 1, 1, 1, 2, 2, 2, 2, 2, 3, 3, 3, 3, 3, 4, 4, 4, 4, 4, 5, 5, 5, 5, 5],
        'dm_it': [0, 1, 1, 1, 0, 1, 1, 0, 1, 1, 0, 1, 1, 1, 1, 0, 1, 0, 0, 1, 1, 1, 1, 0, 1]}

df = pd.DataFrame(data)


dm_off_group = {}
dm_off_date = {}
dm_transition_count = {}
current_state = None

for index, row in df.iterrows():
    if row['id'] not in dm_off_group:
        dm_off_group[row['id']] = 0
        dm_off_date[row['id']] = None
        dm_transition_count[row['id']] = 0
    
    if row['dm_it'] == 1:
        dm_off_group[row['id']] = 1
    
    if row['dm_it'] == 0 and dm_off_group[row['id']] == 1 and dm_off_date[row['id']] is None:
        dm_off_date[row['id']] = row['date_collected']
    
    if row['dm_it'] == 1 and dm_off_group[row['id']] == 1:
        if current_state is None or current_state == 0:
            dm_transition_count[row['id']] += 1
        current_state = 1
    elif row['dm_it'] == 0:
        current_state = 0

dm_transition_ids = [id for id, count in dm_transition_count.items() if count >= 2]

print("IDs with more than two transitions: ", dm_transition_ids)
print(df)

当前输出

IDs with more than two transitions:  [2, 4]
   date_collected  id  dm_it
0      2023-09-01   1      0
1      2023-09-02   1      1
2      2023-09-03   1      1
3      2023-09-04   1      1
4      2023-09-05   1      0
5      2023-09-01   2      1
6      2023-09-02   2      1
7      2023-09-03   2      0
8      2023-09-04   2      1
9      2023-09-05   2      1
10     2023-09-01   3      0
11     2023-09-02   3      1
12     2023-09-03   3      1
13     2023-09-04   3      1
14     2023-09-05   3      1
15     2023-09-01   4      0
16     2023-09-02   4      1
17     2023-09-03   4      0
18     2023-09-04   4      0
19     2023-09-05   4      1
20     2023-09-01   5      1
21     2023-09-02   5      1
22     2023-09-03   5      1
23     2023-09-04   5      0
24     2023-09-05   5      1

问题原因

核心问题是全局的current_state变量被所有ID共用,导致处理ID5时,状态判断被之前的ID(比如ID4)的最终状态干扰:

  • ID4的最后一行dm_it=1,将current_state设为1
  • 处理ID5的前几行dm_it=1时,current_state已经是1,不会触发dm_transition_count的增加
  • ID5的1→0→1转变中,从0回到1的步骤本应触发计数,但因为初始状态被全局变量污染,最终计数不足2,无法被纳入结果

修改方案

将current_state改为按ID维护的字典,确保每个ID的状态独立:

修改后的代码

import pandas as pd

data = {'date_collected': ['2023-09-01', '2023-09-02', '2023-09-03', '2023-09-04', '2023-09-05', '2023-09-01', '2023-09-02', '2023-09-03', '2023-09-04', '2023-09-05', '2023-09-01', '2023-09-02', '2023-09-03', '2023-09-04', '2023-09-05', '2023-09-01', '2023-09-02', '2023-09-03', '2023-09-04', '2023-09-05', '2023-09-01', '2023-09-02', '2023-09-03', '2023-09-04', '2023-09-05'],
        'id': [1, 1, 1, 1, 1, 2, 2, 2, 2, 2, 3, 3, 3, 3, 3, 4, 4, 4, 4, 4, 5, 5, 5, 5, 5],
        'dm_it': [0, 1, 1, 1, 0, 1, 1, 0, 1, 1, 0, 1, 1, 1, 1, 0, 1, 0, 0, 1, 1, 1, 1, 0, 1]}

df = pd.DataFrame(data)


dm_off_group = {}
dm_off_date = {}
dm_transition_count = {}
# 改为字典,每个ID维护自己的当前状态
current_states = {}

for index, row in df.iterrows():
    current_id = row['id']
    if current_id not in dm_off_group:
        dm_off_group[current_id] = 0
        dm_off_date[current_id] = None
        dm_transition_count[current_id] = 0
        # 初始化当前ID的状态为None
        current_states[current_id] = None
    
    if row['dm_it'] == 1:
        dm_off_group[current_id] = 1
    
    if row['dm_it'] == 0 and dm_off_group[current_id] == 1 and dm_off_date[current_id] is None:
        dm_off_date[current_id] = row['date_collected']
    
    if row['dm_it'] == 1 and dm_off_group[current_id] == 1:
        # 使用当前ID对应的状态
        if current_states[current_id] is None or current_states[current_id] == 0:
            dm_transition_count[current_id] += 1
        current_states[current_id] = 1
    elif row['dm_it'] == 0:
        current_states[current_id] = 0

dm_transition_ids = [id for id, count in dm_transition_count.items() if count >= 2]

print("IDs with more than two transitions: ", dm_transition_ids)
print(df)

修改后输出

IDs with more than two transitions:  [2, 4, 5]
...(后续DataFrame输出与原输出一致)

关键修改点

  1. 将全局current_state替换为字典current_states,每个ID对应独立的状态值
  2. 在初始化新ID时,同时初始化该ID的状态为None
  3. 所有状态判断和更新操作,都使用当前ID对应的current_states[current_id],而非全局变量

这样每个ID的状态转变都会被独立追踪,ID5的1→0→1转变会被正确计数,最终被纳入结果列表。

内容的提问来源于stack exchange,提问作者B.Le

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 23:12:10