Python pandas:正确识别dm_it完成1→0→1转换的ID问题排查
问题:如何让ID5被正确纳入目标ID列表?
需求是找出按date_collected排序后,dm_it字段经历1→0→1转变的ID,预期结果为[2,4,5],但当前代码仅返回[2,4]。
当前代码
import pandas as pd data = {'date_collected': ['2023-09-01', '2023-09-02', '2023-09-03', '2023-09-04', '2023-09-05', '2023-09-01', '2023-09-02', '2023-09-03', '2023-09-04', '2023-09-05', '2023-09-01', '2023-09-02', '2023-09-03', '2023-09-04', '2023-09-05', '2023-09-01', '2023-09-02', '2023-09-03', '2023-09-04', '2023-09-05', '2023-09-01', '2023-09-02', '2023-09-03', '2023-09-04', '2023-09-05'], 'id': [1, 1, 1, 1, 1, 2, 2, 2, 2, 2, 3, 3, 3, 3, 3, 4, 4, 4, 4, 4, 5, 5, 5, 5, 5], 'dm_it': [0, 1, 1, 1, 0, 1, 1, 0, 1, 1, 0, 1, 1, 1, 1, 0, 1, 0, 0, 1, 1, 1, 1, 0, 1]} df = pd.DataFrame(data) dm_off_group = {} dm_off_date = {} dm_transition_count = {} current_state = None for index, row in df.iterrows(): if row['id'] not in dm_off_group: dm_off_group[row['id']] = 0 dm_off_date[row['id']] = None dm_transition_count[row['id']] = 0 if row['dm_it'] == 1: dm_off_group[row['id']] = 1 if row['dm_it'] == 0 and dm_off_group[row['id']] == 1 and dm_off_date[row['id']] is None: dm_off_date[row['id']] = row['date_collected'] if row['dm_it'] == 1 and dm_off_group[row['id']] == 1: if current_state is None or current_state == 0: dm_transition_count[row['id']] += 1 current_state = 1 elif row['dm_it'] == 0: current_state = 0 dm_transition_ids = [id for id, count in dm_transition_count.items() if count >= 2] print("IDs with more than two transitions: ", dm_transition_ids) print(df)
当前输出
IDs with more than two transitions: [2, 4] date_collected id dm_it 0 2023-09-01 1 0 1 2023-09-02 1 1 2 2023-09-03 1 1 3 2023-09-04 1 1 4 2023-09-05 1 0 5 2023-09-01 2 1 6 2023-09-02 2 1 7 2023-09-03 2 0 8 2023-09-04 2 1 9 2023-09-05 2 1 10 2023-09-01 3 0 11 2023-09-02 3 1 12 2023-09-03 3 1 13 2023-09-04 3 1 14 2023-09-05 3 1 15 2023-09-01 4 0 16 2023-09-02 4 1 17 2023-09-03 4 0 18 2023-09-04 4 0 19 2023-09-05 4 1 20 2023-09-01 5 1 21 2023-09-02 5 1 22 2023-09-03 5 1 23 2023-09-04 5 0 24 2023-09-05 5 1
问题原因
核心问题是全局的current_state变量被所有ID共用,导致处理ID5时,状态判断被之前的ID(比如ID4)的最终状态干扰:
- ID4的最后一行
dm_it=1,将current_state设为1 - 处理ID5的前几行
dm_it=1时,current_state已经是1,不会触发dm_transition_count的增加 - ID5的
1→0→1转变中,从0回到1的步骤本应触发计数,但因为初始状态被全局变量污染,最终计数不足2,无法被纳入结果
修改方案
将current_state改为按ID维护的字典,确保每个ID的状态独立:
修改后的代码
import pandas as pd data = {'date_collected': ['2023-09-01', '2023-09-02', '2023-09-03', '2023-09-04', '2023-09-05', '2023-09-01', '2023-09-02', '2023-09-03', '2023-09-04', '2023-09-05', '2023-09-01', '2023-09-02', '2023-09-03', '2023-09-04', '2023-09-05', '2023-09-01', '2023-09-02', '2023-09-03', '2023-09-04', '2023-09-05', '2023-09-01', '2023-09-02', '2023-09-03', '2023-09-04', '2023-09-05'], 'id': [1, 1, 1, 1, 1, 2, 2, 2, 2, 2, 3, 3, 3, 3, 3, 4, 4, 4, 4, 4, 5, 5, 5, 5, 5], 'dm_it': [0, 1, 1, 1, 0, 1, 1, 0, 1, 1, 0, 1, 1, 1, 1, 0, 1, 0, 0, 1, 1, 1, 1, 0, 1]} df = pd.DataFrame(data) dm_off_group = {} dm_off_date = {} dm_transition_count = {} # 改为字典,每个ID维护自己的当前状态 current_states = {} for index, row in df.iterrows(): current_id = row['id'] if current_id not in dm_off_group: dm_off_group[current_id] = 0 dm_off_date[current_id] = None dm_transition_count[current_id] = 0 # 初始化当前ID的状态为None current_states[current_id] = None if row['dm_it'] == 1: dm_off_group[current_id] = 1 if row['dm_it'] == 0 and dm_off_group[current_id] == 1 and dm_off_date[current_id] is None: dm_off_date[current_id] = row['date_collected'] if row['dm_it'] == 1 and dm_off_group[current_id] == 1: # 使用当前ID对应的状态 if current_states[current_id] is None or current_states[current_id] == 0: dm_transition_count[current_id] += 1 current_states[current_id] = 1 elif row['dm_it'] == 0: current_states[current_id] = 0 dm_transition_ids = [id for id, count in dm_transition_count.items() if count >= 2] print("IDs with more than two transitions: ", dm_transition_ids) print(df)
修改后输出
IDs with more than two transitions: [2, 4, 5] ...(后续DataFrame输出与原输出一致)
关键修改点
- 将全局
current_state替换为字典current_states,每个ID对应独立的状态值 - 在初始化新ID时,同时初始化该ID的状态为
None - 所有状态判断和更新操作,都使用当前ID对应的
current_states[current_id],而非全局变量
这样每个ID的状态转变都会被独立追踪,ID5的1→0→1转变会被正确计数,最终被纳入结果列表。
内容的提问来源于stack exchange,提问作者B.Le
相关产品推荐
相关产品推荐

