DataFrame切片提取:如何正确提取指定状态周期的子数据集
问题描述
现有如下DataFrame:
time power speed state 1 14.00 29 3 1 2 14.01 30 3 2 3 14.02 29 3 3 4 14.03 30 3 4 5 14.04 29 3 5 6 14.05 30 3 6 7 14.06 29 3 6 8 14.07 30 3 6 9 14.08 29 3 6 10 14.09 30 3 5 11 14.10 29 3 5 12 14.11 30 3 5 13 14.12 29 3 5 14 14.13 30 3 6 15 14.14 31 4 6 16 14.15 32 4 6
需要提取符合以下规则的周期数据,每个周期存为独立DataFrame:
- 周期起始:
state为5,且该5出现在state为6之后(如示例中第10行) - 周期结束:下一次
state为6出现前的最后一行(如示例中第13行)
用户尝试了以下代码,但未得到正确结果:
charge_cycles = [] current_charge_start = None current_drive_start = None total_energy_consumed = 0 drive_data = [] for index, row in data.iterrows(): if row['state'] == '6': if current_drive_start is not None: energy_during_drive = total_energy_consumed charge_cycles.append(energy_during_drive) drive_data.append(data.loc[current_drive_start:index]) current_drive_start = None total_energy_consumed = 0 current_charge_start = row['time'] elif row['state'] == '5': if current_charge_start is not None and current_drive_start is None: current_drive_start = index if current_drive_start is not None: total_energy_consumed += row['power'] * (row['time'] - data.loc[current_drive_start, 'time']) current_drive_start = index # Print the energy consumption during driving between each charge cycle for i, energy in enumerate(charge_cycles, start=1): print(f"Charge Cycle {i}: Energy Consumed During Driving = {energy} units") # Display the DataFrames for each driving cycle for i, drive_df in enumerate(drive_data, start=1): print(f"Driving Cycle {i}:\n{drive_df}")
解决方案
原代码存在几个关键问题:
state判断用了字符串'6',但数据中state是数值类型,导致匹配失败- 周期结束时提取了包含
state=6的行,不符合“到下一次state6出现前结束”的要求 - 能耗计算逻辑错误,
current_drive_start的更新会导致时间差计算重复
以下是修正后的代码:
import pandas as pd # 假设原始数据已加载到data DataFrame中 # 先确保state列为数值类型,避免匹配错误 data['state'] = pd.to_numeric(data['state'], errors='coerce') drive_cycles = [] in_cycle = False cycle_start_idx = None for idx, row in data.iterrows(): # 遇到state=6且当前处于周期中,结束周期(提取到当前行的前一行) if row['state'] == 6 and in_cycle: cycle_df = data.loc[cycle_start_idx:idx-1].copy() drive_cycles.append(cycle_df) in_cycle = False cycle_start_idx = None # 遇到state=5且未处于周期中,检查前一行是否为state=6,满足则启动新周期 elif row['state'] == 5 and not in_cycle: if idx > 0 and data.loc[idx-1, 'state'] == 6: in_cycle = True cycle_start_idx = idx # 处理数据末尾未结束的周期(如果最后是state=5且无后续state=6) if in_cycle and cycle_start_idx is not None: cycle_df = data.loc[cycle_start_idx:].copy() drive_cycles.append(cycle_df) # 输出每个周期的DataFrame for i, cycle in enumerate(drive_cycles, 1): print(f"驾驶周期 {i}:\n{cycle}\n") # 计算并输出每个周期的能耗 for i, cycle in enumerate(drive_cycles, 1): if len(cycle) < 2: energy = 0.0 else: # 计算每行与前一行的时间差 cycle['time_diff'] = cycle['time'].diff() # 累加功率×时间差得到总能耗 energy = (cycle['power'] * cycle['time_diff']).sum() print(f"驾驶周期 {i} 总能耗: {energy:.2f} 单位")
代码说明
- 先统一
state列的数值类型,避免类型不匹配导致的判断错误 - 通过
in_cycle标记当前是否处于周期内,cycle_start_idx记录周期起始索引 - 严格按照规则判断周期的起止:仅当
state=5且前一行是state=6时启动周期;遇到state=6时结束周期,提取到该state=6行的前一行 - 能耗计算采用逐行时间差与功率相乘后累加的方式,结果更准确
内容的提问来源于stack exchange,提问作者Karma_X
相关产品推荐
相关产品推荐

