如何计算特定事件总时长?微秒级timestamp数据集时长统计需求
解决微秒级时序数据中事件/非事件总时长统计问题
Got it, let's break this down with a practical, pandas-based solution—this is exactly the kind of task pandas was built for, especially with its robust datetime handling.
核心思路
Your dataset has microsecond-precision timestamps and a binary Machining state variable. The key steps are:
- Ensure your timestamps are properly parsed as datetime objects (pandas handles microseconds seamlessly with
datetime64[ns]). - Calculate the time interval between consecutive rows (each row's state lasts until the next timestamp).
- Group these intervals by the
Machiningstate and sum them up.
完整代码示例
First, let's start with a reproducible example—adjust the parsing step to match your actual timestamp format (numeric microseconds, string, etc.):
import pandas as pd # 示例数据集:替换成你的真实数据 data = { 'timestamp': [1620000000000000, 1620000001500000, 1620000003000000, 1620000006000000, 1620000009000000], 'Machining': [0, 1, 0, 1, 0] } df = pd.DataFrame(data) # 1. 转换微秒级timestamp为datetime对象 # 如果你的timestamp是字符串格式(比如"2021-05-03 12:00:00.000000"),直接用pd.to_datetime(df['timestamp'])即可 df['timestamp'] = pd.to_datetime(df['timestamp'], unit='us') # 2. 关键:确保数据按时间戳排序(避免计算错误) df = df.sort_values('timestamp').reset_index(drop=True) # 3. 计算每个状态持续到下一个时间点的时长 # shift(-1)把下一行的时间差对应到当前行的状态 df['state_duration'] = df['timestamp'].diff().shift(-1) # 4. 分组统计总时长 total_machining = df[df['Machining'] == 1]['state_duration'].sum() total_non_machining = df[df['Machining'] == 0]['state_duration'].sum() # 转换为易读的单位(秒、分钟等) print(f"总加工事件时长: {total_machining.total_seconds()} 秒") print(f"总非加工状态时长: {total_non_machining.total_seconds()} 秒")
关键细节说明
- Timestamp Parsing: If your timestamps are stored as strings (e.g.,
2021-05-03 14:30:00.123456), skip theunit='us'parameter—pandas will auto-detect the microseconds. - Handling the Final State: The last row's
state_durationwill beNaNbecause there's no subsequent timestamp. If you need to account for the final state's duration (e.g., until a known end time), add a final row to your DataFrame with that end time and the same state as the last row. - Boolean States: If
Machiningis a boolean column (True/False), just adjust the filter todf[df['Machining'] == True]or simplerdf[df['Machining']].
输出示例
For the sample data above, the output would be:
总加工事件时长: 4.5 秒 总非加工状态时长: 4.5 秒
内容的提问来源于stack exchange,提问作者A Das
相关产品推荐
相关产品推荐

