You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现按ID匹配pause/resume时间并计算时间差

Python 整理时间数据并计算时长差

问题需求

给定如下CSV格式的事件数据:

ID, eventName, time
1,pause,2022-10-27T18:15:47Z
1,resume,2022-10-28T08:01:16Z
2,pause,2022-10-26T22:00:01Z
2,resume,2022-10-27T08:00:15Z
3,pause,2022-10-27T04:00:26Z
3,resume,2022-10-27T10:00:20Z
4,pause,2022-10-28T09:09:07Z
4,resume,2022-10-28T07:05:13Z
5,pause,2022-10-27T09:42:14Z
5,resume,2022-10-27T23:01:00Z

需要实现以下逻辑:

  • 按ID分组,将每个ID对应的pause和resume时间整理为time_pause(较早时间)和time_resume(较晚时间)
  • 计算两者的时间差(取整数小时)
  • 输出指定格式的结果:
ID, time_pause, time_resume, time_diff
1,2022-10-27T18:15:47Z,2022-10-28T08:01:16Z,14hr
2,2022-10-26T22:00:01Z,2022-10-27T08:00:15Z,10hr
3,2022-10-27T04:00:26Z,2022-10-27T10:00:20Z,6hr
4,2022-10-28T07:05:13Z,2022-10-28T09:09:07Z,2hr
5,2022-10-27T09:42:14Z,2022-10-27T23:01:00Z,14hr

实现代码

import csv
from datetime import datetime

# 输入数据(可替换为读取文件路径)
input_data = """ID, eventName, time
1,pause,2022-10-27T18:15:47Z
1,resume,2022-10-28T08:01:16Z
2,pause,2022-10-26T22:00:01Z
2,resume,2022-10-27T08:00:15Z
3,pause,2022-10-27T04:00:26Z
3,resume,2022-10-27T10:00:20Z
4,pause,2022-10-28T09:09:07Z
4,resume,2022-10-28T07:05:13Z
5,pause,2022-10-27T09:42:14Z
5,resume,2022-10-27T23:01:00Z"""

# 按ID分组存储时间
id_time_map = {}
reader = csv.DictReader(input_data.splitlines())
for row in reader:
    current_id = row['ID']
    # 转换UTC时间字符串为datetime对象
    event_time = datetime.fromisoformat(row['time'].replace('Z', '+00:00'))
    if current_id not in id_time_map:
        id_time_map[current_id] = []
    id_time_map[current_id].append(event_time)

# 生成结果数据
output = [['ID', 'time_pause', 'time_resume', 'time_diff']]
# 按ID升序处理
for id_str, times in sorted(id_time_map.items(), key=lambda x: int(x[0])):
    # 排序时间,早的在前
    sorted_times = sorted(times)
    # 转换回指定格式的字符串
    pause_time = sorted_times[0].isoformat().replace('+00:00', 'Z')
    resume_time = sorted_times[1].isoformat().replace('+00:00', 'Z')
    # 计算小时差(取整)
    hour_diff = int((sorted_times[1] - sorted_times[0]).total_seconds() // 3600)
    output.append([id_str, pause_time, resume_time, f"{hour_diff}hr"])

# 输出到控制台
print('\n'.join([','.join(line) for line in output]))

# 写入文件(可选)
# with open('result.csv', 'w', newline='') as f:
#     writer = csv.writer(f)
#     writer.writerows(output)

代码说明

  1. 数据读取与分组:用csv.DictReader解析输入数据,将每个ID对应的时间存入字典,实现分组管理。
  2. 时间格式转换:把带Z的UTC时间字符串转为datetime对象,方便后续时间运算。
  3. 时间排序:对每个ID的两个时间进行排序,确保time_pause是较早的时间,不受原始数据中pause/resume的顺序影响。
  4. 时长计算:通过时间差的总秒数除以3600取整,得到整数小时数,格式化为Xhr形式。
  5. 结果输出:支持直接打印到控制台,也可写入CSV文件,完全匹配需求格式。

内容的提问来源于stack exchange,提问作者anup

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 19:05:15