You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Pandas计算固定起止点等间隔时段内的事件出现次数?

实现过去24小时事件时间线(按小时分组含零值)

没问题,我来帮你搞定这个时间线统计的需求!核心思路是先划定过去24小时的时间窗口,生成每个小时的时间桶,再把事件映射到对应的桶里计数,最后补全零值的桶就行。下面给你两种实用的实现方式:

方式一:纯Python原生实现(无第三方依赖)

这种方式不需要额外安装库,适合轻量场景:

import datetime
from collections import defaultdict

# 你的示例事件数据
items = [
    datetime.datetime(2018, 3, 19, 16, 51, 48),
    datetime.datetime(2018, 3, 19, 17, 25, 19),
    datetime.datetime(2018, 3, 20, 6, 33, 35),
    datetime.datetime(2018, 3, 19, 23, 21, 35),
    datetime.datetime(2018, 3, 19, 15, 8, 41),
    datetime.datetime(2018, 3, 19, 21, 10, 5)
]

# 1. 确定时间窗口:这里以最新事件时间为结束点,往前推24小时
end_time = max(items)
start_time = end_time - datetime.timedelta(hours=24)

# 2. 生成24个小时的时间桶(每个整点)
hourly_buckets = []
current_hour = start_time.replace(minute=0, second=0, microsecond=0)
while current_hour <= end_time:
    hourly_buckets.append(current_hour)
    current_hour += datetime.timedelta(hours=1)

# 3. 统计每个小时的事件数
event_counts = defaultdict(int)
for event_time in items:
    # 将事件时间向下取整到整点,匹配对应的时间桶
    bucket = event_time.replace(minute=0, second=0, microsecond=0)
    # 只统计时间窗口内的事件
    if start_time <= event_time <= end_time:
        event_counts[bucket] += 1

# 4. 生成包含零值的完整时间线
timeline = [(bucket, event_counts.get(bucket, 0)) for bucket in hourly_buckets]

# 打印结果(格式可自定义)
for hour, count in timeline:
    print(f"{hour.strftime('%Y-%m-%d %H:00')}: {count} 次事件")

关键细节说明:

  • 时间窗口可以灵活调整:如果想以当前时间为结束点,把end_time = max(items)改成end_time = datetime.datetime.now()即可。
  • replace(minute=0, ...)是把事件时间对齐到当前小时的整点,确保同一个小时内的事件都归到同一个桶。
  • 用event_counts.get(bucket, 0)确保没有事件的小时自动填充0,不会遗漏。

方式二:用Pandas实现(更简洁高效)

如果你的数据量较大,或者已经在使用Pandas做数据分析,这种方式会更省心:

import pandas as pd
import datetime

items = [
    datetime.datetime(2018, 3, 19, 16, 51, 48),
    datetime.datetime(2018, 3, 19, 17, 25, 19),
    datetime.datetime(2018, 3, 20, 6, 33, 35),
    datetime.datetime(2018, 3, 19, 23, 21, 35),
    datetime.datetime(2018, 3, 19, 15, 8, 41),
    datetime.datetime(2018, 3, 19, 21, 10, 5)
]

# 1. 将事件列表转为Pandas Series
s = pd.Series(items)

# 2. 确定时间窗口
end_time = s.max()
start_time = end_time - pd.Timedelta(hours=24)

# 3. 生成连续的小时级时间索引
hourly_index = pd.date_range(start=start_time, end=end_time, freq='H')

# 4. 分组统计+补零
timeline = s.groupby(s.dt.floor('H')).count().reindex(hourly_index, fill_value=0)

# 打印结果
for hour, count in timeline.items():
    print(f"{hour.strftime('%Y-%m-%d %H:00')}: {count} 次事件")

优势说明:

  • Pandas的date_range可以一键生成连续的时间序列,不用手动循环。
  • reindex(hourly_index, fill_value=0)自动帮我们补全所有零值的小时,代码更简洁。
  • 处理大规模数据时,Pandas的性能比纯Python循环好很多。

内容的提问来源于stack exchange,提问作者Jerad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 12:05:58