You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中按5分钟时间分箱分组生成会话ID并统计应用?

如何用Pandas按5分钟窗口划分会话并标记SessionID

没问题,我帮你实现这个需求——把带时间戳的数据集按5分钟划分为会话,生成唯一的SessionID,同时统计会话总数和每个会话里的应用。咱们一步步来:

步骤1:准备数据与预处理

首先得确保你的时间戳列是datetime类型,这样才能计算时间差。这里直接用你给出的示例数据构造DataFrame:

import pandas as pd

# 构造示例数据集
data = {
    'timestamp': ['2018-04-08 09:47:57.849', '2018-04-08 09:48:17.573', '2018-04-08 09:48:28.538',
                  '2018-04-08 09:48:37.381', '2018-04-08 09:48:46.680', '2018-04-08 09:48:56.672',
                  '2018-04-08 09:56:58.880', '2018-04-08 09:57:25.461', '2018-04-08 11:28:38.762',
                  '2018-04-08 12:58:31.455', '2018-04-08 14:31:18.131', '2018-04-08 14:31:29.209',
                  '2018-04-08 14:58:42.875', '2018-04-08 18:18:04.757', '2018-04-08 21:08:41.368',
                  '2018-04-11 10:53:10.744', '2018-04-14 19:54:37.441', '2018-04-14 19:54:59.833',
                  '2018-04-14 19:55:10.844', '2018-04-14 19:55:34.486', '2018-04-14 20:23:00.315',
                  '2018-04-15 08:23:44.873', '2018-04-15 08:24:07.257'],
    'App': ['Chrome', 'YouTube', 'Instagram', 'Maps', 'Netflix', 'Google Play Store',
            'Google', 'DB Navigator', 'Google', 'Google', 'Google', 'Google',
            'Google', 'Chrome', 'Google', 'Google', 'Google', 'Google',
            'YouTube', 'Google', 'Google', 'Google', 'Google']
}

df = pd.DataFrame(data)

# 将timestamp列转换为datetime类型,这是计算时间差的前提
df['timestamp'] = pd.to_datetime(df['timestamp'])

步骤2:计算时间差并生成SessionID

核心逻辑是:如果当前记录和上一条记录的时间差超过5分钟,就开启一个新会话。我们用累积求和的方式生成连续的SessionID:

# 计算相邻两条记录的时间差,第一条记录没有前序,差值设为NaN
time_diff = df['timestamp'].diff()

# 标记需要开启新会话的行:时间差>5分钟,或者是第一条记录
new_session = time_diff > pd.Timedelta(minutes=5)
new_session.iloc[0] = True  # 第一条记录默认属于第一个会话

# 对新会话标记做累积求和,得到连续的SessionID
df['SessionID'] = new_session.cumsum()

运行这段代码后,你得到的DataFrame就会和你期望的输出完全一致,每条记录都对应正确的SessionID。

步骤3:统计会话总数与每个会话的应用

现在可以快速提取你需要的统计信息:

# 统计移动会话总数
total_sessions = df['SessionID'].nunique()
print(f"移动会话总数:{total_sessions}")

# 统计每个会话内启动的应用(去重后用逗号分隔展示)
session_apps = df.groupby('SessionID')['App'].unique().reset_index()
session_apps['App'] = session_apps['App'].apply(lambda x: ', '.join(x))
print("\n每个会话内的应用:")
print(session_apps)

执行后你会看到:

  • 会话总数为12(和示例里的最后一个SessionID一致)
  • 每个会话对应的应用列表,比如会话1包含Chrome, YouTube, Instagram, Maps, Netflix, Google Play Store。

内容的提问来源于stack exchange,提问作者Moh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:25:54