You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Pandas中按Session分组取Datetime列最后值并计算预期结束时间

解决Pandas按Session分组并计算期望结束时间的问题

Hey there, let's tackle this Pandas grouping and calculation problem step by step.

First, let's lay out your original DataFrame clearly:

import pandas as pd

# 原始输入数据
data = {
    'Doctor': ['A']*12,
    'Start': ['2020-01-18 12:00:00', '2020-01-18 12:30:00', '2020-01-18 13:00:00', '2020-01-18 13:00:00',
              '2020-01-18 13:30:00', '2020-01-18 14:00:00', '2020-01-18 14:00:00', '2020-01-18 14:30:00',
              '2020-01-18 14:30:00', '2020-01-19 12:00:00', '2020-01-19 12:30:00', '2020-01-19 14:00:00'],
    'B_ID': [1,2,3,4,5,6,7,8,9,12,13,14],
    'Session': ['S1']*5 + ['S3']*4 + ['S2']*3,
    'Finish': ['2020-01-18 12:33:00', '2020-01-18 12:52:00', '2020-01-18 13:23:00', '2020-01-18 13:37:00',
               '2020-01-18 13:56:00', '2020-01-18 14:15:00', '2020-01-18 14:28:00', '2020-01-18 14:40:00',
               '2020-01-18 15:01:00', '2020-01-19 12:20:00', '2020-01-19 12:40:00', '2020-01-19 14:20:00']
}

df = pd.DataFrame(data)

Your goal is to group by the Session column, extract the last Start and Finish times for each group, then add an expected_finish column equal to the group's last_start plus 30 minutes. Here's how to do it:

Step 1: Convert time columns to datetime type

First, we need to convert the string-formatted time columns to proper datetime objects—otherwise we can't perform time-based calculations:

df['Start'] = pd.to_datetime(df['Start'])
df['Finish'] = pd.to_datetime(df['Finish'])

Step 2: Group by Session and aggregate last values

Use groupby() and agg() to pull the final Start and Finish entries for each Session:

grouped_df = df.groupby('Session').agg(
    last_start=('Start', 'last'),
    last_finish=('Finish', 'last')
).reset_index()

Step 3: Calculate the expected_finish column

Add 30 minutes to each last_start using pd.Timedelta:

grouped_df['expected_finish'] = grouped_df['last_start'] + pd.Timedelta(minutes=30)

Final Output

If you print grouped_df, you'll get exactly the result you wanted:

Sessionlast_startlast_finishexpected_finish
S12020-01-18 13:30:002020-01-18 13:56:002020-01-18 14:00:00
S22020-01-19 14:00:002020-01-19 14:20:002020-01-19 14:30:00
S32020-01-18 14:30:002020-01-18 15:01:002020-01-18 15:00:00

内容的提问来源于stack exchange,提问作者Danish

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 10:22:36