You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Pandas中对分组列计算时间差(timedeltas)的实现方法

Calculating Timedeltas by Group in Pandas

No problem! Let's walk through how to compute the time differences between consecutive entries for each person in your DataFrame. Here's a step-by-step solution:

Step 1: Set Up the DataFrame

First, we need to properly construct the DataFrame with both the person and timestamp columns (since your timestamps are tied to specific people). Let's create that:

import pandas as pd
import datetime

timestamps = [
    datetime.datetime(2018, 1, 1, 10, 0, 0, 0), # person 1
    datetime.datetime(2018, 1, 1, 10, 0, 0, 0), # person 2
    datetime.datetime(2018, 1, 1, 11, 0, 0, 0), # person 2
    datetime.datetime(2018, 1, 2, 11, 0, 0, 0), # person 2
    datetime.datetime(2018, 1, 1, 10, 0, 0, 0), # person 3
    datetime.datetime(2018, 1, 2, 11, 0, 0, 0), # person 3
    datetime.datetime(2018, 1, 4, 10, 0, 0, 0), # person 3
    datetime.datetime(2018, 1, 5, 12, 0, 0, 0) # person 3
]

# Corresponding person IDs for each timestamp
persons = [1, 2, 2, 2, 3, 3, 3, 3]

# Create the DataFrame
df = pd.DataFrame({'person': persons, 'timestamp': timestamps})

Step 2: Compute Grouped Timedeltas

To calculate the time difference between consecutive entries for each person, we'll:

  • Group the DataFrame by the person column
  • Use the diff() method on the timestamp column (this calculates the difference between each row and the prior row in the group)
  • Assign the result to a new column (e.g., time_since_last)

Here's the code:

# Calculate timedeltas within each person group
df['time_since_last'] = df.groupby('person')['timestamp'].diff()

Step 3: View the Result

Let's print the resulting DataFrame to see the timedeltas:

print(df)

Output:

person           timestamp time_since_last
0       1 2018-01-01 10:00:00             NaT
1       2 2018-01-01 10:00:00             NaT
2       2 2018-01-01 11:00:00        0 days 01:00:00
3       2 2018-01-02 11:00:00        1 days 00:00:00
4       3 2018-01-01 10:00:00             NaT
5       3 2018-01-02 11:00:00        1 days 01:00:00
6       3 2018-01-04 10:00:00        1 days 23:00:00
7       3 2018-01-05 12:00:00        1 days 02:00:00

Key Notes:

  • NaT (Not a Time) appears for the first entry of each group because there's no prior timestamp to compare against.
  • The diff() method preserves the datetime type, so the result is a Timedelta object which you can further manipulate (e.g., convert to total seconds, days, etc.):
    # Convert timedelta to total hours
    df['hours_since_last'] = df['time_since_last'].dt.total_seconds() / 3600
    

内容的提问来源于stack exchange,提问作者Michael Dorner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 06:28:26