在Pandas中对分组列计算时间差(timedeltas)的实现方法
Calculating Timedeltas by Group in Pandas
No problem! Let's walk through how to compute the time differences between consecutive entries for each person in your DataFrame. Here's a step-by-step solution:
Step 1: Set Up the DataFrame
First, we need to properly construct the DataFrame with both the person and timestamp columns (since your timestamps are tied to specific people). Let's create that:
import pandas as pd import datetime timestamps = [ datetime.datetime(2018, 1, 1, 10, 0, 0, 0), # person 1 datetime.datetime(2018, 1, 1, 10, 0, 0, 0), # person 2 datetime.datetime(2018, 1, 1, 11, 0, 0, 0), # person 2 datetime.datetime(2018, 1, 2, 11, 0, 0, 0), # person 2 datetime.datetime(2018, 1, 1, 10, 0, 0, 0), # person 3 datetime.datetime(2018, 1, 2, 11, 0, 0, 0), # person 3 datetime.datetime(2018, 1, 4, 10, 0, 0, 0), # person 3 datetime.datetime(2018, 1, 5, 12, 0, 0, 0) # person 3 ] # Corresponding person IDs for each timestamp persons = [1, 2, 2, 2, 3, 3, 3, 3] # Create the DataFrame df = pd.DataFrame({'person': persons, 'timestamp': timestamps})
Step 2: Compute Grouped Timedeltas
To calculate the time difference between consecutive entries for each person, we'll:
- Group the DataFrame by the
personcolumn - Use the
diff()method on thetimestampcolumn (this calculates the difference between each row and the prior row in the group) - Assign the result to a new column (e.g.,
time_since_last)
Here's the code:
# Calculate timedeltas within each person group df['time_since_last'] = df.groupby('person')['timestamp'].diff()
Step 3: View the Result
Let's print the resulting DataFrame to see the timedeltas:
print(df)
Output:
person timestamp time_since_last 0 1 2018-01-01 10:00:00 NaT 1 2 2018-01-01 10:00:00 NaT 2 2 2018-01-01 11:00:00 0 days 01:00:00 3 2 2018-01-02 11:00:00 1 days 00:00:00 4 3 2018-01-01 10:00:00 NaT 5 3 2018-01-02 11:00:00 1 days 01:00:00 6 3 2018-01-04 10:00:00 1 days 23:00:00 7 3 2018-01-05 12:00:00 1 days 02:00:00
Key Notes:
NaT(Not a Time) appears for the first entry of each group because there's no prior timestamp to compare against.- The
diff()method preserves the datetime type, so the result is aTimedeltaobject which you can further manipulate (e.g., convert to total seconds, days, etc.):# Convert timedelta to total hours df['hours_since_last'] = df['time_since_last'].dt.total_seconds() / 3600
内容的提问来源于stack exchange,提问作者Michael Dorner
相关产品推荐
相关产品推荐

